The integrator state was already device only, but some other memory is
also frequently accessed and should be on the GPU for best performance.
This is a follow up for a pre-existing issue found reviewing #163437 and
#163930.
Pull Request: https://projects.blender.org/blender/blender/pulls/164306
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.
Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.
The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.
For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.
Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.
This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.
The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from
build_files/config/pipeline_config.yaml
to
build_files/buildbot/config/blender_version.cmake
Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.
Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:
- There is an easy to follow table of per-architecture or family
toolkits.
- If there is no architectural preference, all provided toolkits are
probed. For SM kernels newer toolkits are tested first, and for
COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
default one is used.
A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).
Pull Request: https://projects.blender.org/blender/blender/pulls/164002
When a GPU device's driver does not meet Blender's minimum required
version, the device is now shown in the preferences as a greyed-out
entry with a message inside the brackets, indicating which driver
version is needed, instead of being silently hidden.
This helps users understand why their GPU is not available for
rendering and what action they can take to resolve the situation.
Pull Request: https://projects.blender.org/blender/blender/pulls/159405
Unfortunately it appears that moving MNEE to another kernel did not
fundamentally fix the apparent compiler bug that breaks this. Another
refactor in 0baa98866c made the bug surface again.
It appears to work fine with HIP-RT, so we leave that case enabled.
HIP-RT is also enabled by default, so it's not as bad.
Pull Request: https://projects.blender.org/blender/blender/pulls/160110
Delay device side reallocation of image_info until load_image_info(),
and add DeviceQueue function to do this on the queue. This will allow
multiple GPU devices asynchronously update their image_info so that
one GPU does not have to stop kernel execution while another GPU
allocates a new image tile.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
Allow updating image_info while rendering for CPU devices, by replacing
the pointer in all per-thread kernel globals. The per-thread kernel globals
are now owned by the device to make this possible.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
Previously a string needed to be stored for each image texture, but we
might as well compute this on demand and avoid the overhead. There is
now a separate log_name() for the logging, and global_name() for
copying to GPU kernel global variable.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
There is a strong coupling between the API version an application
passes to the hiprtCreateContext() and the HIP-RT library: if there
is a mismatch the context will fail to be created.
Additionally, the HIP-RT library does not have static dependencies
as it loads HIP SDK dynamically.
It all makes the wrangler be not so useful, and harmful in cases
when Blender is packaged for different Linux distros which might
use HIP-RT library of a different version from what Cycles is
currently expecting via the bundled hiprtew.h.
The HIPRT_API_VERSION is now coming from the hiprt/hiprt.h so it
is guaranteed that Cycles will request the same API version as
the library is was compiled with.
The slightly annoying part is that this approach requires the
dynamic libraries to have API version in the name of the library
to function properly This is because of the SONAME which still
mentions the original library name. The similar thing happens on
Windows where the hiprt0200564.lib points to the hiprt0200564.dll.
There might be some way to do the rename during install, but it
is a bit tricky, especially due to the manifest on Windows. It
might be the easiest to simply rename precompiled libraries.
Pull Request: https://projects.blender.org/blender/blender/pulls/153892
Use lazy initialization for global variables, so they get destructed before the
guardedalloc leak check and destruction. The same pattern is used elsewhere
in Blender.
There may be more cases, these were the ones I found in testing. These issue
were exposed by f8eec542f4.
Pull Request: https://projects.blender.org/blender/blender/pulls/154173
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
The workaround of forcing BVH building into single thread
execution on the Blender side is not needed anymore,
because the problem was properly fixed in the upstream
since Embree upgrade in Blender 4.5
This reverts commit c0f0e2ca6f.
Pull Request: https://projects.blender.org/blender/blender/pulls/146859
There is a bug in Embree that makes BVH updates crash. Disabling multithreaded
BVH updates after the initial BVH build appears to work around it, at the cost
of some performance.
This will not affect performance of the initial BVH build, transforming objects
or editing a single mesh. It will only affect performance when multiple smaller
meshes are edited together, as those can no longer have their BVH updated in
parallel or benefit from parallellization over many primitives.
Pull Request: https://projects.blender.org/blender/blender/pulls/134747
This may be used for device to do host memory allocation in a way that
is more efficient for copy the host memory to the device.
Also rename and group device memory allocation functions for clarity.
Pull Request: https://projects.blender.org/blender/blender/pulls/134412
Regardless of what mem info reports. We can't move this to the host, so
might as well try because the free memory might not be a reliable predictor
of success.
Pull Request: https://projects.blender.org/blender/blender/pulls/132912
With host mapped memory these can be shared, and we can't get back the
original host pointer unless we make a copy which is inefficient.
Also add asserts to verify this doesn't happen.
Pull Request: https://projects.blender.org/blender/blender/pulls/132912
This was not thread safe. And it's better to do them one by one to avoid
moving more than is needed, when another thread already freed up enough.
Thanks to Jorn Visser for investigating and finding this problem.
Pull Request: https://projects.blender.org/blender/blender/pulls/132912
Should be a bit more efficient, and it fixes host memory fallback bugs,
where host memory was incorrectly freed during re-copy. For the case
where memory should get reallocated on the host, a new mem_move_to_host
was added.
Thanks to Jorn Visser for investigating and finding this problem.
Pull Request: https://projects.blender.org/blender/blender/pulls/132912
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.
Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
This is because with the addition of new features to Cycles, these GPUs
experienced significant performance regressions and bugs, all stemming
from bugs in the Metal GPU driver/compiler. The only reasonable way to
work around these issues was to disable parts of Cycles code on
these GPUs to avoid the driver/compiler bugs.
This resulted in increased development time maintaining these platforms
while being unable to deliver feature parity with other
GPU backends.
It has been decided that this development time is better spent
maintaining platforms that are still actively maintained by
hardware/software vendors, and so AMD and Intel GPU support will be
removed from the Metal backend for Cycles.
Pull Request: https://projects.blender.org/blender/blender/pulls/123551
Since #118841 there are more cases where Cycles would check for the
graphics interop support. This could lead to a crash when graphics
interop functions are called without having active graphics context.
This change makes it so there is no graphics interop calls when doing
headless render. In order to achieve this the device creation is now
aware of the headless mode.
Pull Request: https://projects.blender.org/blender/blender/pulls/122844
Better to disable than crashing, as we are not expecting a quick fix. The cause
is likely similar to issues with the light tree, which was already disabled.
Ref #104013
HIP RT enables AMD hardware ray tracing on RDNA2 and above, and falls back to a
to shader implementation for older graphics cards. It offers an average 25%
sample rendering rate improvement in Cycles benchmarks, on a W6800 card.
The ray tracing feature functions are accessed through HIP RT SDK, available on
GPUOpen. HIP RT traversal functionality is pre-compiled in bitcode format and
shipped with the SDK.
This is not yet enabled as there are issues to be resolved, but landing the
code now makes testing and further changes easier.
Known limitations:
* Not working yet with current public AMD drivers.
* Visual artifact in motion blur.
* One of the buffers allocated for traversal has a static size. Allocating it
dynamically would reduce memory usage.
* This is for Windows only currently, no Linux support.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Ref #105538
Updated Embree 4 library with GPU support is required for it to be
compiled - compatiblity with Embree 3 and Embree 4 without GPU support
is maintained.
Enabling hardware raytracing is an opt-in user setting for now.
Pull Request: https://projects.blender.org/blender/blender/pulls/106266
For example
```
OIIOOutputDriver::~OIIOOutputDriver()
{
}
```
becomes
```
OIIOOutputDriver::~OIIOOutputDriver() {}
```
Saves quite some vertical space, which is especially handy for
constructors.
Pull Request: https://projects.blender.org/blender/blender/pulls/105594
Host memory fallback in CUDA and HIP devices is almost identical.
We remove duplicated code and create a shared generic version that
other devices (oneAPI) will be able to use.
Reviewed By: brecht
Differential Revision: https://developer.blender.org/D17173
Uses a light tree to more effectively sample scenes with many lights. This can
significantly reduce noise, at the cost of a somewhat longer render time per
sample.
Light tree sampling is enabled by default. It can be disabled in the Sampling >
Lights panel. Scenes using light clamping or ray visibility tricks may render
different as these are biased techniques that depend on the sampling strategy.
The implementation is currently disabled on AMD HIP. This is planned to be fixed
before the release.
Implementation by Jeffrey Liu, Weizhen Huang, Alaska and Brecht Van Lommel.
Ref T77889