Commit graph

189 commits

Author SHA1 Message Date
Brecht Van Lommel
067ed4498f Fix: Cycles: Ensure integrator working memory is not host mapped
The integrator state was already device only, but some other memory is
also frequently accessed and should be on the GPU for best performance.

This is a follow up for a pre-existing issue found reviewing #163437 and
#163930.

Pull Request: https://projects.blender.org/blender/blender/pulls/164306
2026-09-25 10:49:13 +02:00
Brecht Van Lommel
5c328388f3 Cycles: Add concurrent states growth and shrinking for all GPU backends
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.

Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.

The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.

For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.

Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>

Pull Request: https://projects.blender.org/blender/blender/pulls/163930
2026-09-23 15:22:57 +02:00
Brecht Van Lommel
677de09a15 Fix: Cycles: MNEE enabled on HIP without HIP-RT in multi-device render
The change in 61ca1b5856 only handled the single device case properly.

Pull Request: https://projects.blender.org/blender/blender/pulls/164137
2026-09-19 23:14:16 +02:00
Sergey Sharybin
6b15792508 Cycles: Switch to CUDA 13 and enable DLSS
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.

This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.

The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from

  build_files/config/pipeline_config.yaml

to

  build_files/buildbot/config/blender_version.cmake

Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.

Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:

- There is an easy to follow table of per-architecture or family
  toolkits.
- If there is no architectural preference, all provided toolkits are
  probed. For SM kernels newer toolkits are tested first, and for
  COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
  default one is used.

A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).

Pull Request: https://projects.blender.org/blender/blender/pulls/164002
2026-09-18 19:30:10 +02:00
Nikita Sirgienko
4f07d79bcc Cycles: Show devices with outdated drivers as disabled in preferences
When a GPU device's driver does not meet Blender's minimum required
version, the device is now shown in the preferences as a greyed-out
entry with a message inside the brackets, indicating which driver
version is needed, instead of being silently hidden.

This helps users understand why their GPU is not available for
rendering and what action they can take to resolve the situation.

Pull Request: https://projects.blender.org/blender/blender/pulls/159405
2026-06-25 10:51:09 +02:00
Brecht Van Lommel
61ca1b5856 Fix #160089: Cycles: Disable MNEE on HIP without hardware ray-tracing
Unfortunately it appears that moving MNEE to another kernel did not
fundamentally fix the apparent compiler bug that breaks this. Another
refactor in 0baa98866c made the bug surface again.

It appears to work fine with HIP-RT, so we leave that case enabled.
HIP-RT is also enabled by default, so it's not as bad.

Pull Request: https://projects.blender.org/blender/blender/pulls/160110
2026-06-23 12:33:04 +02:00
Brecht Van Lommel
9c736864af Refactor: Cycles: Generalize tile request mask to tile access state
To extend it with a USED state for cache eviction later.

Pull Request: https://projects.blender.org/blender/blender/pulls/157244
2026-04-20 17:51:59 +02:00
Brecht Van Lommel
56a90cac85 Refactor: Cycles: Support allocating GPU device_image during render
Delay device side reallocation of image_info until load_image_info(),
and add DeviceQueue function to do this on the queue. This will allow
multiple GPU devices asynchronously update their image_info so that
one GPU does not have to stop kernel execution while another GPU
allocates a new image tile.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Brecht Van Lommel
105c7ac6d4 Refactor: Cycles: Support allocating CPU device_image during render
Allow updating image_info while rendering for CPU devices, by replacing
the pointer in all per-thread kernel globals. The per-thread kernel globals
are now owned by the device to make this possible.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Brecht Van Lommel
e25f0c4c7e Refactor: Cycles: Compute device_memory name on demand for logging
Previously a string needed to be stored for each image texture, but we
might as well compute this on demand and avoid the overhead. There is
now a separate log_name() for the logging, and global_name() for
copying to GPU kernel global variable.

Pull Request: https://projects.blender.org/blender/blender/pulls/154668
2026-02-25 17:07:24 +01:00
Sergey Sharybin
05a3b98542 Cycles: Remove the HIP-RT extension wrangler
There is a strong coupling between the API version an application
passes to the hiprtCreateContext() and the HIP-RT library: if there
is a mismatch the context will fail to be created.

Additionally, the HIP-RT library does not have static dependencies
as it loads HIP SDK dynamically.

It all makes the wrangler be not so useful, and harmful in cases
when Blender is packaged for different Linux distros which might
use HIP-RT library of a different version from what Cycles is
currently expecting via the bundled hiprtew.h.

The HIPRT_API_VERSION is now coming from the hiprt/hiprt.h so it
is guaranteed that Cycles will request the same API version as
the library is was compiled with.

The slightly annoying part is that this approach requires the
dynamic libraries to have API version in the name of the library
to function properly This is because of the SONAME which still
mentions the original library name. The similar thing happens on
Windows where the hiprt0200564.lib points to the hiprt0200564.dll.
There might be some way to do the rename during install, but it
is a bit tricky, especially due to the manifest on Windows. It
might be the easiest to simply rename precompiled libraries.

Pull Request: https://projects.blender.org/blender/blender/pulls/153892
2026-02-13 17:13:57 +01:00
Brecht Van Lommel
05e2f09fa7 Fix #146642: Memory leak warnings when killing render with Ctrl + C
Use lazy initialization for global variables, so they get destructed before the
guardedalloc leak check and destruction. The same pattern is used elsewhere
in Blender.

There may be more cases, these were the ones I found in testing. These issue
were exposed by f8eec542f4.

Pull Request: https://projects.blender.org/blender/blender/pulls/154173
2026-02-10 12:00:54 +01:00
Brecht Van Lommel
527f9ea306 Refactor: Cycles: More consistent naming of image functions and structs
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.

Pull Request: https://projects.blender.org/blender/blender/pulls/152665
2026-01-14 17:57:46 +01:00
Nikita Sirgienko
d36eea7fa7 Fix #129596: Always use multithreaded BVH building in Cycles again
The workaround of forcing BVH building into single thread
execution on the Blender side is not needed anymore,
because the problem was properly fixed in the upstream
since Embree upgrade in Blender 4.5

This reverts commit c0f0e2ca6f.

Pull Request: https://projects.blender.org/blender/blender/pulls/146859
2025-09-27 21:07:39 +02:00
Brecht Van Lommel
2615cecf10 Refactor: Cycles: Align log levels with CLOG
WORK -> DEBUG
DEBUG, STATS -> TRACE

Pull Request: https://projects.blender.org/blender/blender/pulls/144490
2025-08-18 20:22:44 +02:00
Brecht Van Lommel
73fe848e07 Fix: Cycles log levels conflict with macros on some platforms
In particular DEBUG, but prefix all of them to be sure.

Pull Request: https://projects.blender.org/blender/blender/pulls/141749
2025-07-10 19:44:14 +02:00
Brecht Van Lommel
8cf031ba95 Fix: Wrong Cycles NanoVDB memory alignment on Windows
This was not a problem in practice so far, but will be with upcoming changes.

Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 20:59:27 +02:00
Brecht Van Lommel
cf36acbc0c Refactor: Cycles: Replace remaining fprintf with logging
Pull Request: https://projects.blender.org/blender/blender/pulls/140244
2025-07-09 20:59:25 +02:00
Brecht Van Lommel
fb4e3c8167 Refactor: Cycles: Remove distinction between severity and verbosity
Only use LOG() and LOG_IS_ON() macros, no more VLOG_.

Pull Request: https://projects.blender.org/blender/blender/pulls/140244
2025-07-09 20:59:24 +02:00
Brecht Van Lommel
cf7f276d49 Refactor: Cycles: Tweak logging to prepare for dropping glog
* Implement own simple ScopedMockLog
* Always use names instead of numbers
* Avoid logging in header files

Pull Request: https://projects.blender.org/blender/blender/pulls/140244
2025-07-09 20:59:24 +02:00
Brecht Van Lommel
c0f0e2ca6f Fix #129596: Cycles oneAPI crash with interactive BVH updates
There is a bug in Embree that makes BVH updates crash. Disabling multithreaded
BVH updates after the initial BVH build appears to work around it, at the cost
of some performance.

This will not affect performance of the initial BVH build, transforming objects
or editing a single mesh. It will only affect performance when multiple smaller
meshes are edited together, as those can no longer have their BVH updated in
parallel or benefit from parallellization over many primitives.

Pull Request: https://projects.blender.org/blender/blender/pulls/134747
2025-02-18 16:20:59 +01:00
Brecht Van Lommel
f99f958c47 Refactor: Cycles: Add host_alloc/free to device API
This may be used for device to do host memory allocation in a way that
is more efficient for copy the host memory to the device.

Also rename and group device memory allocation functions for clarity.

Pull Request: https://projects.blender.org/blender/blender/pulls/134412
2025-02-13 19:58:56 +01:00
Brecht Van Lommel
c4c0c23c5a Fix: Cycles: Always try to alloc MEM_DEVICE_ONLY on device
Regardless of what mem info reports. We can't move this to the host, so
might as well try because the free memory might not be a reliable predictor
of success.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:12:25 +01:00
Brecht Van Lommel
e8ebcb3ee3 Fix: Cycles: Check if memory is host mapped without access to device_mem_map
This avoids concurrency issues.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:12:23 +01:00
Brecht Van Lommel
8b7fce492e Refactor: Cycles: Change API so host and device memory are freed together
With host mapped memory these can be shared, and we can't get back the
original host pointer unless we make a copy which is inefficient.

Also add asserts to verify this doesn't happen.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:12:19 +01:00
Brecht Van Lommel
1ec04e0eec Fix: Cycles: Only move textures to host on one device at a time
This was not thread safe. And it's better to do them one by one to avoid
moving more than is needed, when another thread already freed up enough.

Thanks to Jorn Visser for investigating and finding this problem.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:12:09 +01:00
Brecht Van Lommel
cd3d3b2646 Refactor: Cycles: Delay load_texture_info() to enqueue
Doing it immediately after moving textures to the host is less efficient, and
interacts in confusing ways.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:12:06 +01:00
Brecht Van Lommel
2cfe2e0bfe Fix: Cycles: Re-copy memory from host to device without realloc
Should be a bit more efficient, and it fixes host memory fallback bugs,
where host memory was incorrectly freed during re-copy. For the case
where memory should get reallocated on the host, a new mem_move_to_host
was added.

Thanks to Jorn Visser for investigating and finding this problem.

Pull Request: https://projects.blender.org/blender/blender/pulls/132912
2025-01-29 14:11:50 +01:00
Brecht Van Lommel
9971648783 Refactor: Cycles: Replace new/delete by unique_ptr, in simple cases
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:30 +01:00
Brecht Van Lommel
a8654a1dbe Refactor: Cycles: Make CPU kernel globals storage more sane
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:27 +01:00
Brecht Van Lommel
57ff24cb99 Refactor: Cycles: Add const keyword to more function parameters
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:24 +01:00
Brecht Van Lommel
dd51c8660b Refactor: Cycles: Add const keyword where possible, using clang-tidy
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.

Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:20 +01:00
Brecht Van Lommel
60bec183cb Refactor: Cycles: Replace foreach() by range based for loops
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:05 +01:00
Brecht Van Lommel
d0c2e68e5f Refactor: Cycles: Automated clang-tidy fixups in Cycles
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:55 +01:00
Brecht Van Lommel
3c2a6fbb9c Refactor: Cycles: Use nullptr instead of NULL
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:43 +01:00
Alaska
c8340cf754 Cycles: Remove AMD and Intel GPU support from Metal backend
This is because with the addition of new features to Cycles, these GPUs
experienced significant performance regressions and bugs, all stemming
from bugs in the Metal GPU driver/compiler. The only reasonable way to
work around these issues was to disable parts of Cycles code on
these GPUs to avoid the driver/compiler bugs.

This resulted in increased development time maintaining these platforms
while being unable to deliver feature parity with other
GPU backends.

It has been decided that this development time is better spent
maintaining platforms that are still actively maintained by
hardware/software vendors, and so AMD and Intel GPU support will be
removed from the Metal backend for Cycles.

Pull Request: https://projects.blender.org/blender/blender/pulls/123551
2024-06-26 17:16:20 +02:00
Sergey Sharybin
b803d7fabb Fix: Command line Cycles render crash on multi-CUDA device
Since #118841 there are more cases where Cycles would check for the
graphics interop support. This could lead to a crash when graphics
interop functions are called without having active graphics context.

This change makes it so there is no graphics interop calls when doing
headless render. In order to achieve this the device creation is now
aware of the headless mode.

Pull Request: https://projects.blender.org/blender/blender/pulls/122844
2024-06-07 17:53:44 +02:00
Campbell Barton
5b9740c913 Cleanup: use braces for sources in intern/
Omitted intern/itasc as some of these sources are from KDL:
https://www.orocos.org/kdl.html
2023-09-17 09:05:40 +10:00
Nikita Sirgienko
7977247995 Cycles: Avoid printing capabilities when there are no devices
Pull Request: https://projects.blender.org/blender/blender/pulls/109511
2023-07-03 13:09:41 +02:00
Brecht Van Lommel
45b9542e6c Merge branch 'blender-v3.6-release' into main 2023-06-15 16:45:15 +02:00
Brecht Van Lommel
0ab58864f3 Fix Cycles Metal AMD crash with shadow caustics, by disabling it
Better to disable than crashing, as we are not expecting a quick fix. The cause
is likely similar to issues with the light tree, which was already disabled.

Ref #104013
2023-06-15 16:33:21 +02:00
Campbell Barton
c12994612b License headers: use SPDX-FileCopyrightText in intern/cycles 2023-06-14 16:53:23 +10:00
Campbell Barton
6859bb6e67 Cleanup: format (with BraceWrapping::AfterControlStatement "MultiLine") 2023-05-02 09:37:49 +10:00
Sahar A. Kashi
557a245dd5 Cycles: add HIP RT device, for AMD hardware ray tracing on Windows
HIP RT enables AMD hardware ray tracing on RDNA2 and above, and falls back to a
to shader implementation for older graphics cards. It offers an average 25%
sample rendering rate improvement in Cycles benchmarks, on a W6800 card.

The ray tracing feature functions are accessed through HIP RT SDK, available on
GPUOpen. HIP RT traversal functionality is pre-compiled in bitcode format and
shipped with the SDK.

This is not yet enabled as there are issues to be resolved, but landing the
code now makes testing and further changes easier.

Known limitations:
* Not working yet with current public AMD drivers.
* Visual artifact in motion blur.
* One of the buffers allocated for traversal has a static size. Allocating it
  dynamically would reduce memory usage.
* This is for Windows only currently, no Linux support.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>

Ref #105538
2023-04-25 20:19:43 +02:00
Nikita Sirgienko
3f8c995109 Cycles: add hardware raytracing support to oneAPI device
Updated Embree 4 library with GPU support is required for it to be
compiled - compatiblity with Embree 3 and Embree 4 without GPU support
is maintained.
Enabling hardware raytracing is an opt-in user setting for now.

Pull Request: https://projects.blender.org/blender/blender/pulls/106266
2023-04-18 22:09:42 +02:00
Sergey Sharybin
d32d787f5f Clang-Format: Allow empty functions to be single-line
For example

```
OIIOOutputDriver::~OIIOOutputDriver()
{
}
```

becomes

```
OIIOOutputDriver::~OIIOOutputDriver() {}
```

Saves quite some vertical space, which is especially handy for
constructors.

Pull Request: https://projects.blender.org/blender/blender/pulls/105594
2023-03-29 16:50:54 +02:00
Brecht Van Lommel
cc6d8cd573 Fix #105442: Cycles CUDA and HIP host memory fallback not working
Transforming the host pointer should not be done in an assert, it only works
in debug builds then. Caused by 6dcfb6d.
2023-03-17 21:52:29 +01:00
Lucas Tadeu
a1282ab015 Fix Cycles debug build error after host falback changes
Introduced in dcfb6df9ce6.

Co-authored-by: Lucas Tadeu Teixeira <lucas@lucastadeu.com>

Pull Request #104454
2023-02-08 19:27:40 +01:00
Nikita Sirgienko
6dcfb6df9c Cycles: Abstract host memory fallback for GPU devices
Host memory fallback in CUDA and HIP devices is almost identical.
We remove duplicated code and create a shared generic version that
other devices (oneAPI) will be able to use.

Reviewed By: brecht

Differential Revision: https://developer.blender.org/D17173
2023-02-06 22:19:32 +01:00
Weizhen Huang
ee89f213de Cycles: improve many lights sampling using light tree
Uses a light tree to more effectively sample scenes with many lights. This can
significantly reduce noise, at the cost of a somewhat longer render time per
sample.

Light tree sampling is enabled by default. It can be disabled in the Sampling >
Lights panel. Scenes using light clamping or ray visibility tricks may render
different as these are biased techniques that depend on the sampling strategy.

The implementation is currently disabled on AMD HIP. This is planned to be fixed
before the release.

Implementation by Jeffrey Liu, Weizhen Huang, Alaska and Brecht Van Lommel.

Ref T77889
2022-12-05 16:09:03 +01:00