Commit graph

156 commits

Author SHA1 Message Date
Xavier Hallade
73cfbda0a0 Cycles: oneAPI: Increase free memory reserve for states allocation to 2%
This is a bit empirical at the moment. On Linux+Intel Arc A750,
volume_instance test is running out of memory with 0% and all tests are
passing with 1%, 2% sounds reasonable for now.

Pull Request: https://projects.blender.org/blender/blender/pulls/164329
2026-09-25 23:02:55 +02:00
Brecht Van Lommel
5c328388f3 Cycles: Add concurrent states growth and shrinking for all GPU backends
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.

Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.

The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.

For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.

Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>

Pull Request: https://projects.blender.org/blender/blender/pulls/163930
2026-09-23 15:22:57 +02:00
Brecht Van Lommel
3edfd6e396 Cleanup: Cycles: Fix some mistakes in error messages and comments
These issues were introduced in 527f9ea306 and 41a27b4eed.

Pull Request: https://projects.blender.org/blender/blender/pulls/163579
2026-09-07 14:34:08 +02:00
Christoph Neuhauser
6d864c8f0a Fix #159584: Cycles, Compositor: Do not bundle UR Level Zero v1 adapter
A bug that existed in Level Zero loader versions up to 1.28.2 (fixed by
https://github.com/oneapi-src/level-zero/pull/435). It causes a crash
when calling `zeInitDrivers` a second time when no Level Zero drivers
are available. In Cycles and the Compositor module, this behavior is
triggered when `sycl::platform::get_platforms()` is called (in the case
of the Compositor module transitively via OIDN). SYCL loads both the
Level Zero v1 and v2 adapters, which both call the problematic function.

As a workaround, the SYCL UR Level Zero v1 adapter library files are
no longer bundled and the `SYCL_UR_USE_LEVEL_ZERO_V2=1` environment
variable is set to force devices to use the v2 adapter that would
otherwise default to v1. The v2 adapter has been shipped by Blender
starting with 5.0 (using DPC++ version 6.2), but has been available
in DPC++ since version 6.0.

Pull Request: https://projects.blender.org/blender/blender/pulls/163087
2026-08-26 11:22:17 +02:00
Sergey Sharybin
a70b34065d Refactor: Cycles: Make kernel_features 64bit
This change adds 32 more bit to store kernel features.

While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.

Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.

The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.

Ref #159470

Pull Request: https://projects.blender.org/blender/blender/pulls/162737
2026-08-18 10:41:27 +02:00
Nikita Sirgienko
4f07d79bcc Cycles: Show devices with outdated drivers as disabled in preferences
When a GPU device's driver does not meet Blender's minimum required
version, the device is now shown in the preferences as a greyed-out
entry with a message inside the brackets, indicating which driver
version is needed, instead of being silently hidden.

This helps users understand why their GPU is not available for
rendering and what action they can take to resolve the situation.

Pull Request: https://projects.blender.org/blender/blender/pulls/159405
2026-06-25 10:51:09 +02:00
Hugh Delaney
b374914881 Fix: Cycles: SYCL_CACHE_THRESHOLD ignored
This environment variable was always overridden due to a typo.

Pull Request: https://projects.blender.org/blender/blender/pulls/160396
2026-06-22 12:15:49 +02:00
Nikita Sirgienko
60dbeb4abb Cycles: oneAPI: Increase minimal Intel Linux driver after IGC upgrade
This is a mandatory step, which is needed to be done after our
recent IGC upgrade for Blender 5.2 LTS - to ensure compatibility
between generated IGC binaries and their execution on the
end user system.

Pull Request: https://projects.blender.org/blender/blender/pulls/159416
2026-06-02 19:10:30 +02:00
Brecht Van Lommel
5fc85d1cb7 Refactor: Cycles: Move MNEE walk into separate kernel
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.

MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.

MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.

This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.

Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
2026-05-27 20:34:06 +02:00
Brecht Van Lommel
56a90cac85 Refactor: Cycles: Support allocating GPU device_image during render
Delay device side reallocation of image_info until load_image_info(),
and add DeviceQueue function to do this on the queue. This will allow
multiple GPU devices asynchronously update their image_info so that
one GPU does not have to stop kernel execution while another GPU
allocates a new image tile.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Xavier Hallade
810648dc41 Fix: Cycles: oneAPI: add missing ; after d15d20d51e
Fixing a typo introduced while rebasing when landing the PR.
2026-03-25 14:03:32 +01:00
Nikita Sirgienko
ecdb7fdcaf Cycles: oneAPI: Remove separate context workaround
The minimum supported driver version for Blender 5.1 and
5.2 already includes a fix for the underlying issue, so
there is no need to keep the workaround that isolated each
device into a separate DPC++ context when multiple
descrete GPUs or Level-Zero platforms were detected.
It is now removed.

Ref #138384

Pull Request: https://projects.blender.org/blender/blender/pulls/156249
2026-03-25 13:19:44 +01:00
Nikita Sirgienko
d15d20d51e Fix #155964: Crash when rendering on Intel dGPU on a system with Intel iGPU
On systems with both an Intel iGPU and dGPU using different
drivers (e.g., legacy 11th-14th Gen driver alongside the
Arc driver), the oneAPI Level-Zero copy optimization extension
can cause crashes during host-to-device memory transfers.

Detect when multiple Level-Zero platforms are present, which
indicates separate Intel drivers in the system, and disable
the copy optimization extension in such configurations to
prevent crashes.

This workaround can be removed once the minimum supported
driver version includes a future fix for the underlying
issue.
2026-03-25 13:19:43 +01:00
Brecht Van Lommel
84d4a2e9cf Refactor: Cycles: Split KernelImageTexture and KernelImageInfo
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.

Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.

Pull Request: https://projects.blender.org/blender/blender/pulls/154668
2026-02-25 17:07:39 +01:00
Brecht Van Lommel
e25f0c4c7e Refactor: Cycles: Compute device_memory name on demand for logging
Previously a string needed to be stored for each image texture, but we
might as well compute this on demand and avoid the overhead. There is
now a separate log_name() for the logging, and global_name() for
copying to GPU kernel global variable.

Pull Request: https://projects.blender.org/blender/blender/pulls/154668
2026-02-25 17:07:24 +01:00
Nikita Sirgienko
b105a3bc9a Cycles: oneAPI: use ocloc 101.8424 on Windows
This new version of the graphics compiler brings a fix to recently
discovered issue with Ahead-Of-Time binaries, used right now by
Blender for upcoming Intel® Core™ Ultra Series 3 which would
lead to the rejection of the GPU binaries by the future drivers
on this platform. In order to avoid such situation, and spare
users time to recompile the GPU binaries, this upgrade is
necessary.

Previously set minimal driver version 101.8306 was not increased,
and compatibility was manually tested internally at Intel, to
ensure no problems with it.

Pull Request: https://projects.blender.org/blender/blender/pulls/154647
2026-02-24 09:58:44 +01:00
Brecht Van Lommel
ab0cb89d89 Cleanup: Cycles: Fix various compiler warnings
* Uninitialized variable warning in oneAPI.
* Use std::copy_n instead of memcpy.
* Use simpler bit packing for normal map convention that avoids
  signed/unsigned warning.
* Unnecessary device keyword for default constructor.
* Unused variables in Principled BSDF due to constexpr.
* Hydra function that should be static.

Pull Request: https://projects.blender.org/blender/blender/pulls/154430
2026-02-16 12:57:53 +01:00
Stefan Werner
63607b7bfc Cycles: Removed OneAPI host device support
Host execution of OneAPI devices was used for development/debugging. It hasn't been working lately and only adds complexity.

Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/153650
2026-01-30 14:00:11 +01:00
Nikita Sirgienko
d4147c046e Cycles: oneAPI: Generate ARL-H AoT binaries and mark it as optimized
These changes are enabling generation of AoT binaries for the
ARL-H architecture, which are iGPU architecture in Intel CPUs
such as Intel Core Ultra 9 285H, Intel Core Ultra 7 265H/255H
and Intel Core Ultra 5 225H. We are also marking this
architecture as optimized one. In addition, we are also
refactoring the oneAPI table for the optimization status - making
it more easy to compare and maintain against the list of the
recognized, by DPC++ runtime, Intel architectures.

Pull Request: https://projects.blender.org/blender/blender/pulls/153528
2026-01-29 15:29:52 +01:00
Xavier Hallade
8abc0f2e57 Cycles: oneAPI: Use copy engine
The use of Intel GPU's Copy Engine was initially disabled for improved
stability but it seems more mature now.
Enabling it fixes a performance regression seen on Linux+A750,
indirectly introduced by #152649 that led to an increased use of
zeCommandListAppendMemoryFill.
Performance remains unchanged on Windows+newer GPUs.

Pull Request: https://projects.blender.org/blender/blender/pulls/153121
2026-01-21 15:06:03 +01:00
Brecht Van Lommel
4b34743b4e Cycles: Perform direct light shader eval in own kernel
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.

Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.

A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.

The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.

LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.

The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.

Pull Request: https://projects.blender.org/blender/blender/pulls/152649
2026-01-20 20:34:16 +01:00
Brecht Van Lommel
527f9ea306 Refactor: Cycles: More consistent naming of image functions and structs
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.

Pull Request: https://projects.blender.org/blender/blender/pulls/152665
2026-01-14 17:57:46 +01:00
Nikita Sirgienko
421a061f16 Cycles: oneAPI: use ocloc 101.8331 on Windows
This new version of the graphics compiler brings on average
no performance change for the currently supported Intel devices
and adds small performance improvements for the upcoming Intel
hardware. Such an upgrade also requires an increase in the
minimal supported driver version on Windows, which is why these
changes are combined together with the ocloc upgrade.

Previously set minimal version 101.8132 was increased to 101.8331.

Pull Request: https://projects.blender.org/blender/blender/pulls/152535
2026-01-09 15:04:44 +01:00
Campbell Barton
d1c9ef117d Cleanup: spelling in comments, strings & vars (make check_spelling_*) 2025-11-29 22:12:24 +11:00
Nikita Sirgienko
23b8053416 Cycles: oneAPI: Lower the minimal driver version for Intel® Arc™ Pro
We were expecting the compute-runtime version to be 34938
on Windows, which is not too limiting for Intel Client GPUs.
But for the latest workstation (drivers for Intel® Arc™ Pro GPU)
driver, it is right now are 34177. Yet, it is close enough to be
compatible with our AoT GPU binaries, which we created using
ocloc with version 34938. So, in order to allow Arc Pro users to
use AoT GPU binaries, I am lowering the minimal version, that
will be accepted, to 34177.
2025-10-18 21:50:10 +02:00
Nikita Sirgienko
38adb8f1a4 Cycles: oneAPI: Fix duplicated GPU device entries on some setups
In some hardware configurations, it is possible that DPC++ or
Intel Drivers wrongfully report all devices twice. It is already
being worked on internally, and the fixes will be available in
the future - but for now, we need a workaround for this problem
in Blender as well, to ensure that our end-users are not impacted.

Pull Request: https://projects.blender.org/blender/blender/pulls/147731
2025-10-10 17:25:29 +02:00
Nikita Sirgienko
b133019f9f Cycles: oneAPI: use ocloc 101.8132 on Windows
This new version of the graphics compiler improves performance
for the majority of supported Intel devices and adds support
for upcoming Intel hardware. Such an upgrade also requires
an increase in the minimal supported driver version on Windows,
which is why these changes are combined together with
the ocloc upgrade.

Previously set minimal version 101.6557 was increased to 101.8132.

Pull Request: https://projects.blender.org/blender/blender/pulls/147460
2025-10-08 13:36:08 +02:00
Christoph Neuhauser
72f098248d Cycles: Add Vulkan/oneAPI graphics interop
This PR adds Vulkan/oneAPI graphics interop to Cycles. Just like for
CUDA and HIP interop, persistent memory mapping is used, as there could
potentially be some overhead of continuously mapping/unmapping buffers.

Pull Request: https://projects.blender.org/blender/blender/pulls/144442
2025-10-06 18:16:56 +02:00
Nikita Sirgienko
49414a72f6 Cycles: oneAPI: Add new arch codes for upcoming Intel hardware
Pull Request: https://projects.blender.org/blender/blender/pulls/147221
2025-10-04 22:34:54 +02:00
Thomas Dinges
66224d69b0 Deps: Library changes for Blender 5.0
This commit includes the changes to the build system, updated hashes to the actual new libraries as well as a required test update.

* DPC++ 6.2.0 RC
* freetype 2.13.3
* HIP 6.4.5010
* IGC 2.16.0
* ISPC 1.28.0
* libharu  2.4.5
* libpng 1.6.50
* libvpx 1.15.2
* libxml2 2.14.5
* LLVM 20.1.8
* Manifold 3.2.1
* MaterialX 1.39.3
* OpenColorIO 2.4.2
* openexr 3.3.5
* OpenImageIO 3.0.9.1
* openjpeg 2.5.3
* OpenShadingLanguage 1.14.7.0
* openssl 3.5.2
* Python 3.11.13
* Rubber Band 4.0.0
* ShaderC 2025.3
* sqlite 3.50.4
* USD 25.08
* Wayland 1.24.0

Ref #138940

Co-authored-by: Ray Molenkamp <github@lazydodo.com>
Co-authored-by: Jesse Yurkovich <jesse.y@gmail.com>
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Co-authored-by: Nikita Sirgienko <nikita.sirgienko@intel.com>
Co-authored-by: Sybren A. Stüvel <sybren@blender.org>
Co-authored-by: Kace <lakacey03@gmail.com>
Co-authored-by: Sebastian Parborg <sebastian@blender.org>
Co-authored-by: Anthony Roberts <anthony.roberts@linaro.org>
Co-authored-by: Jonas Holzman <jonas@holzman.fr>

Pull Request: https://projects.blender.org/blender/blender/pulls/144479
2025-10-02 18:34:11 +02:00
Weizhen Huang
2b0a1cae06 Cycles: Add an option to use ray marching for volume rendering
Null Scattering currently has performance and noise issues, and it will
take time to address them. For now add the previous Ray Marching back as
an option.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/146317
2025-09-26 12:14:45 +02:00
Nikita Sirgienko
5efeb06613 Fix #145449: Workaround wrongly generated Intel Linux driver version
There are several Driver versions which are constructing the wrong,
semantically, version which would force Blender to decline the Intel
device for oneAPI backend usage, based on this. Unfortunately,
the upstream fix is taking a long time to be finally delivered to
the distros and end-users, so it is better if Blender will detect
this wrong version string and parse it properly, allowing these
devices to be used - as the wrong driver version string is the only
issue here, besides this the driver functionality is fine.

Pull Request: https://projects.blender.org/blender/blender/pulls/145658
2025-09-03 19:26:05 +02:00
Brecht Van Lommel
2615cecf10 Refactor: Cycles: Align log levels with CLOG
WORK -> DEBUG
DEBUG, STATS -> TRACE

Pull Request: https://projects.blender.org/blender/blender/pulls/144490
2025-08-18 20:22:44 +02:00
Nikita Sirgienko
21cba7024c Cycles: oneAPI: Disable L0 copy optimization for several dGPUs
Currently, it was discovered that in the case of several different
Intel dGPUs being present in the system, the experimental L0 copy
optimization does not work correctly in the Intel Driver, which is
causing crashes in the driver and Blender application. So, to avoid
this situation and restore functionality on these platforms,
a workaround was added to disable this extension from being used if
such a configuration is detected. In the future, when this problem is
fully fixed in all Intel Drivers, this workaround can be removed from
the Blender source code to restore some performance that was lost on
configurations of several dGPUs because of this workaround.

Pull Request: https://projects.blender.org/blender/blender/pulls/144262
2025-08-14 12:14:51 +02:00
Weizhen Huang
b2b2d9a4f3 Cycles: Render volume by ray marching through octrees
One octree per volume per shader based on the density. In preparation
for the null scattering
2025-08-13 10:28:50 +02:00
Campbell Barton
cccc2c77c5 Cleanup: consistent for C-style comment blocks 2025-08-08 07:37:33 +10:00
Stefan Werner
c81e1d95c1 Cycles: Fixed typo in my last commit 2025-07-29 10:53:13 +02:00
Stefan Werner
e7312b1ad5 Cycles: Explicitly setting SYCL device for Embree
This fixes issues when using Embree on mutliple GPUs.
A previous workaround used separate contexts, this one now
lets us keep a single context for all GPUs.

Pull Request: https://projects.blender.org/blender/blender/pulls/143089
2025-07-29 10:40:28 +02:00
Hans Goudey
c3181490f3 Cleanup: Formatting 2025-07-14 10:22:46 -04:00
Nikita Sirgienko
609f8ddbef Cycles: oneAPI: Fix DPC++ level issues for multi GPU execution
These changes introduce modifications to the SYCL queue creation
in OneapiDevice::create_queue. In case several DPC++ devices are
detected by Blender and exposed through it, we are now creating
a new SYCL context for each device, which allows us to prevent
execution failures due to some known issues in the DPC++ runtime
regarding multi GPU support. As this would have some small
performance impact, few percents, it is only applied to
multi GPU configurations, while the behavior for a single
GPU configuration remains the same.

Pull Request: https://projects.blender.org/blender/blender/pulls/141834
2025-07-14 14:33:42 +02:00
Brecht Van Lommel
73fe848e07 Fix: Cycles log levels conflict with macros on some platforms
In particular DEBUG, but prefix all of them to be sure.

Pull Request: https://projects.blender.org/blender/blender/pulls/141749
2025-07-10 19:44:14 +02:00
Xavier Hallade
94e9203713 Fix previous 4.5 merge 2025-07-10 17:47:03 +02:00
Xavier Hallade
48f89ff1c3 Merge branch 'blender-v4.5-release' 2025-07-10 17:43:30 +02:00
Xavier Hallade
05f27f594e Fix #141661: Crash when selecting oneAPI in preferences with legacy drivers
On systems with multiple Intel GPUs with a mix of recent and old
unsupported drivers (such as 101.3302), the Level-Zero stack may have
troubles initializing, leading to a crash while enumerating devices.

Luckily this condition actually leads to an exception we can catch,
as implemented here in this commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/141674
2025-07-10 17:36:00 +02:00
Brecht Van Lommel
b6c4233b28 Refactor: Cycles: Remove now unused 3D image texture support
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Brecht Van Lommel
7978799e6f Cycles: Always render volume as NanoVDB
All GPU backends now support NanoVDB, using our own kernel side code
that is easily portable. This simplifies kernel and device code.

Volume bounds are now built from the NanoVDB grid instead of OpenVDB,
to avoid having to keep around the OpenVDB grid after loading.

While this reduces memory usage, it does have a performance impact,
particularly for the Cubic filter. That will be addressed by
another commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Brecht Van Lommel
fb4e3c8167 Refactor: Cycles: Remove distinction between severity and verbosity
Only use LOG() and LOG_IS_ON() macros, no more VLOG_.

Pull Request: https://projects.blender.org/blender/blender/pulls/140244
2025-07-09 20:59:24 +02:00
Xavier Hallade
7691e6520b Fix #141171: oneAPI: Rendering artifacts in barbershop scene
max_shaders was not updated when Embree was disabled.

Pull Request: https://projects.blender.org/blender/blender/pulls/141175
2025-06-30 16:39:53 +02:00
Xavier Hallade
2df163a648 Fix: Cycles low performance with scenes with many shaders on Arc B570
The performance of the sorted_paths_array kernel on B570 is problematic.
Relying on local sorting+partitioning instead gives a 25% overall rendering
speedup and no regression in shade_surface when rendering Agent 327 Barbershop scene.
On Arc A770, it still gives a 2% speedup when rendering Barbershop.

Pull Request: https://projects.blender.org/blender/blender/pulls/140308
2025-06-18 08:21:19 +02:00
Brecht Van Lommel
e84fad92ea Fix #139986: Cycles crash on some scene updates, after Embree upgrade
Device::const_copy_to is sometimes called when the Embree BVH has been freed
and not replaced yet. Previously this was a simpler pointer copy, now there is
a function call. Make sure it's just a function copy.

Thanks to Nikita Sirgienko for figuring this out.

Pull Request: https://projects.blender.org/blender/blender/pulls/140457
2025-06-16 17:59:57 +02:00