This is a bit empirical at the moment. On Linux+Intel Arc A750,
volume_instance test is running out of memory with 0% and all tests are
passing with 1%, 2% sounds reasonable for now.
Pull Request: https://projects.blender.org/blender/blender/pulls/164329
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.
Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.
The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.
For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.
Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
A bug that existed in Level Zero loader versions up to 1.28.2 (fixed by
https://github.com/oneapi-src/level-zero/pull/435). It causes a crash
when calling `zeInitDrivers` a second time when no Level Zero drivers
are available. In Cycles and the Compositor module, this behavior is
triggered when `sycl::platform::get_platforms()` is called (in the case
of the Compositor module transitively via OIDN). SYCL loads both the
Level Zero v1 and v2 adapters, which both call the problematic function.
As a workaround, the SYCL UR Level Zero v1 adapter library files are
no longer bundled and the `SYCL_UR_USE_LEVEL_ZERO_V2=1` environment
variable is set to force devices to use the v2 adapter that would
otherwise default to v1. The v2 adapter has been shipped by Blender
starting with 5.0 (using DPC++ version 6.2), but has been available
in DPC++ since version 6.0.
Pull Request: https://projects.blender.org/blender/blender/pulls/163087
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
When a GPU device's driver does not meet Blender's minimum required
version, the device is now shown in the preferences as a greyed-out
entry with a message inside the brackets, indicating which driver
version is needed, instead of being silently hidden.
This helps users understand why their GPU is not available for
rendering and what action they can take to resolve the situation.
Pull Request: https://projects.blender.org/blender/blender/pulls/159405
This is a mandatory step, which is needed to be done after our
recent IGC upgrade for Blender 5.2 LTS - to ensure compatibility
between generated IGC binaries and their execution on the
end user system.
Pull Request: https://projects.blender.org/blender/blender/pulls/159416
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.
MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.
MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.
This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
Delay device side reallocation of image_info until load_image_info(),
and add DeviceQueue function to do this on the queue. This will allow
multiple GPU devices asynchronously update their image_info so that
one GPU does not have to stop kernel execution while another GPU
allocates a new image tile.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
The minimum supported driver version for Blender 5.1 and
5.2 already includes a fix for the underlying issue, so
there is no need to keep the workaround that isolated each
device into a separate DPC++ context when multiple
descrete GPUs or Level-Zero platforms were detected.
It is now removed.
Ref #138384
Pull Request: https://projects.blender.org/blender/blender/pulls/156249
On systems with both an Intel iGPU and dGPU using different
drivers (e.g., legacy 11th-14th Gen driver alongside the
Arc driver), the oneAPI Level-Zero copy optimization extension
can cause crashes during host-to-device memory transfers.
Detect when multiple Level-Zero platforms are present, which
indicates separate Intel drivers in the system, and disable
the copy optimization extension in such configurations to
prevent crashes.
This workaround can be removed once the minimum supported
driver version includes a future fix for the underlying
issue.
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.
Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
Previously a string needed to be stored for each image texture, but we
might as well compute this on demand and avoid the overhead. There is
now a separate log_name() for the logging, and global_name() for
copying to GPU kernel global variable.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
This new version of the graphics compiler brings a fix to recently
discovered issue with Ahead-Of-Time binaries, used right now by
Blender for upcoming Intel® Core™ Ultra Series 3 which would
lead to the rejection of the GPU binaries by the future drivers
on this platform. In order to avoid such situation, and spare
users time to recompile the GPU binaries, this upgrade is
necessary.
Previously set minimal driver version 101.8306 was not increased,
and compatibility was manually tested internally at Intel, to
ensure no problems with it.
Pull Request: https://projects.blender.org/blender/blender/pulls/154647
* Uninitialized variable warning in oneAPI.
* Use std::copy_n instead of memcpy.
* Use simpler bit packing for normal map convention that avoids
signed/unsigned warning.
* Unnecessary device keyword for default constructor.
* Unused variables in Principled BSDF due to constexpr.
* Hydra function that should be static.
Pull Request: https://projects.blender.org/blender/blender/pulls/154430
Host execution of OneAPI devices was used for development/debugging. It hasn't been working lately and only adds complexity.
Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/153650
These changes are enabling generation of AoT binaries for the
ARL-H architecture, which are iGPU architecture in Intel CPUs
such as Intel Core Ultra 9 285H, Intel Core Ultra 7 265H/255H
and Intel Core Ultra 5 225H. We are also marking this
architecture as optimized one. In addition, we are also
refactoring the oneAPI table for the optimization status - making
it more easy to compare and maintain against the list of the
recognized, by DPC++ runtime, Intel architectures.
Pull Request: https://projects.blender.org/blender/blender/pulls/153528
The use of Intel GPU's Copy Engine was initially disabled for improved
stability but it seems more mature now.
Enabling it fixes a performance regression seen on Linux+A750,
indirectly introduced by #152649 that led to an increased use of
zeCommandListAppendMemoryFill.
Performance remains unchanged on Windows+newer GPUs.
Pull Request: https://projects.blender.org/blender/blender/pulls/153121
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.
Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.
A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.
The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.
LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.
The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.
Pull Request: https://projects.blender.org/blender/blender/pulls/152649
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
This new version of the graphics compiler brings on average
no performance change for the currently supported Intel devices
and adds small performance improvements for the upcoming Intel
hardware. Such an upgrade also requires an increase in the
minimal supported driver version on Windows, which is why these
changes are combined together with the ocloc upgrade.
Previously set minimal version 101.8132 was increased to 101.8331.
Pull Request: https://projects.blender.org/blender/blender/pulls/152535
We were expecting the compute-runtime version to be 34938
on Windows, which is not too limiting for Intel Client GPUs.
But for the latest workstation (drivers for Intel® Arc™ Pro GPU)
driver, it is right now are 34177. Yet, it is close enough to be
compatible with our AoT GPU binaries, which we created using
ocloc with version 34938. So, in order to allow Arc Pro users to
use AoT GPU binaries, I am lowering the minimal version, that
will be accepted, to 34177.
In some hardware configurations, it is possible that DPC++ or
Intel Drivers wrongfully report all devices twice. It is already
being worked on internally, and the fixes will be available in
the future - but for now, we need a workaround for this problem
in Blender as well, to ensure that our end-users are not impacted.
Pull Request: https://projects.blender.org/blender/blender/pulls/147731
This new version of the graphics compiler improves performance
for the majority of supported Intel devices and adds support
for upcoming Intel hardware. Such an upgrade also requires
an increase in the minimal supported driver version on Windows,
which is why these changes are combined together with
the ocloc upgrade.
Previously set minimal version 101.6557 was increased to 101.8132.
Pull Request: https://projects.blender.org/blender/blender/pulls/147460
This PR adds Vulkan/oneAPI graphics interop to Cycles. Just like for
CUDA and HIP interop, persistent memory mapping is used, as there could
potentially be some overhead of continuously mapping/unmapping buffers.
Pull Request: https://projects.blender.org/blender/blender/pulls/144442
Null Scattering currently has performance and noise issues, and it will
take time to address them. For now add the previous Ray Marching back as
an option.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/146317
There are several Driver versions which are constructing the wrong,
semantically, version which would force Blender to decline the Intel
device for oneAPI backend usage, based on this. Unfortunately,
the upstream fix is taking a long time to be finally delivered to
the distros and end-users, so it is better if Blender will detect
this wrong version string and parse it properly, allowing these
devices to be used - as the wrong driver version string is the only
issue here, besides this the driver functionality is fine.
Pull Request: https://projects.blender.org/blender/blender/pulls/145658
Currently, it was discovered that in the case of several different
Intel dGPUs being present in the system, the experimental L0 copy
optimization does not work correctly in the Intel Driver, which is
causing crashes in the driver and Blender application. So, to avoid
this situation and restore functionality on these platforms,
a workaround was added to disable this extension from being used if
such a configuration is detected. In the future, when this problem is
fully fixed in all Intel Drivers, this workaround can be removed from
the Blender source code to restore some performance that was lost on
configurations of several dGPUs because of this workaround.
Pull Request: https://projects.blender.org/blender/blender/pulls/144262
This fixes issues when using Embree on mutliple GPUs.
A previous workaround used separate contexts, this one now
lets us keep a single context for all GPUs.
Pull Request: https://projects.blender.org/blender/blender/pulls/143089
These changes introduce modifications to the SYCL queue creation
in OneapiDevice::create_queue. In case several DPC++ devices are
detected by Blender and exposed through it, we are now creating
a new SYCL context for each device, which allows us to prevent
execution failures due to some known issues in the DPC++ runtime
regarding multi GPU support. As this would have some small
performance impact, few percents, it is only applied to
multi GPU configurations, while the behavior for a single
GPU configuration remains the same.
Pull Request: https://projects.blender.org/blender/blender/pulls/141834
On systems with multiple Intel GPUs with a mix of recent and old
unsupported drivers (such as 101.3302), the Level-Zero stack may have
troubles initializing, leading to a crash while enumerating devices.
Luckily this condition actually leads to an exception we can catch,
as implemented here in this commit.
Pull Request: https://projects.blender.org/blender/blender/pulls/141674
All GPU backends now support NanoVDB, using our own kernel side code
that is easily portable. This simplifies kernel and device code.
Volume bounds are now built from the NanoVDB grid instead of OpenVDB,
to avoid having to keep around the OpenVDB grid after loading.
While this reduces memory usage, it does have a performance impact,
particularly for the Cubic filter. That will be addressed by
another commit.
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
The performance of the sorted_paths_array kernel on B570 is problematic.
Relying on local sorting+partitioning instead gives a 25% overall rendering
speedup and no regression in shade_surface when rendering Agent 327 Barbershop scene.
On Arc A770, it still gives a 2% speedup when rendering Barbershop.
Pull Request: https://projects.blender.org/blender/blender/pulls/140308
Device::const_copy_to is sometimes called when the Embree BVH has been freed
and not replaced yet. Previously this was a simpler pointer copy, now there is
a function call. Make sure it's just a function copy.
Thanks to Nikita Sirgienko for figuring this out.
Pull Request: https://projects.blender.org/blender/blender/pulls/140457