This is a bit empirical at the moment. On Linux+Intel Arc A750,
volume_instance test is running out of memory with 0% and all tests are
passing with 1%, 2% sounds reasonable for now.
Pull Request: https://projects.blender.org/blender/blender/pulls/164329
The integrator state was already device only, but some other memory is
also frequently accessed and should be on the GPU for best performance.
This is a follow up for a pre-existing issue found reviewing #163437 and
#163930.
Pull Request: https://projects.blender.org/blender/blender/pulls/164306
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.
Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.
The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.
For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.
Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
Some values don't need full float precision, and with millions of states
reducing memory usage is important, while the math to pack/unpack these
is quite cheap. On CPU full precision is used since there are few states
and the extra conversion cost only hurts.
This gives a 2% reduction in state size with just
KERNEL_FEATURE_PATH_TRACING, and 9% reduction when enabling more
features like LIGHT_PASSES + DENOISING or SUBSURFACE + VOLUME.
Benchmarks do not show any significant impact on GPU render time either
way.
Pull Request: https://projects.blender.org/blender/blender/pulls/161876
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.
This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.
The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from
build_files/config/pipeline_config.yaml
to
build_files/buildbot/config/blender_version.cmake
Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.
Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:
- There is an easy to follow table of per-architecture or family
toolkits.
- If there is no architectural preference, all provided toolkits are
probed. For SM kernels newer toolkits are tested first, and for
COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
default one is used.
A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).
Pull Request: https://projects.blender.org/blender/blender/pulls/164002
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.
This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.
Pull Request: https://projects.blender.org/blender/blender/pulls/162785
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
Commit 976ca9d244 fixed the "Hardware
Ray-Tracing" state message being printed to the log always showing off
for OptiX. This improves that further by querying whether the device
actually supports it, since OptiX has a software fallback on devices
without hardware ray tracing support.
Pull Request: https://projects.blender.org/blender/blender/pulls/163807
Adds the option to use DLSS Ray Reconstruction for viewport denoising
in Cycles. For this to work, scheduling is adjusted to continuously
reset samples (so that independent frames are rendered), pixel jitter
is forced on and the required denoising passes (color, depth, diffuse
albedo, specular albedo, normals, roughness, motion vectors, specular
motion vectors) are enabled.
DLSS expects those inputs in the form of CUDA textures, while Cycles
keeps passes in an interleaved buffer layout. The data therefore has to
be converted, for which specialized versions of the existing denoising
filter kernels are introduced, which read/write directly to temporary
CUDA textures that are managed in denoiser_dlss.cpp.
The integration of DLSS itself is done in a similar fashion to OptiX:
The DLSS SDK is pulled in for the type definitions, but the DLSS
implementation is loaded by the NVIDIA driver installed on the system.
Pull Request: https://projects.blender.org/blender/blender/pulls/153077
Both shadow kernels seems to be expensive enough and having them in the
main module leads to very long module creation times.
These new modules are always created (at least for now), but they are
created in parallel with other modules.
Overall it drastically reduces the time it takes to run OptiX OSL tests
for the first time.
Combined with the previous commits this reduces:
- cycles_bake_optix_osl from 1330.44 sec to 177.30 sec
- cycles_attributes_optix_osl from 2809.71 sec to 501.62 sec
Measured on i9-11900K, NVIDIA RTX 6000 Ada.
Pull Request: https://projects.blender.org/blender/blender/pulls/163706
There was seemingly a mistake in the code which handled the first
task on the same thread from which module was requested to load.
According to the documentation it is only the optixModuleCreateWithTasks()
that needs to be called outside of threading, and it'll try to
exit as soon as possible, leaving with a task that could be run
on a thread.
This should allow creating main module, MNEE, etc in sepaarte
threads.
Enabled with --debug-cycles.
The goal is to provide more tooling for looking into which modules
takes the longest. Comes from the need to investigate OptiX OSL tests
time.
In D16235[^1], Cycles MNEE was only enabled on macOS 13 and above due to
an inefficiency in the calculation of spill requirements (quoting from
the patch), following the minimum macOS deployment target bump to 13.0
in !163627, remove the now unneeded checks this patch introduced, both
on the Cycles feature side, UI side and Cycles test blocklist.
Additionally, as the `__KERNEL_METAL_MACOS__` global define introduced
by this patch didn't have any other additional users, it was also
removed.
[^1]: https://archive.blender.org/developer/D16235
Pull Request: https://projects.blender.org/blender/blender/pulls/163665
Following-up on !163627, which bumped the macOS minimum deployment
target to 13.0, this commit removes the Metal workaround for macOS
versions older than 12, where the 64-bit Cycles kernel features value
was split into two 32-bit function constants.
Now that 64-bit constants are always available, the remaining
`KernelData_kernel_features_64bit` was also renammed to
`KernelData_kernel_features`.
Pull Request: https://projects.blender.org/blender/blender/pulls/163663
Following-up on !163627 the GPUAddressHelper present in
device/metal/util.mm was (a back-compatibiliy helper for getting
gpuAddress & gpuResourceID on macOS < 13.0) isn't required anymore since
the deployment target bump to 13.0.
Remove it in this sense, moving its 13.0+ variants (direct usage of the
newer gpuAddress/gpuResourceID APIs) to the write_resource() templated
functions in queue.mm.
Pull Request: https://projects.blender.org/blender/blender/pulls/163628
This commit raises the mimium macOS deployment target from the current
11.2 to 13.0, allowing broader use of the macOS SDK API functions
without @available checks, and removing @available blocks <= 13.0.
This also aligns our minimum deployment on our "official" macOS minimum
version[^1]. To be specific, this change was sparked by the Jolt Physics
library Metal Compute backend requiring API introduced in 13.0 with no
easy @available workaround.
[^1]: https://www.blender.org/download/requirements/
Pull Request: https://projects.blender.org/blender/blender/pulls/163627
Set it in the device info, the logging is the only place that uses it
for OptiX currently as CUDA is a separate device type, unlike other
backends.
Also don't print it for denoising devices, it's irrelevant.
This issue was introduced in 86374b6b10.
Pull Request: https://projects.blender.org/blender/blender/pulls/163579
The pack only contains bands 1 to 3. The evaluation function has
been renamed as part of the original review. It makes sense to be
consistent, although a bit more verbose.
Pull Request: https://projects.blender.org/blender/blender/pulls/163446
Includes possibility of adding spherical harmonics on geometry, as well
as includes spherical harmonics evaluation.
Spherical harmonics are stored quantized to 8 bit integer, and always
use degree 3. The reason for the limited degree is the memory: degree
3 requires relatively small amount of coefficients (45 int8_t, which
is similar to how realtime renderers pack them into float3x4). Going
higher dimension adds a lot more coefficients to store, and so far it
doesn't seem there are gsplat data-sets that are trained to a degree
higher than 3.
The attribute is currently unused. It is preparation work foe the
gsplat project to make reviews easier.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163409
The main goal is to bring a bit of structure to the use of quaternions
in the Cycles kernel.
Prior to this change there was no dedicated quaternion type and float4
was used instead, and quaternions are stored as (x, y, z, w) matching
naming between the quaternion and float4 fields. However, Blender uses
(w, x, y, z) quaternion order for attributes, which is different from
what Cycles uses. Without dedicated type this either leads to different
quaternion orders depending whether it comes from attribute, or makes
it intrinsically not possible to use implicit sharing.
This change introduces Cycles Quaternion type which is compatible with
Blender attribute math::Quaternion in both layout and alignment, making
it possible to benefit from implicit sharing and use a nice structure
in the kernel.
This change does not modify the existing quaternion usage in the kernel
which is currently used for transform decomposition and interpolation.
There is currently no SIMD for the Quaternion type, which allows to
avoid any special alignment requirement and share attributes with
Blender, but performance might not be ideal.
This attribute type will be used for representing gaussian splat
rotation.
The attribute access and interpolation matches behavior prior to this
change. Added some basic tests for quaternion attribute access for RGB
and alpha, mesh and volume attributes.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163342
A bug that existed in Level Zero loader versions up to 1.28.2 (fixed by
https://github.com/oneapi-src/level-zero/pull/435). It causes a crash
when calling `zeInitDrivers` a second time when no Level Zero drivers
are available. In Cycles and the Compositor module, this behavior is
triggered when `sycl::platform::get_platforms()` is called (in the case
of the Compositor module transitively via OIDN). SYCL loads both the
Level Zero v1 and v2 adapters, which both call the problematic function.
As a workaround, the SYCL UR Level Zero v1 adapter library files are
no longer bundled and the `SYCL_UR_USE_LEVEL_ZERO_V2=1` environment
variable is set to force devices to use the v2 adapter that would
otherwise default to v1. The v2 adapter has been shipped by Blender
starting with 5.0 (using DPC++ version 6.2), but has been available
in DPC++ since version 6.0.
Pull Request: https://projects.blender.org/blender/blender/pulls/163087
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
It is possible that the Metal acceleration structure does not exist
for a previously created BVHMetal. This happens when the BVH is first
created for empty mesh, and then vertices are added to it while the
viewport is in rendered mode. The assert might also happen in other
cases: when render is cancelled prior to BVH is fully built.
Pull Request: https://projects.blender.org/blender/blender/pulls/161790
MultiDevice::has_unified_image_memory needs to return true only if all
devices have unified image memory, since this is used to determine if
we need to copy image memory to the devices at all.
Renamed functions and added a comment to clarify things.
Regression from 44b4fd061e.
Pull Request: https://projects.blender.org/blender/blender/pulls/161682
Using balanced for all geometry type gives good memory saving, but
increases render time for Attic and Bistro due to spatial splitting.
This change mitigates it by using High Quality build for triangles
that keeps spatial splits enabled. For the hair it uses balanced
build as hair BVH is the most memory hungry.
The goal of the change is to reduce memory usage and potentially lower
the memory bandwidth in the BVH traversal.
The bandwidth is a bit tricky to predict, as the number of fetches is
probably the same, but of a lower size (int instead of int2). At a very
least it'll be possible to remove duplicated information.
While the core idea of the implementation matches BVH2, there are a few
extra indirection arrays with offset and primitive types.
Having it a more built-in feature to the BVH itself would be much more
preferred, as without it it is quite hard to have good quality BVH with
a low memory footprint.
As for the BVH2, implementing heuristics from the STBVH paper could be
nice. Or, make the inner nodes aware of the time range. However, the
relevance of BVH2 is becoming lower and lower. Perhaps, it is time to
simplify it as well, and fully rely on the good quality HW-RT.
These checks would lead into primitive indices mismatch between the
actual BVH state and the objects primitives offset calculation.
It is not currently a problem as the HIP-RT BVH uses its own index
arrays. However, these arrays do require extra memory, and could
lead to extra fetch indirections from the BVH traversal.
The need of these checks is not really clear, and it doesn't seem that
other backends perform them. These checks exists from the beginning of
the HIP-RT integration.
The reasoning for the curves control points check is also unclear,
and it would be good to have an example where it is actually needed
and know what problem it solves. Is it some sort of memory optimization
or was it intended to solve numerical intersection problem.
Adds a new specular motion vector pass, for use by denoisers. The
algorithm is described in Ray Tracing Gems chapter 32.4.3 (basically a
reflected object is projected through the reflector, so that the image
of that object can be used like a normal primary hit for motion
calculation).
Pull Request: https://projects.blender.org/blender/blender/pulls/159256
When a GPU device's driver does not meet Blender's minimum required
version, the device is now shown in the preferences as a greyed-out
entry with a message inside the brackets, indicating which driver
version is needed, instead of being silently hidden.
This helps users understand why their GPU is not available for
rendering and what action they can take to resolve the situation.
Pull Request: https://projects.blender.org/blender/blender/pulls/159405
This particular combination was not computing the stack size correctly,
since shader raytrace was moved to a separate OptiX module. It needs to
be set manually now that it's not part of the same module.
Regression from c51fcf73a7.
Pull Request: https://projects.blender.org/blender/blender/pulls/160318
Unfortunately it appears that moving MNEE to another kernel did not
fundamentally fix the apparent compiler bug that breaks this. Another
refactor in 0baa98866c made the bug surface again.
It appears to work fine with HIP-RT, so we leave that case enabled.
HIP-RT is also enabled by default, so it's not as bad.
Pull Request: https://projects.blender.org/blender/blender/pulls/160110
After some refactoring in #159499, disabling OSL now results in compile
errors because the code tries to use "osl_volume_module" that is only
available when OSL support is enabled.
Expand the "WITH_OSL" guard to include this code as well.
Pull Request: https://projects.blender.org/blender/blender/pulls/159649
This helps especially for OSL, where load time is very long. By
splitting off shader raytrace, MNEE and volumes, the first time
rendering is much faster.
The downside is that noinline functions will be duplicated. However OSL
startup performance is very bad currently and this seems the better
trade-off for now. There are ways to make this work if we do not mark
noinline functions as static, but this will require some bigger code
reorganization.
Without OSL, shader raytrace and MNEE were combined in a single module,
and that has been split into two as well.
Besides a better user experience, This will fix OptiX OSL test timeouts
on the buildbot, where some tests need to compute a few different
specializations and the 600s timeout is exceeded, with the longest test
run time being around 300s.
Pull Request: https://projects.blender.org/blender/blender/pulls/159499
This is a mandatory step, which is needed to be done after our
recent IGC upgrade for Blender 5.2 LTS - to ensure compatibility
between generated IGC binaries and their execution on the
end user system.
Pull Request: https://projects.blender.org/blender/blender/pulls/159416
Introduce `MTLResidencySet` to explicitly manage GPU memory residency on macOS 15.0+ devices.
This provides the Metal/GPU driver with a clear list of resources to optimise memory handling, potentially improving performance by reducing overhead from hundreds of `useResource` calls.
Memory allocation and deallocation operations are now routed through new wrapper functions that conditionally add or remove resources from the residency set. Explicit `useResource` calls are bypassed when residency sets are active, as resource residency is now managed through the `MTLResidencySet` API. A debug flag is added to enable or disable this feature at runtime (via env var `CYCLES_METAL_RESIDENCY_SETS=0`).
Pull Request: https://projects.blender.org/blender/blender/pulls/158558
Motion is now part of the position attribute. On the kernel side, a
combined position + radius attribute is now stored, replacing the
previous motion only attribute.
Legacy motion attributes are now removed.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728