Commit graph

111 commits

Author SHA1 Message Date
Patrick Mours
9433b655c3 Cycles: GPU: Optimize parallel prefix sum kernel
Each thread in the launched block runs over a portion of the counter
array. First to calculate a sum of each portion, then summing those across
all threads to get the base offset of each portion, before each thread
does the prefix sum of its local portion again including that base offset.

This could be improved further, but it already solves negative performance
impact when increasing the number of states in the following commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/163930
2026-09-23 15:22:44 +02:00
Alex Fuller
3ff1f508e6 Cycles: Fallback value for missing image or attribute in shader
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.

This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.

Pull Request: https://projects.blender.org/blender/blender/pulls/162785
2026-09-15 19:45:36 +02:00
Patrick Mours
bfb5775388 Cycles: Enable local shader sorting for CUDA/OptiX
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.

Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.

Pull Request: https://projects.blender.org/blender/blender/pulls/163436
2026-09-15 12:39:35 +02:00
Patrick Mours
e1fa0c2805 Cycles: Add NVIDIA DLSS support for viewport denoising
Adds the option to use DLSS Ray Reconstruction for viewport denoising
in Cycles. For this to work, scheduling is adjusted to continuously
reset samples (so that independent frames are rendered), pixel jitter
is forced on and the required denoising passes (color, depth, diffuse
albedo, specular albedo, normals, roughness, motion vectors, specular
motion vectors) are enabled.
DLSS expects those inputs in the form of CUDA textures, while Cycles
keeps passes in an interleaved buffer layout. The data therefore has to
be converted, for which specialized versions of the existing denoising
filter kernels are introduced, which read/write directly to temporary
CUDA textures that are managed in denoiser_dlss.cpp.

The integration of DLSS itself is done in a similar fashion to OptiX:
The DLSS SDK is pulled in for the type definitions, but the DLSS
implementation is loaded by the NVIDIA driver installed on the system.

Pull Request: https://projects.blender.org/blender/blender/pulls/153077
2026-09-10 14:17:46 +02:00
Brecht Van Lommel
90e93e577e Fix #161492: Cycles CUDA/OptiX wrong bicubic texture sample on Blackwell
Workaround an apparent compiler bug in CUDA 12.8, change noinline to
inline. This appears already fixed in CUDA 12.9, but for backporting to
5.2 LTS a workaround is safer.

Pull Request: https://projects.blender.org/blender/blender/pulls/162161
2026-08-03 14:45:16 +02:00
Brecht Van Lommel
5fc85d1cb7 Refactor: Cycles: Move MNEE walk into separate kernel
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.

MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.

MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.

This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.

Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
2026-05-27 20:34:06 +02:00
Rafal Bielski
d454d502b2 Cycles: oneAPI: Use sycl::inclusive_scan_over_group instead of ballot
Use sycl::inclusive_scan_over_group instead of the group_ballot
extension in adaptive_sampling_convergence_check and
integrator_shadow_catcher_count_possible_splits.

The current DPC++ implementation of the extension only supports
devices with sub-group sizes up to 64. The inclusive  scan works
also on devices with larger sub-group sizes.

Fixes the behaviour of the two affected kernels on the Snapdragon X
Elite GPU.

Pull Request: https://projects.blender.org/blender/blender/pulls/155801
2026-04-08 12:58:17 +02:00
Patrick Mours
02171cc356 Cycles: Add fundamental support for upscaling denoisers
Adds basic infrastructure for denoisers that can also upscale for
viewport rendering, by taking advantage of the existing resolution
divider functionality.

The implementation basically tracks two resolution dividers, the normal
one and an additional one that also has a denoiser upscale factor
applied. It then uses the latter one for rendering and processing on
noisy render buffers, and the former one for any processing happening
after denoising. This has the advantage of allowing an additional
resolution divider on top of denoiser upscaling (so can also set the
pixel size to something other than 1x and that is still being
respected). The resolution divider was made into a floating point value
to allow fractional scale factors.

Pull Request: https://projects.blender.org/blender/blender/pulls/151133
2026-03-31 14:05:28 +02:00
Brecht Van Lommel
b68c056e2c Cycles: Texture cache tiled image loading
* Detect tx files associated with images
* Add tiled image loading in ImageCache and OIIOImageLoader
* Pad tiles for fast hardware texture filtering
* Kernel side tile mapping with stochastic mip selection
* Statistics support for mipmaps and tiles loaded

Ref #68917

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Brecht Van Lommel
fa383aa511 Refactor: Cycles: Texture cache miss handling
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Weizhen Huang
f2394a431e Cycles: Support automatic differentiation of shader nodes in SVM
This is an internal change preparing for the texture cache. Only implemented
for surfaces, and currently supports the following nodes:

- Geometry
- Tangent
- Mapping
- Attribute
- Texture Coordinate
- Environment Texture
- Image Texture
- Vector math
- UV Map
- Combine/Separate XYZ
- Bump

This has some impact on GPU rendering performance. Various changes were
made to optimize this, but rendering can still be a few % slower on some
GPUs. Some of the optimizations done:

* Use different node types enum for derivative nodes, consecutive to ensure the
  jump table works.
* Template various SVM derivative nodes to separate them from the
  non-derivative case, and avoid using dual types in those implementations.
* Use template on return type for stack_store and stack_load to make the above
  easier to implement.
* Use template on return type of primitive attribute reading to make derivative
  and non-derivative variations.
* Unify derivative and bump dx/dy nodes. Now it's a single derivative node
  that handles both cases.
* Derivative nodes are disabled in volume shaders for now.
* Tweak inlining on a few functions.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>

Pull Request: https://projects.blender.org/blender/blender/pulls/155706
2026-03-16 18:20:54 +01:00
Brecht Van Lommel
74cec5a868 Merge branch 'blender-v5.1-release' 2026-03-06 22:32:17 +01:00
Sean Stirling
4cbe9ce2fe Fix: Cycles: GPU rendering error with more than 1024 shaders
The local atomic sort kernels initialize shared memory arrays sized
max_shaders, but used if (local_id < max_shaders) to do so, meaning
only one entry per thread was initialized.

With smaller workgroup sizes, further entries were left uninitialized
and caused out of bounds writes.

Replace the single-slot if-guard with a strided for loop so all
max_shaders entries are correctly initialized regardless of
workgroup size.

Pull Request: https://projects.blender.org/blender/blender/pulls/155269
2026-03-06 22:31:21 +01:00
Brecht Van Lommel
1ffb43fc69 Refactor: Cycles: Move UDIM code out of SVM, refactor ImageHandle
Refactor ImageHandle to used pointers to a new ImageSlot istead of slot
indices. ImageSlot has subclasses ImageSingle and ImageUDIM, for
individual images and UDIMs. ImageUDIM has a list of handles for
ImageSingle.

UDIM tile lookup was moved out of SVM and OSL code, and is now shared.

Early loading of images for displacement and volumes was refactored.
Now a set of ImageSingle pointers is collected, which are then loaded
by the image manager.

Pull Request: https://projects.blender.org/blender/blender/pulls/154668
2026-02-25 17:07:49 +01:00
Brecht Van Lommel
84d4a2e9cf Refactor: Cycles: Split KernelImageTexture and KernelImageInfo
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.

Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.

Pull Request: https://projects.blender.org/blender/blender/pulls/154668
2026-02-25 17:07:39 +01:00
Stefan Werner
63607b7bfc Cycles: Removed OneAPI host device support
Host execution of OneAPI devices was used for development/debugging. It hasn't been working lately and only adds complexity.

Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/153650
2026-01-30 14:00:11 +01:00
Sergey Sharybin
bc90c64ec8 Cleanup: Unify HW-RT defines in Cycles
This change unifies the define names that are used to indicate HW-RT
nature of kernels:

- __HIPRT__ is renamed to __KERNEL_HIPRT__
- __METALRT__ is renamed to __KERNEL_METALRT__

This makes them named similar to __KERNEL_OPTIX__.

One might argue that it is weird to have multiple KERNEL defines at
the same time (like __KERNEL_METAL__ and __KERNEL_METALRT__, or
__KERNEL_SSE4__ and __KERNEL_AVX2__) More correct name is more like
KERNEL_CAPABILITY or KERNEL_FEATURE, but such rename is outside of
the scope of this commit. For now we just follow existing naming.
2026-01-26 14:20:24 +01:00
Brecht Van Lommel
4b34743b4e Cycles: Perform direct light shader eval in own kernel
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.

Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.

A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.

The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.

LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.

The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.

Pull Request: https://projects.blender.org/blender/blender/pulls/152649
2026-01-20 20:34:16 +01:00
Brecht Van Lommel
527f9ea306 Refactor: Cycles: More consistent naming of image functions and structs
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.

Pull Request: https://projects.blender.org/blender/blender/pulls/152665
2026-01-14 17:57:46 +01:00
Sergey Sharybin
0230a528f8 Refactor: Cycles, remove non-raytacing kernels from HIP-RT
Initially, KIP-RT module essentially included all kernels, not only the
ones that do intersection checks. This changes makes it so the HIP-RT
module only includes kernels that use intersections checks, making it
behave closer to how CUDA and OptiX modules are structured.

There does not seem to be direct impact on the performance, but perhaps
it could still help compiler to avoid making bad decisions. It does
help with the compilation: building the HIP-RT kernel for gfx1201 goes
down from 1min 47sec to 58sec on the test machine (i5-8600K).

The uncompressed kernel_rt_gfx1201.hipfb goes down from 6.5Mb to 3.2Mb,
and compressed size goes down from 1Mb to 790Kb.

The HIP-RT's kernel.cpp is now structured closer to how gpu/kernel.h
is structured, potentially making it easier to move the signatures to
the common place.

There is also potential to split shader ray-trace kernel to separate
module, following OptiX for potentially better performance as well.
However, from quick tests that does not give any immediate speedup
measured when removing shade raytrace code from the module.

Pull Request: https://projects.blender.org/blender/blender/pulls/152081
2025-12-29 12:16:01 +01:00
Weizhen Huang
2b0a1cae06 Cycles: Add an option to use ray marching for volume rendering
Null Scattering currently has performance and noise issues, and it will
take time to address them. For now add the previous Ray Marching back as
an option.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/146317
2025-09-26 12:14:45 +02:00
Patrick Mours
b4bb075285 Cycles: Flip image vertically before passing to OptiX denoiser to improve result quality
Experiments have shown that the OptiX denoiser performs best when
operating on images that have their origin at the top-left corner,
while Blender renders with the origin at the bottom-left corner.
Simply flipping the image vertically before and after denoising is a
relatively trivial operation, so this patch introduces this as an
additional preprocessing and postprocessing step for denoising when the
OptiX denoiser is used. Additionally, this patch also removes an unused
helper function, now that OptiX 8.0 is the minimum.

Pull Request: https://projects.blender.org/blender/blender/pulls/145358
2025-09-04 16:04:23 +02:00
Weizhen Huang
a4f8e0bfa2 Cycles: Use RGBE for denoised guiding buffers to reduce memory usage
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
2025-08-13 10:28:50 +02:00
Weizhen Huang
5cb6014efd Cycles: Volume Scattering Probability Guiding
Guide the probability to scatter in or transmit through the volume.
Only applied for primary rays.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>
2025-08-13 10:28:50 +02:00
Weizhen Huang
b2b2d9a4f3 Cycles: Render volume by ray marching through octrees
One octree per volume per shader based on the density. In preparation
for the null scattering
2025-08-13 10:28:50 +02:00
Hugh Delaney
930a942dd0 Refactor: Cycles: Move block sizes into common header
This change puts all the block size macros in the same common header, so
they can be included in host side code without needing to also include
the kernels that are defined in the device headers that contained these
values.

This change also removes a magic number used to enqueue a kernel, which
happened to agree with the GPU_PARALLEL_SORT_BLOCK_SIZE macro.

Pull Request: https://projects.blender.org/blender/blender/pulls/143646
2025-08-01 13:26:02 +02:00
Brecht Van Lommel
4c25b49875 Refactor: Cycles: Deduplicate 3D texture sampling between devices
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Brecht Van Lommel
b6c4233b28 Refactor: Cycles: Remove now unused 3D image texture support
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Brecht Van Lommel
7978799e6f Cycles: Always render volume as NanoVDB
All GPU backends now support NanoVDB, using our own kernel side code
that is easily portable. This simplifies kernel and device code.

Volume bounds are now built from the NanoVDB grid instead of OpenVDB,
to avoid having to keep around the OpenVDB grid after loading.

While this reduces memory usage, it does have a performance impact,
particularly for the Cubic filter. That will be addressed by
another commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Campbell Barton
07121d44ae Cleanup: use braces (follow own style guide) 2025-06-11 09:05:26 +00:00
Brecht Van Lommel
0e7a696819 Cleanup: Unused arguments in Cycles kernel
And add back the compiler flag that hid them.

Pull Request: https://projects.blender.org/blender/blender/pulls/139497
2025-05-27 21:30:45 +02:00
Nikita Sirgienko
a0b7ad436b Cleanup: Cycles: oneAPI: Switch to non-experimental work item API
There is now a non-experimental API for this_work_item functionality, so
let's use it for better code quality and also to avoid the deprecation
warning during compilation.

No functional or performance changes are expected.

Pull Request: https://projects.blender.org/blender/blender/pulls/133472
2025-02-12 21:46:22 +01:00
Brecht Van Lommel
57ff24cb99 Refactor: Cycles: Add const keyword to more function parameters
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:24 +01:00
Brecht Van Lommel
dd51c8660b Refactor: Cycles: Add const keyword where possible, using clang-tidy
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.

Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:20 +01:00
Brecht Van Lommel
3c2a6fbb9c Refactor: Cycles: Use nullptr instead of NULL
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:43 +01:00
Thomas Dinges
22e16ca096 Cycles: add make_float4(float3 a, float b) type
This resolves a todo from the code. Part of the Quality Project.

Pull Request: https://projects.blender.org/blender/blender/pulls/131915
2024-12-17 09:11:08 +01:00
Nikita Sirgienko
fb21f3fb56 Cleanup: Cycles: oneAPI: Fix deprecation warnings about get_pointer() 2024-10-01 22:26:15 +02:00
Michael Jones
99f5433445 Cycles: Dormant fixes for adaptive feature compilation
This PR fixes the (currently unused) scene-based selective feature compilation macros. These feature based macros haven't been used for a few years, and enabling them currently results in compilation errors.

The only functional change in this PR is in geom/primitive.h where undef-ing `__HAIR__` had exposed an inconsistency in how pointcloud attributes were being fetched. Using the more general `primitive_surface_attribute_float4` (instead of `curve_attribute_float4`) fixed a compilation error that occurred when rendering pointcloud unit test scenes with adaptive compilation enabled.

Pull Request: https://projects.blender.org/blender/blender/pulls/121216
2024-04-30 12:56:22 +02:00
Stefan Werner
31d55e87f9 Cycles: Metal support for OpenImageDenoise
This is supported on Apple Silicon GPUs and macOS 13.0+.

Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Co-authored-by: Attila Afra <attila.t.afra@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/116124
2024-02-06 21:13:23 +01:00
Brecht Van Lommel
d015e98ee6 Fix Cycles ASAN error with boolean kernel arguments 2023-12-12 13:27:36 +01:00
Brecht Van Lommel
6cdb43195e Refactor: replace NanoVDB kernel side implementation by own code
The NanoVDB headers are not compatible with Metal due to missing address
space qualifiers. We currently have a big patch for NanoVDB header
files, which is difficult to update for OpenVDB 11. Instead extract a
few hundred lines of code from NanoVDB to do just what we need.

Pull Request: https://projects.blender.org/blender/blender/pulls/115992
2023-12-10 19:37:36 +01:00
Brecht Van Lommel
8ba474dc4f Refactor: replace NanoVDB SampleFromVoxels by own code
This makes the GPU tricubic implementation more efficient. The dense
grid code implemented this in terms of trilinear lookups that are
hardware accelerated, but for NanoVDB this just causes unnecessary voxel
reads. Instead match the CPU code.

Pull Request: https://projects.blender.org/blender/blender/pulls/115992
2023-12-10 19:37:36 +01:00
Campbell Barton
7f34ad736a Cleanup: spelling in comments 2023-08-05 13:54:25 +10:00
Campbell Barton
c12994612b License headers: use SPDX-FileCopyrightText in intern/cycles 2023-06-14 16:53:23 +10:00
Sergey Sharybin
ba3f26fac5 Cycles: light and shadow linking
With light linking, lights can be set to affect only specific objects in the
scene. Shadow linking additionally gives control over which objects acts a
shadow blockers for a light.

Usage:
https://wiki.blender.org/wiki/Reference/Release_Notes/4.0/Cycles

Implementation:
https://wiki.blender.org/wiki/Source/Render/Cycles/LightLinking

Ref #104972
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
2023-05-24 14:11:47 +02:00
Campbell Barton
bf36a61e62 Cleanup: spelling in comments & some corrections 2023-05-20 21:17:09 +10:00
Nikita Sirgienko
bafd82c9c1 Cycles: oneAPI: use local memory for faster shader sorting
Co-authored-by: Stefan Werner <stefan.werner@intel.com>

Pull Request: https://projects.blender.org/blender/blender/pulls/107994
2023-05-17 11:07:57 +02:00
Campbell Barton
6859bb6e67 Cleanup: format (with BraceWrapping::AfterControlStatement "MultiLine") 2023-05-02 09:37:49 +10:00
Sahar A. Kashi
557a245dd5 Cycles: add HIP RT device, for AMD hardware ray tracing on Windows
HIP RT enables AMD hardware ray tracing on RDNA2 and above, and falls back to a
to shader implementation for older graphics cards. It offers an average 25%
sample rendering rate improvement in Cycles benchmarks, on a W6800 card.

The ray tracing feature functions are accessed through HIP RT SDK, available on
GPUOpen. HIP RT traversal functionality is pre-compiled in bitcode format and
shipped with the SDK.

This is not yet enabled as there are issues to be resolved, but landing the
code now makes testing and further changes easier.

Known limitations:
* Not working yet with current public AMD drivers.
* Visual artifact in motion blur.
* One of the buffers allocated for traversal has a static size. Allocating it
  dynamically would reduce memory usage.
* This is for Windows only currently, no Linux support.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>

Ref #105538
2023-04-25 20:19:43 +02:00
Xavier Hallade
70892e82ac Cycles: oneAPI: use specialization constant to compile with/without Embree on GPU 2023-04-18 22:09:42 +02:00