Commit graph

73 commits

Author SHA1 Message Date
Sergey Sharybin
819d52f3d2 GSplat: Initial rendering support for Cycles
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.

Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.

For the Embree and hardware intersection it is implemented as a custom
primitive type.

The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.

There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.

Ref #159470

Pull Request: https://projects.blender.org/blender/blender/pulls/163103
2026-09-16 16:52:10 +02:00
Sergey Sharybin
cb737eb519 Cycles: HIP-RT: Always use balanced build for BLAS
Value predictable performance and memory usage over potential speedup
that comes with a huge memory penalty over small performance boost.
2026-07-14 09:17:45 +02:00
Sergey Sharybin
a11363e86e Cycles: HIP-RT: Remove custom arrays for primitive indices 2026-07-14 09:17:45 +02:00
Sergey Sharybin
0c23bccf43 Cycles: HIP-RT: Reduce indirections in the BVH traversal
This change switches BVH traversal/intersection code from using custom
index arrays that are specific to HIP-RT builder to the general data
arrays that are available on all backends. This is possible since there
are no primitive splitting based on the BVH time steps.

Some functions (like set_intersect_point) still have the same number of
fetches, but some of the fetches are now int and not int2.

Other hot functions (like motion_triangle_custom_intersect) now have
only 2 fetches instead of 4.
2026-07-14 09:17:45 +02:00
Sergey Sharybin
823b172a47 Cycles: HIP-RT: Remove BVH steps
The goal of the change is to reduce memory usage and potentially lower
the memory bandwidth in the BVH traversal.

The bandwidth is a bit tricky to predict, as the number of fetches is
probably the same, but of a lower size (int instead of int2). At a very
least it'll be possible to remove duplicated information.

While the core idea of the implementation matches BVH2, there are a few
extra indirection arrays with offset and primitive types.

Having it a more built-in feature to the BVH itself would be much more
preferred, as without it it is quite hard to have good quality BVH with
a low memory footprint.

As for the BVH2, implementing heuristics from the STBVH paper could be
nice. Or, make the inner nodes aware of the time range. However, the
relevance of BVH2 is becoming lower and lower. Perhaps, it is time to
simplify it as well, and fully rely on the good quality HW-RT.
2026-07-14 09:17:45 +02:00
Brecht Van Lommel
5fc85d1cb7 Refactor: Cycles: Move MNEE walk into separate kernel
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.

MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.

MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.

This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.

Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
2026-05-27 20:34:06 +02:00
Sergey Sharybin
d61fa843f5 Refactor: Cycles: Rename path visibility flags
Prefix all bits that are mapped to the BVH visibility mask with
`PATH_RAY_VISIBILITY_`.

Should be no functional changes.

Ref !157822
2026-05-27 19:22:53 +02:00
salipour
2bc06ee047 Fix #150055: Cycles HIP-RT subsurface scattering and motion blur issue
The scene_intersect_local function was using the wrong macro copied from
the BVH2 implementation.

Pull Request: https://projects.blender.org/blender/blender/pulls/158663
2026-05-18 16:14:25 +02:00
Campbell Barton
4df665a2fd Cleanup: blank space around C-style comments
Either use no spaces (common for disabled arguments) `/*arg*/`
or spaces `/* Regular comment. */`

Cleanup lop-sided comments such as `/* X*/` or `/*X */`.

Also use doxy-sections for bmesh_structure.hh,
the ad-hoc section comments had become outdated.

Ref !158467
2026-05-12 16:42:20 +10:00
Brecht Van Lommel
bac976d38f Refactor: Cycles: Add writeable global GPU data arrays
This will be used for tile requests in the texture cache.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Sergey Sharybin
6c3a61abb3 Refactor: Cycles, allow granular intersection tests in filter function
Should be no functional changes, preparing for an upcoming refactor.
2026-02-13 17:18:56 +01:00
Brecht Van Lommel
51ff328467 Fix: Cycles HIP-RT incorrect volume intersect filter with motion triangle
The intersect functions should return false to ignore this, unlike the filter
functions which return true to skip.

This issue was introduced in f7407a1cb0.

Pull Request: https://projects.blender.org/blender/blender/pulls/153934
2026-02-05 15:25:31 +01:00
Sahar A. Kashi
1062b57da7 Cycles: HIP-RT, copy missing kernel files to the binary folder
When `WITH_CYCLES_HIP_BINARIES` is disabled, Cycles compiles kernels directly
from the sources located in `scripts\addons_core\cycles\source\kernel`.
Some of the files required for HIP-RT compilation were missing from the CMake
configuration; this pull request adds those missing files.

A regression from !152241

Pull Request: https://projects.blender.org/blender/blender/pulls/153628
2026-02-02 09:06:57 +01:00
Campbell Barton
5549bc5723 CMake: use set(VAR "") for empty initialization, unset() for cleanup
In CMake, set(VAR) without a value unsets the variable rather than
initializing it to empty. This causes issues with --warn-uninitialized
and can lead to unexpected behavior when the variable is used later.

Use set(VAR "") when initializing variables that will be appended to
or used in expressions. Use unset() to explicitly remove temporary
variables after use.

Ref !153503
2026-01-29 11:48:33 +11:00
Sergey Sharybin
f7407a1cb0 Refactor: De-duplicate volume intersection filtering in Cycles
Move the logic to a common place.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
5b5d9f2f0c Refactor: De-duplicate Cycles shadow_all HW-RT function
The goal is to re-use as much of non-trivial logic across devices as
possible, ideally including data-structures and their initialization.

The payload is a bit bigger than it was and it does not utilize OptiX
registers. It seems to be a bit hard to get reliable numbers as they
fluctuate quite a bit and depend on the order of renders, but the
worst performance impact is about 1.5% on the Bistro scene.

It feels there is a room for improvements, like relying on OptiX
itself to keep track on the maximum intersection distance (the code
does accept intersection, but it still tracks the maximum distance in
the filter function), and, maybe, putting payload back on the register.
However, it'll be a bit harder to replicate such behavior on other
platforms, so perhaps it is not a bad trade-off to have.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
9fac41b046 Refactor: Cycles HIP-RT file structure
- Consolidate everything into a single bvh.h file.

  There is no need to have a separate header file with the intersection
  implementation.

- Remove use of macros.

  Either use force-inlined function, or inline the core which macro was
  calling. Solves inconsistency in the way how traversal is invoked:
  some functions were using macro, others had explicit HIP-RT call.

  This change makes it so there is hidden dependency on the local
  variables.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
eef459e74b Refactor: Remove duplicate function in Cycles HIP-RT
Use intersection_ray_valid() which is defined in the bvh/util.h
2026-01-26 14:20:24 +01:00
Sergey Sharybin
3a5f47b6af Refactor: Simplify transparent shadow API in Cycles
Before this change the API consisted of two parts: return value which
indicated whether an opaque surface was hit, and throughput which was
used to accumulate curves transparency.

This change makes it so throughput=0 indicates that an opaque hit was
found, allowing to simplify intersection payload and state tracking:
the payload for Embree and MetalRT now has one less boolean flag.
For OptiX there is no direct affect on the payload size, as there is
some explicit rule about Ray using registers p6 and p7 (while the
boolean flag was on p5). It does, however, avoids need to track extra
boolean flag regardless.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
bc90c64ec8 Cleanup: Unify HW-RT defines in Cycles
This change unifies the define names that are used to indicate HW-RT
nature of kernels:

- __HIPRT__ is renamed to __KERNEL_HIPRT__
- __METALRT__ is renamed to __KERNEL_METALRT__

This makes them named similar to __KERNEL_OPTIX__.

One might argue that it is weird to have multiple KERNEL defines at
the same time (like __KERNEL_METAL__ and __KERNEL_METALRT__, or
__KERNEL_SSE4__ and __KERNEL_AVX2__) More correct name is more like
KERNEL_CAPABILITY or KERNEL_FEATURE, but such rename is outside of
the scope of this commit. For now we just follow existing naming.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
0d189d5bcc Refactor: Remove KernelGlobals from HIP-RT payload
The KernelGlobals are "fake" on the most of the GPU backends,
consisting of only a padding member in them, with hopes that
the compiler optimizes out KernelGlobals function arguments.

The HIP-RT is the only backend where it is used for something:
it stores memory for the stack traversal, which is only needed
by the HIP-RT library.

This change removes KernelGlobals from the payload, making it
smaller and more closely matching payload of, say, MetalRT.
The code creates an explicit local variable, keeping function
calls locally readable, hoping that the compiler optimizes
these variables out.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
96d326c01c Refactor: Remove primitive_type from HIP-RT payload
Move handling related to packed segment type to the conversion of
hiprtHit to Intersection. The benefits of doing so include:

- Reduces the size of the payload
- Solves confusion with point clouds: the previous code that was
  handling isect->type applied it for all non-triangle primitives
  while comment was only referring to curves.
- Makes payload look more like the one for Metal-RT, making it
  easier to eventually unify the payload and code.

There are a bit of extra fetching involved into intersection after
this change. For now it is not a concern, as the code is planned to
be more heavily refactored.
2026-01-26 14:20:24 +01:00
Brecht Van Lommel
527f9ea306 Refactor: Cycles: More consistent naming of image functions and structs
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.

Pull Request: https://projects.blender.org/blender/blender/pulls/152665
2026-01-14 17:57:46 +01:00
Sergey Sharybin
b2cee9709f Refactor: Split Cycles kernel CMakeLists.txt
Split the device-specific logic into individual files that reside
in kernel/device/<device>. Should be no functional change, but it
should make it easier to work with individual devices easier.

Some minor changes compared to prior to this change:

- There is an explicit target cycles_kernel_cpu.

  It makes it easier pass header dependencies to GPU backends, but
  also makes it possible to only try compile CPU kernels when GPU
  binaries are enabled.

- There is a CMake option to allow building different device backends
  in parallel: WITH_CYCLES_PARALLEL_DEVICE_KERNEL_BUILD. It is set
  to OFF by default, matching old behavior.

  Setting it to ON helps in situations when memory is not a concern,
  or when it is only a couple of GPU architectures enabled.

Pull Request: https://projects.blender.org/blender/blender/pulls/152241
2026-01-02 17:29:29 +01:00
Sergey Sharybin
0230a528f8 Refactor: Cycles, remove non-raytacing kernels from HIP-RT
Initially, KIP-RT module essentially included all kernels, not only the
ones that do intersection checks. This changes makes it so the HIP-RT
module only includes kernels that use intersections checks, making it
behave closer to how CUDA and OptiX modules are structured.

There does not seem to be direct impact on the performance, but perhaps
it could still help compiler to avoid making bad decisions. It does
help with the compilation: building the HIP-RT kernel for gfx1201 goes
down from 1min 47sec to 58sec on the test machine (i5-8600K).

The uncompressed kernel_rt_gfx1201.hipfb goes down from 6.5Mb to 3.2Mb,
and compressed size goes down from 1Mb to 790Kb.

The HIP-RT's kernel.cpp is now structured closer to how gpu/kernel.h
is structured, potentially making it easier to move the signatures to
the common place.

There is also potential to split shader ray-trace kernel to separate
module, following OptiX for potentially better performance as well.
However, from quick tests that does not give any immediate speedup
measured when removing shade raytrace code from the module.

Pull Request: https://projects.blender.org/blender/blender/pulls/152081
2025-12-29 12:16:01 +01:00
Sergey Sharybin
2c4477de04 Cleanup: Cycles, sign conversion
A series of commits which reduces the number of sign conversions
(int <-> uint) in the Cycles kernel.

While it is not expected that the conversion emits any instructions,
it is quite confusing to follow the code and choose proper type.

Additionally, from some development in !151540 it seemed that such
mismatch was responsible for the performance drop in HIP-RT.

Pull Request: https://projects.blender.org/blender/blender/pulls/152009
2025-12-29 12:13:11 +01:00
Weizhen Huang
188294d8ad Fix #151249: Cycles recording non-volume triangles when intersecting volume
Should check the shader instead of just the object, since an object can
have multiple shaders

Pull Request: https://projects.blender.org/blender/blender/pulls/151499
2025-12-15 12:34:47 +01:00
Brecht Van Lommel
4ae19df885 Cleanup: Cycles: Fix building without some kernel features for debugging
Pull Request: https://projects.blender.org/blender/blender/pulls/151307
2025-12-08 14:09:26 +01:00
Brecht Van Lommel
0e7a696819 Cleanup: Unused arguments in Cycles kernel
And add back the compiler flag that hid them.

Pull Request: https://projects.blender.org/blender/blender/pulls/139497
2025-05-27 21:30:45 +02:00
Weizhen Huang
23c762e388 Fix: Cycles: Do not count volume bounds bounce as transparent
In forward path tracing, when we pass volume bounding meshes, we
accumulate `volume_bounds_bounce`. We should match this behaviour in NEE
instead of accumulating `transparent_bounce`.

Pull Request: https://projects.blender.org/blender/blender/pulls/137556
2025-04-24 13:10:33 +02:00
Sergey Sharybin
36559fd89f Fix #136811: HIP-RT performance regression in 4.5
Reduce the register pressure and branching in the switch() by using
subclass and cast from void* to the base class.

This ensures intersection functions are not inlined multiple times,
bringing performance back.

Alternative could be to avoid functions (they are quite large) but
that only partially resolves the performance regression.

Pull Request: https://projects.blender.org/blender/blender/pulls/136823
2025-04-01 17:59:44 +02:00
Campbell Barton
42ad772a1f Cleanup: spelling & repeated terms (make check_spelling_*)
Also use comment blocks for English text.
2025-03-27 01:13:34 +00:00
Sergey Sharybin
2ab231d802 Refactor: Pass proper KernelGlobals
HIP-RT functions do have access to kg, and it was used inconsistently:
some functions were passed actual kg, other were passed nullptr.

This change makes it consistent and passes kg everywhere.

Pull Request: https://projects.blender.org/blender/blender/pulls/136503
2025-03-26 11:07:06 +01:00
Sergey Sharybin
709371b278 Refactor: Avoid creation of local copy of RaySelfPrimitives 2025-03-26 11:07:04 +01:00
Sergey Sharybin
888c7e1df9 Cleanup: Avoid redundant data fetch 2025-03-26 11:07:04 +01:00
Sergey Sharybin
3d882acee2 Cleanup: Else after return 2025-03-26 11:07:04 +01:00
Sergey Sharybin
b2dd523d0d Cleanup: Avoid default hit initialization
The entire object is assigned later on, no need to initialize it.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
323e27d825 Cleanup: Remove redundant assignment
The payload stores pointers, no need to restore pointer
of the function argument to the same value.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
e92a8042c3 Refactor: Payload for shadow intersection and filter in HIP-RT
The code before this change was relying on the ShadowPayload have
the same "header" as RayPayload for some of the primitive types
(curve, motion triangle, point): intersection functions were shared
between "regular" and shadow rays (shadow in this case is shadow_all),
but extra filter function was used for shadow rays.

This is fragile if someone changes one of these structures. What is
worse is that compiler might actually decide to shuffle things in
some structs, or remove unused fields.

This change also solves confusion about ShadowPayload::prim_type
seemingly only being assigned to PRIMITIVE_NONE. With time it is
not impossible that compiler will also see this, and constant-fold
some checks, or even remove the field. If that happens then the
render result will be wrong. Maybe it is already happening as there
are some GPU and driver and optimization flag specific bugs in the
area.

It is unclear whether it was causing any actual problem: W7800
seems to render all hair correctly on Linux.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
cdb3f34944 Cleanup: Use full name for the primitive_type
Makes it extra clear locally type of what the variable contains:
primitive, ray, or something else.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
72542f3bb4 Cleanup: Follow Blender style and use more const
Also make some style decisions more consistent: for example,
the way how stop/continue search return value is commented.
Prefer lower vertical space for those.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
bf9c95f164 Cleanup: Move payload type cast to caller in HIP-RT
Mainly readability purposes:
- Having variables called local_payload is ambiguous: does it refer to
  LocalPayload type or to a variable be local in a function?
- Some of the functions are used for different ray types, so having the
  type case in intersectFunc and filterFunc makes it easier to scan.

For the latter: now it is more obvious that Curve_Intersect_Shadow
expects RayPayload, but Curve_Filter_Shadow expects ShadowPayload.
It might not be a problem currently as ShadowPayload has the same
"header" RayPayload, but it might change in the future. Also, compiler
might optimize fields out from one but not from the other.
2025-03-26 11:07:04 +01:00
Sergey Sharybin
3daaf21bab Cleanup: Remove unused function argument in HIP-RT 2025-03-26 11:07:04 +01:00
Sergey Sharybin
50180283e9 Fix #117527: Spatial split leads to artifacts on transparent shadows
The reason for this to happen is because when spatial split is used
the same intersection could be recorded twice (via different BVH nodes).

This change introduces check for the intersection being already recoded,
similar to the check in the local BVH. The check is done during BVH
intersection which allows to properly ignore intersections even for the
maximum bounce number check. A faster approach would be to do such
filtering after sorting, but then we can not keep bounce check in the
BVH code consistent with and without spatial splits.

Intuitively it seems that it should be possible to merge the new loop
with the one that checks for which intersection to keep. But it is not
so trivial in practice: it doesn't run for all intersections, and also
it is formulated in a way that updates isect_index for the next record.

Pull Request: https://projects.blender.org/blender/blender/pulls/136251
2025-03-21 13:56:50 +01:00
Sergey Sharybin
bf65b64708 Refactor: De-duplicate local intersection reservoir sampling logic
The code which was checking whether local intersection is to be
recorded, and under which index was duplicated for triangles,
motion triangles, and HIP-RT triangle filter function.

This change moves the common logic to an utility function which
is reused from all the places mentioned above.

Pull Request: https://projects.blender.org/blender/blender/pulls/136244
2025-03-20 17:19:31 +01:00
Sergey Sharybin
7165146fb2 Cleanup: More spelling fixes in comments 2025-03-20 10:37:09 +01:00
Sergey Sharybin
ae4f6026dc Cleanup: Spelling in comments 2025-03-20 10:36:12 +01:00
Sahar A. Kashi
9ad3b74867 Fix: SSS and Motion Blur or Curves not working on HIP-RT
This change fixes the remaining failing tests with SSS when using HIP-RT.
This includes crash when SSS is used on curves, and objects with motion
blur and SSS rendering black.

The root cause for both cases was the fact that traversal was always
assuming regular BVH (built for triangles), while curves and motion
triangles are using custom primitives, which requires specialized BVH
traversal.

This change includes:

- Early output from `scene_intersect_local()` for non-triangle and
  non-motion-triangle primitives. This fixes `sss_hair.blend` test,
  and also avoids unnecessary BVH traversal when the local intersection
  is requested from curve object. The same early-output could be added
  to other BVH traversal implementation.

- Use `hiprtGeomCustomTraversalAnyHitCustomStack` for motion triangles
  primitives. This fixes motion blur on objects with SSS render black.

Fixes #135856

Co-authored-by: Sahar A. Kashi <sahar.alipourkashi@amd.com>
Co-authored-by: Sergey Sharybin <sergey@blender.org>

Pull Request: https://projects.blender.org/blender/blender/pulls/135943
2025-03-14 18:17:54 +01:00
Sergey Sharybin
a3eb0faa3f Fix: Incorrect ray time used for HIP-RT local intersections
It was always hard-coded to be 0.

It does not seem to result in any extra tests passing, but they are
probably not sophisticated enough.

Noticed while looking into details for the #135856.

Pull Request: https://projects.blender.org/blender/blender/pulls/135878
2025-03-12 19:23:38 +01:00
Alaska
d840d249b3 Cycles: Re-enable HIPRT point cloud rendering
Previously point cloud rendering was disabled on the HIPRT backend due
to unexpected performance regressions introduce by it.

With the recent update to HIP SDK 6.3 and HIPRT 2.5, these performance
regressions have been resolved and so this commit re-enables
point cloud rendering on HIPRT.

Pull Request: https://projects.blender.org/blender/blender/pulls/134902
2025-02-27 00:01:35 +01:00