Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
This change switches BVH traversal/intersection code from using custom
index arrays that are specific to HIP-RT builder to the general data
arrays that are available on all backends. This is possible since there
are no primitive splitting based on the BVH time steps.
Some functions (like set_intersect_point) still have the same number of
fetches, but some of the fetches are now int and not int2.
Other hot functions (like motion_triangle_custom_intersect) now have
only 2 fetches instead of 4.
The goal of the change is to reduce memory usage and potentially lower
the memory bandwidth in the BVH traversal.
The bandwidth is a bit tricky to predict, as the number of fetches is
probably the same, but of a lower size (int instead of int2). At a very
least it'll be possible to remove duplicated information.
While the core idea of the implementation matches BVH2, there are a few
extra indirection arrays with offset and primitive types.
Having it a more built-in feature to the BVH itself would be much more
preferred, as without it it is quite hard to have good quality BVH with
a low memory footprint.
As for the BVH2, implementing heuristics from the STBVH paper could be
nice. Or, make the inner nodes aware of the time range. However, the
relevance of BVH2 is becoming lower and lower. Perhaps, it is time to
simplify it as well, and fully rely on the good quality HW-RT.
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.
MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.
MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.
This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
Either use no spaces (common for disabled arguments) `/*arg*/`
or spaces `/* Regular comment. */`
Cleanup lop-sided comments such as `/* X*/` or `/*X */`.
Also use doxy-sections for bmesh_structure.hh,
the ad-hoc section comments had become outdated.
Ref !158467
When `WITH_CYCLES_HIP_BINARIES` is disabled, Cycles compiles kernels directly
from the sources located in `scripts\addons_core\cycles\source\kernel`.
Some of the files required for HIP-RT compilation were missing from the CMake
configuration; this pull request adds those missing files.
A regression from !152241
Pull Request: https://projects.blender.org/blender/blender/pulls/153628
In CMake, set(VAR) without a value unsets the variable rather than
initializing it to empty. This causes issues with --warn-uninitialized
and can lead to unexpected behavior when the variable is used later.
Use set(VAR "") when initializing variables that will be appended to
or used in expressions. Use unset() to explicitly remove temporary
variables after use.
Ref !153503
The goal is to re-use as much of non-trivial logic across devices as
possible, ideally including data-structures and their initialization.
The payload is a bit bigger than it was and it does not utilize OptiX
registers. It seems to be a bit hard to get reliable numbers as they
fluctuate quite a bit and depend on the order of renders, but the
worst performance impact is about 1.5% on the Bistro scene.
It feels there is a room for improvements, like relying on OptiX
itself to keep track on the maximum intersection distance (the code
does accept intersection, but it still tracks the maximum distance in
the filter function), and, maybe, putting payload back on the register.
However, it'll be a bit harder to replicate such behavior on other
platforms, so perhaps it is not a bad trade-off to have.
- Consolidate everything into a single bvh.h file.
There is no need to have a separate header file with the intersection
implementation.
- Remove use of macros.
Either use force-inlined function, or inline the core which macro was
calling. Solves inconsistency in the way how traversal is invoked:
some functions were using macro, others had explicit HIP-RT call.
This change makes it so there is hidden dependency on the local
variables.
Before this change the API consisted of two parts: return value which
indicated whether an opaque surface was hit, and throughput which was
used to accumulate curves transparency.
This change makes it so throughput=0 indicates that an opaque hit was
found, allowing to simplify intersection payload and state tracking:
the payload for Embree and MetalRT now has one less boolean flag.
For OptiX there is no direct affect on the payload size, as there is
some explicit rule about Ray using registers p6 and p7 (while the
boolean flag was on p5). It does, however, avoids need to track extra
boolean flag regardless.
This change unifies the define names that are used to indicate HW-RT
nature of kernels:
- __HIPRT__ is renamed to __KERNEL_HIPRT__
- __METALRT__ is renamed to __KERNEL_METALRT__
This makes them named similar to __KERNEL_OPTIX__.
One might argue that it is weird to have multiple KERNEL defines at
the same time (like __KERNEL_METAL__ and __KERNEL_METALRT__, or
__KERNEL_SSE4__ and __KERNEL_AVX2__) More correct name is more like
KERNEL_CAPABILITY or KERNEL_FEATURE, but such rename is outside of
the scope of this commit. For now we just follow existing naming.
The KernelGlobals are "fake" on the most of the GPU backends,
consisting of only a padding member in them, with hopes that
the compiler optimizes out KernelGlobals function arguments.
The HIP-RT is the only backend where it is used for something:
it stores memory for the stack traversal, which is only needed
by the HIP-RT library.
This change removes KernelGlobals from the payload, making it
smaller and more closely matching payload of, say, MetalRT.
The code creates an explicit local variable, keeping function
calls locally readable, hoping that the compiler optimizes
these variables out.
Move handling related to packed segment type to the conversion of
hiprtHit to Intersection. The benefits of doing so include:
- Reduces the size of the payload
- Solves confusion with point clouds: the previous code that was
handling isect->type applied it for all non-triangle primitives
while comment was only referring to curves.
- Makes payload look more like the one for Metal-RT, making it
easier to eventually unify the payload and code.
There are a bit of extra fetching involved into intersection after
this change. For now it is not a concern, as the code is planned to
be more heavily refactored.
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
Split the device-specific logic into individual files that reside
in kernel/device/<device>. Should be no functional change, but it
should make it easier to work with individual devices easier.
Some minor changes compared to prior to this change:
- There is an explicit target cycles_kernel_cpu.
It makes it easier pass header dependencies to GPU backends, but
also makes it possible to only try compile CPU kernels when GPU
binaries are enabled.
- There is a CMake option to allow building different device backends
in parallel: WITH_CYCLES_PARALLEL_DEVICE_KERNEL_BUILD. It is set
to OFF by default, matching old behavior.
Setting it to ON helps in situations when memory is not a concern,
or when it is only a couple of GPU architectures enabled.
Pull Request: https://projects.blender.org/blender/blender/pulls/152241
Initially, KIP-RT module essentially included all kernels, not only the
ones that do intersection checks. This changes makes it so the HIP-RT
module only includes kernels that use intersections checks, making it
behave closer to how CUDA and OptiX modules are structured.
There does not seem to be direct impact on the performance, but perhaps
it could still help compiler to avoid making bad decisions. It does
help with the compilation: building the HIP-RT kernel for gfx1201 goes
down from 1min 47sec to 58sec on the test machine (i5-8600K).
The uncompressed kernel_rt_gfx1201.hipfb goes down from 6.5Mb to 3.2Mb,
and compressed size goes down from 1Mb to 790Kb.
The HIP-RT's kernel.cpp is now structured closer to how gpu/kernel.h
is structured, potentially making it easier to move the signatures to
the common place.
There is also potential to split shader ray-trace kernel to separate
module, following OptiX for potentially better performance as well.
However, from quick tests that does not give any immediate speedup
measured when removing shade raytrace code from the module.
Pull Request: https://projects.blender.org/blender/blender/pulls/152081
A series of commits which reduces the number of sign conversions
(int <-> uint) in the Cycles kernel.
While it is not expected that the conversion emits any instructions,
it is quite confusing to follow the code and choose proper type.
Additionally, from some development in !151540 it seemed that such
mismatch was responsible for the performance drop in HIP-RT.
Pull Request: https://projects.blender.org/blender/blender/pulls/152009
In forward path tracing, when we pass volume bounding meshes, we
accumulate `volume_bounds_bounce`. We should match this behaviour in NEE
instead of accumulating `transparent_bounce`.
Pull Request: https://projects.blender.org/blender/blender/pulls/137556
Reduce the register pressure and branching in the switch() by using
subclass and cast from void* to the base class.
This ensures intersection functions are not inlined multiple times,
bringing performance back.
Alternative could be to avoid functions (they are quite large) but
that only partially resolves the performance regression.
Pull Request: https://projects.blender.org/blender/blender/pulls/136823
HIP-RT functions do have access to kg, and it was used inconsistently:
some functions were passed actual kg, other were passed nullptr.
This change makes it consistent and passes kg everywhere.
Pull Request: https://projects.blender.org/blender/blender/pulls/136503
The code before this change was relying on the ShadowPayload have
the same "header" as RayPayload for some of the primitive types
(curve, motion triangle, point): intersection functions were shared
between "regular" and shadow rays (shadow in this case is shadow_all),
but extra filter function was used for shadow rays.
This is fragile if someone changes one of these structures. What is
worse is that compiler might actually decide to shuffle things in
some structs, or remove unused fields.
This change also solves confusion about ShadowPayload::prim_type
seemingly only being assigned to PRIMITIVE_NONE. With time it is
not impossible that compiler will also see this, and constant-fold
some checks, or even remove the field. If that happens then the
render result will be wrong. Maybe it is already happening as there
are some GPU and driver and optimization flag specific bugs in the
area.
It is unclear whether it was causing any actual problem: W7800
seems to render all hair correctly on Linux.
Also make some style decisions more consistent: for example,
the way how stop/continue search return value is commented.
Prefer lower vertical space for those.
Mainly readability purposes:
- Having variables called local_payload is ambiguous: does it refer to
LocalPayload type or to a variable be local in a function?
- Some of the functions are used for different ray types, so having the
type case in intersectFunc and filterFunc makes it easier to scan.
For the latter: now it is more obvious that Curve_Intersect_Shadow
expects RayPayload, but Curve_Filter_Shadow expects ShadowPayload.
It might not be a problem currently as ShadowPayload has the same
"header" RayPayload, but it might change in the future. Also, compiler
might optimize fields out from one but not from the other.
The reason for this to happen is because when spatial split is used
the same intersection could be recorded twice (via different BVH nodes).
This change introduces check for the intersection being already recoded,
similar to the check in the local BVH. The check is done during BVH
intersection which allows to properly ignore intersections even for the
maximum bounce number check. A faster approach would be to do such
filtering after sorting, but then we can not keep bounce check in the
BVH code consistent with and without spatial splits.
Intuitively it seems that it should be possible to merge the new loop
with the one that checks for which intersection to keep. But it is not
so trivial in practice: it doesn't run for all intersections, and also
it is formulated in a way that updates isect_index for the next record.
Pull Request: https://projects.blender.org/blender/blender/pulls/136251
The code which was checking whether local intersection is to be
recorded, and under which index was duplicated for triangles,
motion triangles, and HIP-RT triangle filter function.
This change moves the common logic to an utility function which
is reused from all the places mentioned above.
Pull Request: https://projects.blender.org/blender/blender/pulls/136244
This change fixes the remaining failing tests with SSS when using HIP-RT.
This includes crash when SSS is used on curves, and objects with motion
blur and SSS rendering black.
The root cause for both cases was the fact that traversal was always
assuming regular BVH (built for triangles), while curves and motion
triangles are using custom primitives, which requires specialized BVH
traversal.
This change includes:
- Early output from `scene_intersect_local()` for non-triangle and
non-motion-triangle primitives. This fixes `sss_hair.blend` test,
and also avoids unnecessary BVH traversal when the local intersection
is requested from curve object. The same early-output could be added
to other BVH traversal implementation.
- Use `hiprtGeomCustomTraversalAnyHitCustomStack` for motion triangles
primitives. This fixes motion blur on objects with SSS render black.
Fixes#135856
Co-authored-by: Sahar A. Kashi <sahar.alipourkashi@amd.com>
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/135943
It was always hard-coded to be 0.
It does not seem to result in any extra tests passing, but they are
probably not sophisticated enough.
Noticed while looking into details for the #135856.
Pull Request: https://projects.blender.org/blender/blender/pulls/135878
Previously point cloud rendering was disabled on the HIPRT backend due
to unexpected performance regressions introduce by it.
With the recent update to HIP SDK 6.3 and HIPRT 2.5, these performance
regressions have been resolved and so this commit re-enables
point cloud rendering on HIPRT.
Pull Request: https://projects.blender.org/blender/blender/pulls/134902