Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
Adds the option to use DLSS Ray Reconstruction for viewport denoising
in Cycles. For this to work, scheduling is adjusted to continuously
reset samples (so that independent frames are rendered), pixel jitter
is forced on and the required denoising passes (color, depth, diffuse
albedo, specular albedo, normals, roughness, motion vectors, specular
motion vectors) are enabled.
DLSS expects those inputs in the form of CUDA textures, while Cycles
keeps passes in an interleaved buffer layout. The data therefore has to
be converted, for which specialized versions of the existing denoising
filter kernels are introduced, which read/write directly to temporary
CUDA textures that are managed in denoiser_dlss.cpp.
The integration of DLSS itself is done in a similar fashion to OptiX:
The DLSS SDK is pulled in for the type definitions, but the DLSS
implementation is loaded by the NVIDIA driver installed on the system.
Pull Request: https://projects.blender.org/blender/blender/pulls/153077
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
This PR begins to replace the use of Blender's "BLI_kdopbvh" BVH tree
for meshes with Embree. Embree is a much newer implementation and and
has much better performance characteristics. In particular, building
the BVH is over 5 times faster. Raycasting against the BVH and closest
point sampling is faster as well.
With the extensive usage of the old BVH API across Blender, it's
unrealistic to do a complete replacement in a single PR. This change
focuses on the triangle surface BVH, in cases where the replacement
is relatively obvious. The idea is that enough uses are covered so
that in most use cases only the Embree BVH will be needed.
Because Embree is an optional dependency, there also needs to be a
fallback to use the old API. If we ever wanted to remove the old BVH
implementation we could discuss making it a mandatory dependency.
Test results are different in a few cases because Embree chooses a
different triangle for the arbitrary choice of closest triangle on
either side of an edge. I also had to increase the threshold for the
hair dynamics test, where small differences in any intermediate values
are chaotically amplified.
Next steps are described in #161529.
This was investigated before in !108148.
Pull Request: https://projects.blender.org/blender/blender/pulls/156408
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.
MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.
MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.
This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
Use sycl::inclusive_scan_over_group instead of the group_ballot
extension in adaptive_sampling_convergence_check and
integrator_shadow_catcher_count_possible_splits.
The current DPC++ implementation of the extension only supports
devices with sub-group sizes up to 64. The inclusive scan works
also on devices with larger sub-group sizes.
Fixes the behaviour of the two affected kernels on the Snapdragon X
Elite GPU.
Pull Request: https://projects.blender.org/blender/blender/pulls/155801
Previous implementation assumed that ocloc binary was working by
only checking if the directory for it exists. This was replaced
with proper cascade checks of the ocloc binary functionality to
avoid misleading warnings about
"binaries for X not supported by Intel ocloc" later when we
actually try to use ocloc binary, which existence we had
not even checked beforehand.
Pull Request: https://projects.blender.org/blender/blender/pulls/154669
To suppress unused variable compiler warnings when Cycles is build without
path guiding support the `ccl_attr_maybe_unused` macro is added to the
parameters of the affected function definitions.
Pull Request: https://projects.blender.org/blender/blender/pulls/153193
Host execution of OneAPI devices was used for development/debugging. It hasn't been working lately and only adds complexity.
Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/153650
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.
Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.
A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.
The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.
LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.
The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.
Pull Request: https://projects.blender.org/blender/blender/pulls/152649
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
Split the device-specific logic into individual files that reside
in kernel/device/<device>. Should be no functional change, but it
should make it easier to work with individual devices easier.
Some minor changes compared to prior to this change:
- There is an explicit target cycles_kernel_cpu.
It makes it easier pass header dependencies to GPU backends, but
also makes it possible to only try compile CPU kernels when GPU
binaries are enabled.
- There is a CMake option to allow building different device backends
in parallel: WITH_CYCLES_PARALLEL_DEVICE_KERNEL_BUILD. It is set
to OFF by default, matching old behavior.
Setting it to ON helps in situations when memory is not a concern,
or when it is only a couple of GPU architectures enabled.
Pull Request: https://projects.blender.org/blender/blender/pulls/152241
Making heavier use of specialization constants in SYCL for Embree.
This reduces code size of the intersection kernels and bring
performance improvement up to 9% in some scenes on
Intel GPUs.
Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Co-authored-by: Nikita Sirgienko <nikita.sirgienko@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/141559
Null Scattering currently has performance and noise issues, and it will
take time to address them. For now add the previous Ray Marching back as
an option.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/146317
Experiments have shown that the OptiX denoiser performs best when
operating on images that have their origin at the top-left corner,
while Blender renders with the origin at the bottom-left corner.
Simply flipping the image vertically before and after denoising is a
relatively trivial operation, so this patch introduces this as an
additional preprocessing and postprocessing step for denoising when the
OptiX denoiser is used. Additionally, this patch also removes an unused
helper function, now that OptiX 8.0 is the minimum.
Pull Request: https://projects.blender.org/blender/blender/pulls/145358
Guide the probability to scatter in or transmit through the volume.
Only applied for primary rays.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Add new "Linear 3D Curves" option in the Curves panel in the render
properties. This renders curves as linear segments rather than smooth
curves, for faster render time at the cost of accuracy.
On NVIDIA Blackwell GPUs, this can give a 6x speedup compared to smooth
curves, due to hardware acceleration. On NVIDIA Ada there is still
a 3x speedup, and CPU and other GPU backends will also render this
faster.
A difference with smooth curves is that these have end caps, as this
was simpler to implement and they are usually helpful anyway.
In the future this functionality will also be used to properly support
the CURVE_TYPE_POLY on the new curves object.
Pull Request: https://projects.blender.org/blender/blender/pulls/139735
The code of the "oneapi_load_kernels" function before this modification
was loading kernels and compiling them, if needed, for all devices in
the associated GPU context. This makes sense for one GPU execution
scenario, as well as for execution scenario of multi identical GPU,
but in cases where Blender users have several different GPUs in
render, the previous implementation would compile all kernels
for all devices for each device, unnecessarily doing the same
work multiple times. Because of this, I am changing the
implementation so that now compilation happens only for the used
device per used device, ensuring that no unnecessary work is done.
No render performance changes are expected.
With these changes, we can now mark devices which are expected to work as
performant as possible, and devices which were not optimized for some reason.
For example, because the device was released after the Blender release,
making it impossible for developers to optimize for devices in already
released unchangeable code. This is primarily relevant for the LTS versions,
which are supported for two years and require proper communication about
optimization status for the new devices released during this time.
This is implemented for oneAPI devices. Other device types currently are
marked as optimized for compatibility with old behavior, but may implement
the same in the future.
Pull Request: https://projects.blender.org/blender/blender/pulls/139751
Now ccl_device sets inlining and ccl_device_inline forces inlining.
This matches more closely with what is currently done for cuda and metal
backends.
I've measured from 1% to 6% overall performance improvement in rendering
benchmark scenes on Arc B580, as well as a small decrease in compile
time.
The attribute handling code in the kernel is currently highly duplicated since
it needs to handle five different data types and we couldn't use templates
back then.
We can now, so might as well make use of it and get rid of ~1000 lines.
There are also some small fixes for the GPU OSL code:
- Wrong derivative for .w component when converting float2/float3->float4
- Different conversion for float2->float (CPU averages, GPU used to take .x)
- Removed useless code for converting to float2, not used by OSL
Pull Request: https://projects.blender.org/blender/blender/pulls/134694
The current usage of software-based texture operations in
the oneAPI implementation puts additional register pressure on
the GPU compiler during register allocation. And it also creates
code that requires maintenance. This commit is intended to address
this situation by utilizing a recently productized SYCL bindless
texture API to enable HW-based texture operations using
Intel GPUs' hardware sampler.
This currently translates to 1-11% rendering speedups (scene-specific)
on my Arc A770 and Arc B580. At the moment, there are small
performance regressions with NanoVDB texture operations on Arc B580
and small performance regressions in shade surface MNEE and Raytrace
kernels on Arc A770, but they look recoverable and will be handled
in the future.
Pull Request: https://projects.blender.org/blender/blender/pulls/133457
There is now a non-experimental API for this_work_item functionality, so
let's use it for better code quality and also to avoid the deprecation
warning during compilation.
No functional or performance changes are expected.
Pull Request: https://projects.blender.org/blender/blender/pulls/133472
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.
Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
In 891d71a4d4 this keyword was
dropped due to performance regression after
fdc2962beb, but currently code
does not experience this performance degradation, and in fact
there is minor performance improvement on Lunar Lake GPUs,
along with an expected improvement in compile time.
However, this change brings a minor performance regression to
shade_surface kernel on Intel Arc and Meteor Lake GPUs, which
will be solved later by disabling this keyword for
these platforms only.
Pull Request: https://projects.blender.org/blender/blender/pulls/130299
Previously, when compiling on Rocky Linux 8 with fno-honor-nans, compile
time was more than 5x longer than expected, and there was an unresolved
symbol to __sqrtf_finite in GPU binaries.
Once defining sqrtf in compat.h, both issues are effectively gone, this
was certainly due to problematic interactions with build system's math
library headers.
So we can remove current workaround of defining fhonor-nans, and now
have the same set of flags on both Windows and Linux.
The kernel zeroing memory since we've added host memory fallback didn't
expect large inputs, so with these scenes, it was running into
"Provided range is out of integer limits. Pass
`-fno-sycl-id-queries-fit-in-int' to disable range check" error.
This kernel was used instead of memset to avoid some issues with the
free_memory queries not always being updated.
As we can't reproduce these with recent drivers, we now use memset,
which fixes rendering with BVH2.
This enables scenes with all textures not fitting in GPU
memory to finally render. For scenes that are fitting,
no functional change or performance change is expected.
Pull Request: https://projects.blender.org/blender/blender/pulls/122385
fdc2962beb indirectly introduced a change
in inlining (light_tree_pdf started getting inlined) that led to a 5-10%
drop in performance for most scenes.
Dropping the noinline keyword for oneAPI device recovers it.
It however brings another performance regression to MNEE and Raytrace
kernels, that we'll look into separately.
Along with the 4.1 libraries upgrade, we are bumping the clang-format
version from 8-12 to 17. This affects quite a few files.
If not already the case, you may consider pointing your IDE to the
clang-format binary bundled with the Blender precompiled libraries.
The NanoVDB headers are not compatible with Metal due to missing address
space qualifiers. We currently have a big patch for NanoVDB header
files, which is difficult to update for OpenVDB 11. Instead extract a
few hundred lines of code from NanoVDB to do just what we need.
Pull Request: https://projects.blender.org/blender/blender/pulls/115992
This makes the GPU tricubic implementation more efficient. The dense
grid code implemented this in terms of trilinear lookups that are
hardware accelerated, but for NanoVDB this just causes unnecessary voxel
reads. Instead match the CPU code.
Pull Request: https://projects.blender.org/blender/blender/pulls/115992
OpenImageDenoise V2 comes with GPU support for various backends. This adds a new class, OIDNDenoiserGPU, in order to add this functionality into the existing Cycles post processing pipeline without having to change it much. OptiX and OIDN CPU denoising remain as they are. Rendering on a supported Intel GPU will automatically select the GPU denoiser.
Device support is initially limited to the oneAPI devices that are supported by Cycles, but can be extended.
Ref #115045
Co-authored-by: Stefan Werner <stefan.werner@intel.com>
Co-authored-by: Ray Molenkamp <github@lazydodo.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/108314
Speckles and missing lights were experienced in scenes with Nishita Sky
Texture and a Sun Size smaller than 1.5°, such as in Lone Monk and Attic
scenes.
We previously worked around these by using a more precise
software implementation of cosine.
After recent changes in Cycles, it turns out this workaround isn't
currently needed.
The API for the kernels library is defined, there is no need to
export more than that. This change only affects linux since hidden
visiblity is the default on Windows.