Each thread in the launched block runs over a portion of the counter
array. First to calculate a sum of each portion, then summing those across
all threads to get the base offset of each portion, before each thread
does the prefix sum of its local portion again including that base offset.
This could be improved further, but it already solves negative performance
impact when increasing the number of states in the following commit.
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
The new raycast visibility is not like the other visibility flags, which
cover all possible light transport rays and so PATH_RAY_VISIBILITY_ALL
was a complete mask of those. But raycast rays are separate from that,
and should not affect e.g. volume stack interseciton.
Now define a separate PATH_RAY_VISIBILITY_OBJECT_ALL that indicates all
the possible flags that can be set on an object.
Issue introduced in 6e7a0eee32.
Pull Request: https://projects.blender.org/blender/blender/pulls/164137
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.
This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.
The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from
build_files/config/pipeline_config.yaml
to
build_files/buildbot/config/blender_version.cmake
Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.
Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:
- There is an easy to follow table of per-architecture or family
toolkits.
- If there is no architectural preference, all provided toolkits are
probed. For SM kernels newer toolkits are tested first, and for
COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
default one is used.
A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).
Pull Request: https://projects.blender.org/blender/blender/pulls/164002
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
This appears to have been caused by invalid use of assert() instead of
kernel_assert() in the kernel. I don't think it's actually hitting that
assert, but maybe it generated an unsupported instruction or something
along those lines?
The one that caused the actual problem is in svm/convert.h. I changed
more instances that are currently CPU only but risk becoming enabled on
the CPU with future code changes.
Thanks to Sahar A. Kashi for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/163970
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.
This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.
Pull Request: https://projects.blender.org/blender/blender/pulls/162785
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
Commit 976ca9d244 fixed the "Hardware
Ray-Tracing" state message being printed to the log always showing off
for OptiX. This improves that further by querying whether the device
actually supports it, since OptiX has a software fallback on devices
without hardware ray tracing support.
Pull Request: https://projects.blender.org/blender/blender/pulls/163807
Adds the option to use DLSS Ray Reconstruction for viewport denoising
in Cycles. For this to work, scheduling is adjusted to continuously
reset samples (so that independent frames are rendered), pixel jitter
is forced on and the required denoising passes (color, depth, diffuse
albedo, specular albedo, normals, roughness, motion vectors, specular
motion vectors) are enabled.
DLSS expects those inputs in the form of CUDA textures, while Cycles
keeps passes in an interleaved buffer layout. The data therefore has to
be converted, for which specialized versions of the existing denoising
filter kernels are introduced, which read/write directly to temporary
CUDA textures that are managed in denoiser_dlss.cpp.
The integration of DLSS itself is done in a similar fashion to OptiX:
The DLSS SDK is pulled in for the type definitions, but the DLSS
implementation is loaded by the NVIDIA driver installed on the system.
Pull Request: https://projects.blender.org/blender/blender/pulls/153077
Both shadow kernels seems to be expensive enough and having them in the
main module leads to very long module creation times.
These new modules are always created (at least for now), but they are
created in parallel with other modules.
Overall it drastically reduces the time it takes to run OptiX OSL tests
for the first time.
Combined with the previous commits this reduces:
- cycles_bake_optix_osl from 1330.44 sec to 177.30 sec
- cycles_attributes_optix_osl from 2809.71 sec to 501.62 sec
Measured on i9-11900K, NVIDIA RTX 6000 Ada.
Pull Request: https://projects.blender.org/blender/blender/pulls/163706
Following-up on !163627, which bumped the macOS minimum deployment
target to 13.0, this commit removes the Metal workaround for macOS
versions older than 12, where the 64-bit Cycles kernel features value
was split into two 32-bit function constants.
Now that 64-bit constants are always available, the remaining
`KernelData_kernel_features_64bit` was also renammed to
`KernelData_kernel_features`.
Pull Request: https://projects.blender.org/blender/blender/pulls/163663
It is unclear why there needs to be point_intersect() in the
scene_intersect(): it feels that it mainly does redundant work.
The main part that is needed is to set the prim, type, and u,v
coordinates.
There might be some non-obvious reason for this to exist, but
it really needs to be done explicitly with a comment.
This call was quietly introduced by !111795.
Pull Request: https://projects.blender.org/blender/blender/pulls/162808
Previously the prim_offset was passed as a geometry user data, which
makes it impossible to implement custom primitives: custom geometry
primitive requires a geometry bounds function and this function has
only access to the geometry user data (despite of the custom pointer
being accepted by rtcSetGeometryBoundsFunction: this pointer is being
simply ignored).
This change makes it possible to set geometry user data that can be
accessed from the bounds function in a way that does not conflict
with the previous prim_offset logic.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162738
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
While it makes some code a bit longer, but it also makes parts more
explicit w.r.t where things are coming from.
More importantly, it allows to have Cycles-specific data-types named
the same as Metal standard library calls them. For example, this change
will allow having own float3x3 type, which might use Metal's one for an
underlying implementation.
Ref #159470
There is limited number of bits available for the user-specified OptiX
hit kind, and currently we used of all those bits available.
The easiest is to use a custom enum for custom types as there is no
real need to couple it it PrimitiveType: the hit kind is only used to
check what kind of intersection it is to fetch extra information and
the actual primitive type for intersection is stored on the OptiX
payload.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162647
This PR begins to replace the use of Blender's "BLI_kdopbvh" BVH tree
for meshes with Embree. Embree is a much newer implementation and and
has much better performance characteristics. In particular, building
the BVH is over 5 times faster. Raycasting against the BVH and closest
point sampling is faster as well.
With the extensive usage of the old BVH API across Blender, it's
unrealistic to do a complete replacement in a single PR. This change
focuses on the triangle surface BVH, in cases where the replacement
is relatively obvious. The idea is that enough uses are covered so
that in most use cases only the Embree BVH will be needed.
Because Embree is an optional dependency, there also needs to be a
fallback to use the old API. If we ever wanted to remove the old BVH
implementation we could discuss making it a mandatory dependency.
Test results are different in a few cases because Embree chooses a
different triangle for the arbitrary choice of closest triangle on
either side of an edge. I also had to increase the threshold for the
hair dynamics test, where small differences in any intermediate values
are chaotically amplified.
Next steps are described in #161529.
This was investigated before in !108148.
Pull Request: https://projects.blender.org/blender/blender/pulls/156408
Fix integer overflows that were causing this, in the kernel and in tx
generation. The overflow caused both wrong render results and slowness
due to too many tiles being loaded due to wrong derivatives.
Also fix an additional overflows in CPU image sampling without the cache.
Pull Request: https://projects.blender.org/blender/blender/pulls/162195
Workaround an apparent compiler bug in CUDA 12.8, change noinline to
inline. This appears already fixed in CUDA 12.9, but for backporting to
5.2 LTS a workaround is safer.
Pull Request: https://projects.blender.org/blender/blender/pulls/162161
This change switches BVH traversal/intersection code from using custom
index arrays that are specific to HIP-RT builder to the general data
arrays that are available on all backends. This is possible since there
are no primitive splitting based on the BVH time steps.
Some functions (like set_intersect_point) still have the same number of
fetches, but some of the fetches are now int and not int2.
Other hot functions (like motion_triangle_custom_intersect) now have
only 2 fetches instead of 4.
The goal of the change is to reduce memory usage and potentially lower
the memory bandwidth in the BVH traversal.
The bandwidth is a bit tricky to predict, as the number of fetches is
probably the same, but of a lower size (int instead of int2). At a very
least it'll be possible to remove duplicated information.
While the core idea of the implementation matches BVH2, there are a few
extra indirection arrays with offset and primitive types.
Having it a more built-in feature to the BVH itself would be much more
preferred, as without it it is quite hard to have good quality BVH with
a low memory footprint.
As for the BVH2, implementing heuristics from the STBVH paper could be
nice. Or, make the inner nodes aware of the time range. However, the
relevance of BVH2 is becoming lower and lower. Perhaps, it is time to
simplify it as well, and fully rely on the good quality HW-RT.
Store positions as separate arrays again. This appears to give better
memory access patterns or layout. This was a regression from fc9917352b
and 0baa98866c.
This is a relatively simple change, rerouting the position attribute to
dedicated arrays, still including motion steps in the same array. The rest
of the refactor is preserved.
Pull Request: https://projects.blender.org/blender/blender/pulls/160110
This helps especially for OSL, where load time is very long. By
splitting off shader raytrace, MNEE and volumes, the first time
rendering is much faster.
The downside is that noinline functions will be duplicated. However OSL
startup performance is very bad currently and this seems the better
trade-off for now. There are ways to make this work if we do not mark
noinline functions as static, but this will require some bigger code
reorganization.
Without OSL, shader raytrace and MNEE were combined in a single module,
and that has been split into two as well.
Besides a better user experience, This will fix OptiX OSL test timeouts
on the buildbot, where some tests need to compute a few different
specializations and the 600s timeout is exceeded, with the longest test
run time being around 300s.
Pull Request: https://projects.blender.org/blender/blender/pulls/159499
Motion is now part of the position attribute. On the kernel side, a
combined position + radius attribute is now stored, replacing the
previous motion only attribute.
Legacy motion attributes are now removed.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.
MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.
MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.
This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
Either use no spaces (common for disabled arguments) `/*arg*/`
or spaces `/* Regular comment. */`
Cleanup lop-sided comments such as `/* X*/` or `/*X */`.
Also use doxy-sections for bmesh_structure.hh,
the ad-hoc section comments had become outdated.
Ref !158467
Use sycl::inclusive_scan_over_group instead of the group_ballot
extension in adaptive_sampling_convergence_check and
integrator_shadow_catcher_count_possible_splits.
The current DPC++ implementation of the extension only supports
devices with sub-group sizes up to 64. The inclusive scan works
also on devices with larger sub-group sizes.
Fixes the behaviour of the two affected kernels on the Snapdragon X
Elite GPU.
Pull Request: https://projects.blender.org/blender/blender/pulls/155801
Adds basic infrastructure for denoisers that can also upscale for
viewport rendering, by taking advantage of the existing resolution
divider functionality.
The implementation basically tracks two resolution dividers, the normal
one and an additional one that also has a denoiser upscale factor
applied. It then uses the latter one for rendering and processing on
noisy render buffers, and the former one for any processing happening
after denoising. This has the advantage of allowing an additional
resolution divider on top of denoiser upscaling (so can also set the
pixel size to something other than 1x and that is still being
respected). The resolution divider was made into a floating point value
to allow fractional scale factors.
Pull Request: https://projects.blender.org/blender/blender/pulls/151133
* Detect tx files associated with images
* Add tiled image loading in ImageCache and OIIOImageLoader
* Pad tiles for fast hardware texture filtering
* Kernel side tile mapping with stochastic mip selection
* Statistics support for mipmaps and tiles loaded
Ref #68917
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
Functions scope their variables, macros don't. Use functions
for better variable hygiene, with `PARENT_SCOPE` where output
variables need to propagate.
This avoids the need to `unset(...)` every local variable which is
easy to forget - leaking variables into the callers scope.
Macros that must remain (flag modification, `find_package` forwarding,
paired state, caller-scope `return()`) are annotated with the reason
in their doc-strings.
This is an internal change preparing for the texture cache. Only implemented
for surfaces, and currently supports the following nodes:
- Geometry
- Tangent
- Mapping
- Attribute
- Texture Coordinate
- Environment Texture
- Image Texture
- Vector math
- UV Map
- Combine/Separate XYZ
- Bump
This has some impact on GPU rendering performance. Various changes were
made to optimize this, but rendering can still be a few % slower on some
GPUs. Some of the optimizations done:
* Use different node types enum for derivative nodes, consecutive to ensure the
jump table works.
* Template various SVM derivative nodes to separate them from the
non-derivative case, and avoid using dual types in those implementations.
* Use template on return type for stack_store and stack_load to make the above
easier to implement.
* Use template on return type of primitive attribute reading to make derivative
and non-derivative variations.
* Unify derivative and bump dx/dy nodes. Now it's a single derivative node
that handles both cases.
* Derivative nodes are disabled in volume shaders for now.
* Tweak inlining on a few functions.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/155706
Previous implementation assumed that ocloc binary was working by
only checking if the directory for it exists. This was replaced
with proper cascade checks of the ocloc binary functionality to
avoid misleading warnings about
"binaries for X not supported by Intel ocloc" later when we
actually try to use ocloc binary, which existence we had
not even checked beforehand.
Pull Request: https://projects.blender.org/blender/blender/pulls/154669
The local atomic sort kernels initialize shared memory arrays sized
max_shaders, but used if (local_id < max_shaders) to do so, meaning
only one entry per thread was initialized.
With smaller workgroup sizes, further entries were left uninitialized
and caused out of bounds writes.
Replace the single-slot if-guard with a strided for loop so all
max_shaders entries are correctly initialized regardless of
workgroup size.
Pull Request: https://projects.blender.org/blender/blender/pulls/155269
Refactor ImageHandle to used pointers to a new ImageSlot istead of slot
indices. ImageSlot has subclasses ImageSingle and ImageUDIM, for
individual images and UDIMs. ImageUDIM has a list of handles for
ImageSingle.
UDIM tile lookup was moved out of SVM and OSL code, and is now shared.
Early loading of images for displacement and volumes was refactored.
Now a set of ImageSingle pointers is collected, which are then loaded
by the image manager.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.
Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
This PR fixes the compilation problems when building Cycles without
OSL support and some Windows-related compile problems caused by the
'_USE_MATH_DEFINES' define.
The PR:
- Adds the cycles_util dependencies to the cycles cpu module.
- Adds the `_USE_MATH_DEFINES` as a global compiler definition when
building under Windows.
- Removes unneed setting and definitions of `_USE_MATH_DEFINES` in the code
and CMake files
Pull Request: https://projects.blender.org/blender/blender/pulls/154209