The intel/llvm 7.1 upgrade is giving a large performance regression in
shade_surface kernel on Intel GPUs. I've root-caused it down to a slight
shift in private memory layout due to slightly different inlining.
To fix both this regression and overall performance sensitivity on
private memory layout, that can happen with any code change beyond the
DPC++ upgrade, I'm forcing ShaderData and svm's stack alignments to 64
bytes.
Pull Request: https://projects.blender.org/blender/blender/pulls/164411
This is a bit empirical at the moment. On Linux+Intel Arc A750,
volume_instance test is running out of memory with 0% and all tests are
passing with 1%, 2% sounds reasonable for now.
Pull Request: https://projects.blender.org/blender/blender/pulls/164329
It was notably wrong with volume_instance test.
Example on Linux+Intel Arc A750, before:
GPU queue concurrent states: 14680064, using up to 5.82G
GPU SoA state size: 11.95G
after:
GPU queue concurrent states: 8781824, using up to 7.15G
GPU SoA state size: 7.15G
The integrator state was already device only, but some other memory is
also frequently accessed and should be on the GPU for best performance.
This is a follow up for a pre-existing issue found reviewing #163437 and
#163930.
Pull Request: https://projects.blender.org/blender/blender/pulls/164306
This helps avoid out of memory errors for complex scenes, and improves
performance for smaller scenes with more memory available for states.
Metal already had logic like this, now the logic is centralized and can
be used for all GPU backends.
The parameters have been somewhat tuned per device, based on earlier
work for oneAPI in #163437 and CUDA in #163532. For Metal the behavior
should remain basically the same.
For oneAPI, this enables free_memory queries on iGPUs, as driver have
been exposing this for some time.
Co-authored-by: Patrick Mours <pmours@nvidia.com>
Co-authored-by; Xavier Hallade <xavier.hallade@intel.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
Each thread in the launched block runs over a portion of the counter
array. First to calculate a sum of each portion, then summing those across
all threads to get the base offset of each portion, before each thread
does the prefix sum of its local portion again including that base offset.
This could be improved further, but it already solves negative performance
impact when increasing the number of states in the following commit.
Pull Request: https://projects.blender.org/blender/blender/pulls/163930
Future versions of Intel Open Image Denoise (OIDN) will expect the
denoising depth value to lie between a near and far value. Supplying
values outside of this range can lead to visual issues.
So far, Cycles let a depth of 0 correspond to the near clip plane and
chose the maximum depth value `FLT_MAX`.
The following changes are introduced to make the denoising depth more
robust:
- The near clip plane is added to the depth for bounce 0. This makes
sure that the depth value will never be lower than the near clip
plane. For orthographic cameras this is skipped, as
`nearclip = -farclip / 2` is used and OIDN does not accept negative
near clip planes.
- When writing the background denoising features, `FLT_MAX` is replaced
by the far clip plane depth. For orthogonal cameras, `cliplength` is
used due to the reasons outlined above.
Pull Request: https://projects.blender.org/blender/blender/pulls/163818
This only affects Cycles standalone, as Blender does not use the auto
colorspace. The matrix multiplication order was inverted and decoding
from sRGB was not implemented properly.
Issue introduced in ca4458f6ea.
Pull Request: https://projects.blender.org/blender/blender/pulls/164212
When tessellation is repeated on a mesh in Hydra, preserve the coarse
vertex count used by the subdivision path. Blender integration handles
this by re-creating the mesh, but this should be handled in the core.
Pull Request: https://projects.blender.org/blender/blender/pulls/163988
This mainly affected precision and performance, as far as I can tell
there was no significant impact on e.g. the texture cache or OSL.
Now Cycles uses the same matrix as Blender (transposed), which was
created from dumping the OpenColorIO matrix so it matches exactly.
Thanks to Sergey for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/164197
Add core mechanism for switching the active OpenColorIO configuration at
runtime. This will make it possible to add a per project config.
The main difficulty is that ImBuf stores raw ColorSpace pointers
into the active configuration's color space objects. Scene datablocks
store color space names and can be re-resolved, but some data structures
(icons, previews, thumbnails, brush textures) outlive blend files.
The solution is to switch the configuration in-place, and retain color
spaces from previous configs but hide them in the UI.
Ref #163048
Pull Request: https://projects.blender.org/blender/blender/pulls/162196
Some values don't need full float precision, and with millions of states
reducing memory usage is important, while the math to pack/unpack these
is quite cheap. On CPU full precision is used since there are few states
and the extra conversion cost only hurts.
This gives a 2% reduction in state size with just
KERNEL_FEATURE_PATH_TRACING, and 9% reduction when enabling more
features like LIGHT_PASSES + DENOISING or SUBSURFACE + VOLUME.
Benchmarks do not show any significant impact on GPU render time either
way.
Pull Request: https://projects.blender.org/blender/blender/pulls/161876
The new raycast visibility is not like the other visibility flags, which
cover all possible light transport rays and so PATH_RAY_VISIBILITY_ALL
was a complete mask of those. But raycast rays are separate from that,
and should not affect e.g. volume stack interseciton.
Now define a separate PATH_RAY_VISIBILITY_OBJECT_ALL that indicates all
the possible flags that can be set on an object.
Issue introduced in 6e7a0eee32.
Pull Request: https://projects.blender.org/blender/blender/pulls/164137
This commit fixes the behavior when the shader graph accesses "radiance"
attribute in the case when the geometry has such attribute. The desired
behavior is to output the stored attribute instead of doing the runtime
radiance+base + spherical harmonics evaluation.
This is how OSL backend behaved prior to this change. So in a way it
fixes discrepancy between SVM and OSL.
Pull Request: https://projects.blender.org/blender/blender/pulls/164039
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.
This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.
The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from
build_files/config/pipeline_config.yaml
to
build_files/buildbot/config/blender_version.cmake
Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.
Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:
- There is an easy to follow table of per-architecture or family
toolkits.
- If there is no architectural preference, all provided toolkits are
probed. For SM kernels newer toolkits are tested first, and for
COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
default one is used.
A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).
Pull Request: https://projects.blender.org/blender/blender/pulls/164002
This commit makes it so Cycles renders gsplats the same way as EEVEE when some of
the attributes are missing.
This is achieved by creating missing attributes with some default values that are
obtained empirically. It is an opt-in behavior by the application integration code.
Pull Request: https://projects.blender.org/blender/blender/pulls/164042
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
This appears to have been caused by invalid use of assert() instead of
kernel_assert() in the kernel. I don't think it's actually hitting that
assert, but maybe it generated an unsupported instruction or something
along those lines?
The one that caused the actual problem is in svm/convert.h. I changed
more instances that are currently CPU only but risk becoming enabled on
the CPU with future code changes.
Thanks to Sahar A. Kashi for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/163970
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.
This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.
Pull Request: https://projects.blender.org/blender/blender/pulls/162785
Blender unconditionally requires the Vulkan multiViewport feature and
will fail to create a context on drivers that do not report it.
The feature is now optional. A new GPU_multi_viewport_support()
capability reports whether it can be used. Vulkan requires the
multiViewport feature and at least GPU_MAX_VIEWPORTS viewports;
Metal and OpenGL always report support. Apparently all devices
only support 1 or 16 viewports.
EEVEE uses the capability for shadow rendering. When it is unavailable,
the shadow module falls back to rendering each shadow view separately.
The fallback reads back the number of used views and each size,
restricts the shadow visibility shader to one view at a time and renders
at the origin of the shadow framebuffer.
When the fallback is enabled the performance would be slow, but at least
we can draw the correct shadows on these systems.
Ref: #162564
Pull Request: https://projects.blender.org/blender/blender/pulls/163697
DLSS queues up internal resources for delayed destruction by default
(with a delay of a couple of seconds). That's not necessary in Cycles,
because the denoiser queue is synchronized before release already, so
configure DLSS to free resources immediately.
Pull Request: https://projects.blender.org/blender/blender/pulls/163923
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
Collection has holdout and indirect only restriction toggle in outliner.
These properties exist for object id too. Add them in outliner next to
individual object tree element.
See PR description for images and RCS thread
Pull Request: https://projects.blender.org/blender/blender/pulls/157334
---------
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Commit 976ca9d244 fixed the "Hardware
Ray-Tracing" state message being printed to the log always showing off
for OptiX. This improves that further by querying whether the device
actually supports it, since OptiX has a software fallback on devices
without hardware ray tracing support.
Pull Request: https://projects.blender.org/blender/blender/pulls/163807
When using sample scrambling, independent samples are usually different
enough to trigger DLSS to disregard its temporal history.
This ultimately leads to flickering and artifacts in most situations.
So this commit disables sample scrambling when using DLSS.
Pull Request: https://projects.blender.org/blender/blender/pulls/163765
This adds the Integer Math node to shader nodes. It's the same node that
can already be used in Geometry Nodes and the Compositor.
While creating the regression test, I noticed that the integer math
implementation was a bit broken:
* The integer modulo (`%`) operator on the GPU is undefined for negative
divisors (the added regression test showed that the behavior is not
consistent across platforms).
* Use integer power function instead of using float implementation with
casts.
The MaterialX implementation is omitted for now. I didn't look into it
in detail yet but for the Boolean Math node (#162798) it seemed like
some more work is needed there to support this type.
Pull Request: https://projects.blender.org/blender/blender/pulls/163805
Previously, normal links and internal links both used `bNodeLink`.
However, they are quite different. Internal links are fully runtime
(only computed for muted nodes) and don't need the flags and other stuff
stored on normal links. When `bNodeLink` is extended with e.g. runtime
data, internal link storage shouldn't need to change.
Pull Request: https://projects.blender.org/blender/blender/pulls/163849
This adds the Boolean Math node to shader nodes. It's the same node that
can already be used in Geometry Nodes and the Compositor.
This also fixes the conversion from other types to boolean values so that
it it consistent across different node tree types. Cycles doesn't have
native support for booleans currently (it encodes them in integers), so
for now this conversion is entirely handled in the inliner.
The MaterialX implementation is omitted for now because it seems to need
more work to properly integrate booleans (right now it appears to ignore
all links going from a float to a boolean socket for example).
Pull Request: https://projects.blender.org/blender/blender/pulls/162798
It is possible for users to select the DLSS denoiser, but for it to
not work (E.g. old GPU drivers or missing DLSS files).
In this situation, the viewport sample count is still respected, so
this commit does not disable the viewport sample count UI element if
DLSS is selected but non-functional.
Pull Request: https://projects.blender.org/blender/blender/pulls/163766
This change works around the crash that happens on W7600 when rendering
principled_bsdf_bevel_emission_137420.blend test file.
Somehow the code starts to believe mnee_vertex_count is not 0 in the
middle of the function. Unfortunately, the only workaround seems to be
to un-inline the function.
Pull Request: https://projects.blender.org/blender/blender/pulls/163812