Some values don't need full float precision, and with millions of states
reducing memory usage is important, while the math to pack/unpack these
is quite cheap. On CPU full precision is used since there are few states
and the extra conversion cost only hurts.
This gives a 2% reduction in state size with just
KERNEL_FEATURE_PATH_TRACING, and 9% reduction when enabling more
features like LIGHT_PASSES + DENOISING or SUBSURFACE + VOLUME.
Benchmarks do not show any significant impact on GPU render time either
way.
Pull Request: https://projects.blender.org/blender/blender/pulls/161876
This appears to have been caused by invalid use of assert() instead of
kernel_assert() in the kernel. I don't think it's actually hitting that
assert, but maybe it generated an unsupported instruction or something
along those lines?
The one that caused the actual problem is in svm/convert.h. I changed
more instances that are currently CPU only but risk becoming enabled on
the CPU with future code changes.
Thanks to Sahar A. Kashi for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/163970
Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
This change works around the crash that happens on W7600 when rendering
principled_bsdf_bevel_emission_137420.blend test file.
Somehow the code starts to believe mnee_vertex_count is not 0 in the
middle of the function. Unfortunately, the only workaround seems to be
to un-inline the function.
Pull Request: https://projects.blender.org/blender/blender/pulls/163812
The need of this arose from the gsplat branch where extra attribute
offsets need to be cached for fast lookup. The naive way of adding
them would increase the KernelObject by 4 integers, which is not ideal.
Since some fields in the KernelObject are not used by point cloud
objects it is possible to use union to save the save.
This change introduces the union in the KernelObejct and moves some
code around so that fields in there are only written and accessed
from meshes and volumes.
Since this change could potentially cause regressions that wouldn't
be easy to track it is extracted from the gsplat branch.
Pull Request: https://projects.blender.org/blender/blender/pulls/163441
According to OpenPBR spec v1.1.1 Eq. (9), the correct layering for a
coated substrate is
$$f_\mathrm{layer} = f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub}.$$
For fractional coat weight, we apply Eq. (14), and it becomes
$$
\begin{aligned}
f_\mathrm{weighted-layer}&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}f_\mathrm{layer}\\
&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}(f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub})\\
&=w_\mathrm{coat}f_\mathrm{coat}+(1-w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat})))f_\mathrm{sub}
\end{aligned}
$$
Here, \(E_\mathrm{coat}\) does not contain the coat weight nor the
overall weight of the coated closure \(w_\mathrm{layer}\), but our
current way of using `bsdf_albedo()` contains both weight, after which
we apply \(T_\mathrm{coat}\), which is wrong.
To fix this issue, we allow RGB weight for the layer operation, and
directly return a albedo of
$$
A_\mathrm{coat} = w_\mathrm{layer}w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat}))
$$
Thus, when we apply the layer operation, we get
$$
\begin{aligned}
&w_\mathrm{layer}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat},w_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat} + (w_\mathrm{layer}w_\mathrm{coat}-A_\mathrm{coat})S_\mathrm{sub}\\
=&w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat}+(w_\mathrm{layer}-A_\mathrm{coat})S_\mathrm{sub},
\end{aligned}
$$
Which matches the equation for \(f_\mathrm{weighted-layer}\).
A modification to the sheen BSDF was needed to make it not tint the sub
layers.
Pull Request: https://projects.blender.org/blender/blender/pulls/162471
Mixing multiple volume phase closures can make world background
lighting slightly too bright. The phase-sampled path stores the PDF of
only the selected closure, while the background-sampled path evaluates
the marginal mixture PDF. The resulting MIS weights are inconsistent.
Store the marginal PDF over all phase closures for forward MIS. When
volume path guiding is active, combine it with the guiding PDF so the
stored PDF remains equal to the actual sampling distribution.
Pull Request: https://projects.blender.org/blender/blender/pulls/161690
Dispersion is defined by two parameters:
Abbe Number and Dispersion Scale.
This corresponds to OpenPBR v1.1.1
Pull Request: https://projects.blender.org/blender/blender/pulls/162041
Note: dispersion is temporarily disabled when using MNEE (shadow caustics) on oneAPI, due to a compiler bug. Waiting for a fix from the oneAPI side.

Co-authored-by: Sebastian Herholz <sebastian.herholz@gmail.com>
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
The ray differential widening for the texture cache blurred out caustics
both from procedural textures and image textures. The widening helps
reduce texture cache memory, but is also fundamentally biased.
The solution now is to tie this to the Filter Glossy setting that must
already be lowered from the default to make caustics work. When left at
the default the texture cache will be most efficient, when set to small
value or zero there will be little or no widening.
So this setting now becomes a more general bias control for shading
after a rough glossy or diffuse bounce, and the tooltip was updated.
Pull Request: https://projects.blender.org/blender/blender/pulls/162153
When using a Principled BSDF node with a fully black diffuse color on
input, Cycles would raise an assert error in
surface_shader_bsdf_sample_closure() when baking a diffuse pass.
This fix removes the SD_BSDF flag from the shader in
surface_shader_prepare_closures(), if no BSDF closure is left.
Pull Request: https://projects.blender.org/blender/blender/pulls/162059
We are running out of bits. Split the original `ShaderDataFlag` into
`ShaderRuntimeFlag`, which is determined by closures in the shader and
set up during rendering, and `ShaderDataFlag`, which is the same for the
whole shader graph and hence determined during shader compilation
The size of `ShaderData` did not change due to padding
Pull Request: https://projects.blender.org/blender/blender/pulls/161856
The problem was that next event estimation and forward sampling did not
have the same ray differential, which caused differnet mip level
selection. For multiple importance both of these need to converge to the
exact same result. The visible seam is especially an issue when there
are some discontinuous decision in sampling like the tree.
For forward sampling, previously it used the roughness of the sampled
BSDF. Instead we now use a weighted average of roughness that depends on
the direction and can be computed for both cases, as well as path guiding.
It uses the MIS weights for the weighted average, to ensure a mix of
sharp specular BSDFs and diffuse BDFs still gives a fixed differential
per direction that is suitable for both. MIS weights for sharp BSDFs
will be high in the reflection directions, while in other directions
the diffuse BSDF roughness will dominate.
Computing this weighted average roughness has some overhead. It may be
beneficial for the texture cache, as each direction will now only need
a single mip level instead of multiple.
Pull Request: https://projects.blender.org/blender/blender/pulls/159824
When the texture cache evaluates the background shader at a low mipmap level
for next event estimation, it can hit a bright pixel from e.g. very bright
blurred. But in the importance map generated at a higher resolution, the
corresponding pixels may not be bright at all and get a low importance
sampling probability. This leads to significant noise.
To avoid this, clamp the ray differentials so they are at least the ray
differentials used for generating the importance map. This both reduces
noise and mipmap bias. It comes at some extra memory cost, but the
importance map resolution is limited by default and overall it's worth it.
Pull Request: https://projects.blender.org/blender/blender/pulls/159824
This commit fixes a bug when using light linking in combination with
path guiding. Before it happened that linked infinite light source
generated training samples at positions outside the numerical stable
range. This PR avoids this behavior by treating the contribution of
linked lights as scattered contribution at the current path segment and
generates an additional directional training sample towards the
direction of the linked light source.
In addition this fix clamps the positions fed to the guiding structures
into a numerical safe range.
This PR fixes#158755
Pull Request: https://projects.blender.org/blender/blender/pulls/159804
This helps especially for OSL, where load time is very long. By
splitting off shader raytrace, MNEE and volumes, the first time
rendering is much faster.
The downside is that noinline functions will be duplicated. However OSL
startup performance is very bad currently and this seems the better
trade-off for now. There are ways to make this work if we do not mark
noinline functions as static, but this will require some bigger code
reorganization.
Without OSL, shader raytrace and MNEE were combined in a single module,
and that has been split into two as well.
Besides a better user experience, This will fix OptiX OSL test timeouts
on the buildbot, where some tests need to compute a few different
specializations and the 600s timeout is exceeded, with the longest test
run time being around 300s.
Pull Request: https://projects.blender.org/blender/blender/pulls/159499
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.
MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.
MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.
This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.
Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
The issue this change aims to solve is the fact that we are out of bits
in the PathRayFlag. There is a TODO in the code about the fact that we
might need a flag to record primary scatter for volumes. Additionally,
there are plans to introduce visibility flags for the raycast node
rays.
The choice of the underlying type for the PathRayVisibilityFlag keeps
the raycast visibility in mind.
There is now a dedicated type alias to pass visiiblity around in the
kernel. Currently it is 32 bit, keeping in mind possibility of shift
introduced by the shadow catcher. It is possible to limit it to 16,
but then we'd be limited to only 8 visibility bits (which would be
fine until we introduce other flag after the raycast visibility).
Although, it is not really clear it will brings an actual performance
improvement.
The main path state is using 16 to store path visibility, as per the
underlying type of the PathRayVisibilityFlag. It is possible to limit
it to 8 bit to keep state size small until even more ray visibility
bits be needed in the future (the unaligned node bit is not used by the
path visibility).
The shadow path state also needs path visibility to check which path
type it deals with.
Ref !157822
Currently matches the `uint` that is used everywhere. Intended to
be used everywhere in the kernel and in the host where visibility
is passed around.
It might become a narrower type in the future if visibility fits
into 8 bits (it could become uint16_t then, to have space for the
shadow catcher shifts).
Should be no functional changes,
Ref !157822
- Use spaces between commas when multiple args were passed in.
- Use colon after arg(s) `\param arg: ...` or `\param a, b: ...`.
- Use \param not \params (which isn't a recognized command).
Ref !158691
Either use no spaces (common for disabled arguments) `/*arg*/`
or spaces `/* Regular comment. */`
Cleanup lop-sided comments such as `/* X*/` or `/*X */`.
Also use doxy-sections for bmesh_structure.hh,
the ad-hoc section comments had become outdated.
Ref !158467
Caused by 4d6c7718c5.
Under certain conditions it is possible that the offset logic in the
AttributeTableBuilder::add() assigns offset of -1. This value just
happens to match ATTR_STD_NOT_FOUND, leading to false-positive check
that attribute is not found (for example, svm_node_attr_init() that
checks desc.offset != ATTR_STD_NOT_FOUND.
The simplest solution is to tweak the value of the ATTR_STD_NOT_FOUND
to make it very big negative value.
There is now an utility function to check whether AttributeDescriptor
points to an attribute that is found, to help possibly tweaking this
check in the future, and to make code a bit more semantically clear.
Pull Request: https://projects.blender.org/blender/blender/pulls/157410
Following OpenPBR spec, subsurface anisotropy now goes from -1 to 1.
-1 is strong backward scatter, 1 is strong forward scatter, 0 is
isotropic.
For compatibility, previous Random Walk model is preserved and renamed
to Random Walk (Legacy).
Both Random Walk (Skin) and Random Walk (Legacy) preserve their looks.
For the new negative range and the new Random Walk model, Van de Hulst
mapping is used.
Ref: #145127
Pull Request: https://projects.blender.org/blender/blender/pulls/156355
Blender crashed during rendering when path guiding is enabled because
the number of recorded path segments during training exceeded the
reserved size of the PathSegmentStorage. This problem only occurs if
the unbiased volume rendering method is used. The problem is fixed by
adding checks if the we generated a valid new path segment before
adding data to it.
This is a hotfix to avoid crashes.
Further investigation has to be done: Why does the number of generated
path segments exceed the reserved size.
Pull Request: https://projects.blender.org/blender/blender/pulls/157185
A few functions were wrapped into an auto-address space naming, with
no actual need with the modern structure of the kernel.
It is a residual state from the original OpenCL split kernel where
the same function was used on a data from different memory depending
on a device or kernel.
Pull Request: https://projects.blender.org/blender/blender/pulls/157194
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913