The intel/llvm 7.1 upgrade is giving a large performance regression in
shade_surface kernel on Intel GPUs. I've root-caused it down to a slight
shift in private memory layout due to slightly different inlining.
To fix both this regression and overall performance sensitivity on
private memory layout, that can happen with any code change beyond the
DPC++ upgrade, I'm forcing ShaderData and svm's stack alignments to 64
bytes.
Pull Request: https://projects.blender.org/blender/blender/pulls/164411
This commit fixes the behavior when the shader graph accesses "radiance"
attribute in the case when the geometry has such attribute. The desired
behavior is to output the stored attribute instead of doing the runtime
radiance+base + spherical harmonics evaluation.
This is how OSL backend behaved prior to this change. So in a way it
fixes discrepancy between SVM and OSL.
Pull Request: https://projects.blender.org/blender/blender/pulls/164039
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
This appears to have been caused by invalid use of assert() instead of
kernel_assert() in the kernel. I don't think it's actually hitting that
assert, but maybe it generated an unsupported instruction or something
along those lines?
The one that caused the actual problem is in svm/convert.h. I changed
more instances that are currently CPU only but risk becoming enabled on
the CPU with future code changes.
Thanks to Sahar A. Kashi for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/163970
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.
This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.
Pull Request: https://projects.blender.org/blender/blender/pulls/162785
This adds the Integer Math node to shader nodes. It's the same node that
can already be used in Geometry Nodes and the Compositor.
While creating the regression test, I noticed that the integer math
implementation was a bit broken:
* The integer modulo (`%`) operator on the GPU is undefined for negative
divisors (the added regression test showed that the behavior is not
consistent across platforms).
* Use integer power function instead of using float implementation with
casts.
The MaterialX implementation is omitted for now. I didn't look into it
in detail yet but for the Boolean Math node (#162798) it seemed like
some more work is needed there to support this type.
Pull Request: https://projects.blender.org/blender/blender/pulls/163805
This adds the Boolean Math node to shader nodes. It's the same node that
can already be used in Geometry Nodes and the Compositor.
This also fixes the conversion from other types to boolean values so that
it it consistent across different node tree types. Cycles doesn't have
native support for booleans currently (it encodes them in integers), so
for now this conversion is entirely handled in the inliner.
The MaterialX implementation is omitted for now because it seems to need
more work to properly integrate booleans (right now it appears to ignore
all links going from a float to a boolean socket for example).
Pull Request: https://projects.blender.org/blender/blender/pulls/162798
The pack only contains bands 1 to 3. The evaluation function has
been renamed as part of the original review. It makes sense to be
consistent, although a bit more verbose.
Pull Request: https://projects.blender.org/blender/blender/pulls/163446
Includes possibility of adding spherical harmonics on geometry, as well
as includes spherical harmonics evaluation.
Spherical harmonics are stored quantized to 8 bit integer, and always
use degree 3. The reason for the limited degree is the memory: degree
3 requires relatively small amount of coefficients (45 int8_t, which
is similar to how realtime renderers pack them into float3x4). Going
higher dimension adds a lot more coefficients to store, and so far it
doesn't seem there are gsplat data-sets that are trained to a degree
higher than 3.
The attribute is currently unused. It is preparation work foe the
gsplat project to make reviews easier.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163409
The main goal is to bring a bit of structure to the use of quaternions
in the Cycles kernel.
Prior to this change there was no dedicated quaternion type and float4
was used instead, and quaternions are stored as (x, y, z, w) matching
naming between the quaternion and float4 fields. However, Blender uses
(w, x, y, z) quaternion order for attributes, which is different from
what Cycles uses. Without dedicated type this either leads to different
quaternion orders depending whether it comes from attribute, or makes
it intrinsically not possible to use implicit sharing.
This change introduces Cycles Quaternion type which is compatible with
Blender attribute math::Quaternion in both layout and alignment, making
it possible to benefit from implicit sharing and use a nice structure
in the kernel.
This change does not modify the existing quaternion usage in the kernel
which is currently used for transform decomposition and interpolation.
There is currently no SIMD for the Quaternion type, which allows to
avoid any special alignment requirement and share attributes with
Blender, but performance might not be ideal.
This attribute type will be used for representing gaussian splat
rotation.
The attribute access and interpolation matches behavior prior to this
change. Added some basic tests for quaternion attribute access for RGB
and alpha, mesh and volume attributes.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163342
According to OpenPBR spec v1.1.1 Eq. (9), the correct layering for a
coated substrate is
$$f_\mathrm{layer} = f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub}.$$
For fractional coat weight, we apply Eq. (14), and it becomes
$$
\begin{aligned}
f_\mathrm{weighted-layer}&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}f_\mathrm{layer}\\
&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}(f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub})\\
&=w_\mathrm{coat}f_\mathrm{coat}+(1-w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat})))f_\mathrm{sub}
\end{aligned}
$$
Here, \(E_\mathrm{coat}\) does not contain the coat weight nor the
overall weight of the coated closure \(w_\mathrm{layer}\), but our
current way of using `bsdf_albedo()` contains both weight, after which
we apply \(T_\mathrm{coat}\), which is wrong.
To fix this issue, we allow RGB weight for the layer operation, and
directly return a albedo of
$$
A_\mathrm{coat} = w_\mathrm{layer}w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat}))
$$
Thus, when we apply the layer operation, we get
$$
\begin{aligned}
&w_\mathrm{layer}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat},w_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat} + (w_\mathrm{layer}w_\mathrm{coat}-A_\mathrm{coat})S_\mathrm{sub}\\
=&w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat}+(w_\mathrm{layer}-A_\mathrm{coat})S_\mathrm{sub},
\end{aligned}
$$
Which matches the equation for \(f_\mathrm{weighted-layer}\).
A modification to the sheen BSDF was needed to make it not tint the sub
layers.
Pull Request: https://projects.blender.org/blender/blender/pulls/162471
Also use the MaterialX-compatible `sheen_bsdf()` internally for the old
`sheen()` OSL closure, as they are essentially the same.
Should remove `sheen()` in 6.0
Dispersion is defined by two parameters:
Abbe Number and Dispersion Scale.
This corresponds to OpenPBR v1.1.1
Pull Request: https://projects.blender.org/blender/blender/pulls/162041
Note: dispersion is temporarily disabled when using MNEE (shadow caustics) on oneAPI, due to a compiler bug. Waiting for a fix from the oneAPI side.

Co-authored-by: Sebastian Herholz <sebastian.herholz@gmail.com>
This change adds 32 more bit to store kernel features.
While for a short term it might be possible to make a space for one or
two extra bits, it seems going 64bit is inevitable.
Expanding the field to 64bit might introduce some slowdown due to less
optimal cache, but so is consolidation of existing flags could also
lead to performance drop in certain configurations.
The main tricky part of the change is Metal where function constants
are used to store kernel_features, and 64bit constants are only
available on macOS 12. There is a runtime check for it. On older macOS
versions the flags are stored as a pair of 32bit values. It is slower,
but there are unlikely to be many Cycles users on macOS 11.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162737
This adds a new simple node which retrieves the X, Y or Z component of a vector
by index. This is simpler and more efficient than using workarounds like using a
Separate XYZ with an Index Switch node.
The original motivation for this node came from #149091 where we wanted to
support this using expressions (`vector[i]`).
Co-authored-by: Tibo Stans <stanstibo@gmail.com>
We are running out of bits. Split the original `ShaderDataFlag` into
`ShaderRuntimeFlag`, which is determined by closures in the shader and
set up during rendering, and `ShaderDataFlag`, which is the same for the
whole shader graph and hence determined during shader compilation
The size of `ShaderData` did not change due to padding
Pull Request: https://projects.blender.org/blender/blender/pulls/161856
Regression from dcac328db3. svm_node_attr_init() set offset to 0 which
doesn't work with is_attribute_found() introduced in that commit, it
expects ATTR_STD_NOT_FOUND.
I traced all users of svm_node_attr_init() and found they all check
either (element & ATTR_ELEMENT_X) or is_attribute_found()` which both
work.
Pull Request: https://projects.blender.org/blender/blender/pulls/161012
Add missing derivative SVM node for built-in texture coordinate mapping
of texture nodes. Because this mapping is built into the node, not only
were the derivatives missing but SVM stack memory was read uninitialized.
An existing mipmap test was updated to include this type of mapping.
Pull Request: https://projects.blender.org/blender/blender/pulls/160518
Store positions as separate arrays again. This appears to give better
memory access patterns or layout. This was a regression from fc9917352b
and 0baa98866c.
This is a relatively simple change, rerouting the position attribute to
dedicated arrays, still including motion steps in the same array. The rest
of the refactor is preserved.
Pull Request: https://projects.blender.org/blender/blender/pulls/160110
The geo nodes implementations of the `mix`, `addition`, `multiplication`
and `subtraction` color mix modes are both faster and have better
numerical properties than the ones we use in Cycles, OSL and EEVEE.
This PR ports the geo nodes implementations of those color mix modes
to Cycles, OSL and EEVEE.
Additionally, the SVM mix function implementation is templatized and
a new function called "endvalue_preserving_mix" is added for the
alternative variant of linear interpolation.
Pull Request: https://projects.blender.org/blender/blender/pulls/158022
OpenPBR can also have emission, so the compiler hint for optimizing out
other BSDFs doesn't work anymore.
Instead, we extract the function to compute emission, and share it when
evaluating BSDF and surface emission
Motion is now part of the position attribute. On the kernel side, a
combined position + radius attribute is now stored, replacing the
previous motion only attribute.
Legacy motion attributes are now removed.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
Use packed_float3 instead of float3 for geometry positions and various
other arrays.
Add and use unified accessor methods get_position() and get_radius() for
all geometry, as well as num_verts(), num_points() and num_keys().
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
The issue this change aims to solve is the fact that we are out of bits
in the PathRayFlag. There is a TODO in the code about the fact that we
might need a flag to record primary scatter for volumes. Additionally,
there are plans to introduce visibility flags for the raycast node
rays.
The choice of the underlying type for the PathRayVisibilityFlag keeps
the raycast visibility in mind.
There is now a dedicated type alias to pass visiiblity around in the
kernel. Currently it is 32 bit, keeping in mind possibility of shift
introduced by the shadow catcher. It is possible to limit it to 16,
but then we'd be limited to only 8 visibility bits (which would be
fine until we introduce other flag after the raycast visibility).
Although, it is not really clear it will brings an actual performance
improvement.
The main path state is using 16 to store path visibility, as per the
underlying type of the PathRayVisibilityFlag. It is possible to limit
it to 8 bit to keep state size small until even more ray visibility
bits be needed in the future (the unaligned node bit is not used by the
path visibility).
The shadow path state also needs path visibility to check which path
type it deals with.
Ref !157822
Currently matches the `uint` that is used everywhere. Intended to
be used everywhere in the kernel and in the host where visibility
is passed around.
It might become a narrower type in the future if visibility fits
into 8 bits (it could become uint16_t then, to have space for the
shadow catcher shifts).
Should be no functional changes,
Ref !157822
- Thin glass:
An infinitesimally thin sheet of dielectric, approximated by a reflected
lobe and a trasmitted lobe, the respective weights of both lobes are
analytically computed by summing up infinite geometric series that
account for internal reflections.
- Thin subsurface:
An infinitesimally thin sheet of dense scattering material, approximated
by a diffuse lobe and a translucent lobe, the respective weights of both
lobes are given by subsurface anisotropy, with specifies the relative
amount of backward and forward scattering.
Co-authored-by: Jesse Yurkovich <jesse.y@gmail.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/157469