This patch improves the multi-scatter approximations for isotropic GGX
models to pass the furnace tests. The resolution of the look up tables
is increased as well as the mu (cosine(theta)) parametrization is
optimized to increase resolution at grazing angles.
Pull Request: https://projects.blender.org/blender/blender/pulls/159635
According to OpenPBR spec v1.1.1 Eq. (9), the correct layering for a
coated substrate is
$$f_\mathrm{layer} = f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub}.$$
For fractional coat weight, we apply Eq. (14), and it becomes
$$
\begin{aligned}
f_\mathrm{weighted-layer}&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}f_\mathrm{layer}\\
&=(1-w_\mathrm{coat})f_\mathrm{sub}+w_\mathrm{coat}(f_\mathrm{coat} + T_\mathrm{coat}(1 - E_\mathrm{coat})f_\mathrm{sub})\\
&=w_\mathrm{coat}f_\mathrm{coat}+(1-w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat})))f_\mathrm{sub}
\end{aligned}
$$
Here, \(E_\mathrm{coat}\) does not contain the coat weight nor the
overall weight of the coated closure \(w_\mathrm{layer}\), but our
current way of using `bsdf_albedo()` contains both weight, after which
we apply \(T_\mathrm{coat}\), which is wrong.
To fix this issue, we allow RGB weight for the layer operation, and
directly return a albedo of
$$
A_\mathrm{coat} = w_\mathrm{layer}w_\mathrm{coat}(1-T_\mathrm{coat}(1-E_\mathrm{coat}))
$$
Thus, when we apply the layer operation, we get
$$
\begin{aligned}
&w_\mathrm{layer}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat},w_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}\mathrm{layer}(S_\mathrm{sub},S_\mathrm{coat})\\
=&w_\mathrm{layer}(1-w_\mathrm{coat})S_\mathrm{sub}+w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat} + (w_\mathrm{layer}w_\mathrm{coat}-A_\mathrm{coat})S_\mathrm{sub}\\
=&w_\mathrm{layer}w_\mathrm{coat}S_\mathrm{coat}+(w_\mathrm{layer}-A_\mathrm{coat})S_\mathrm{sub},
\end{aligned}
$$
Which matches the equation for \(f_\mathrm{weighted-layer}\).
A modification to the sheen BSDF was needed to make it not tint the sub
layers.
Pull Request: https://projects.blender.org/blender/blender/pulls/162471
Also use the MaterialX-compatible `sheen_bsdf()` internally for the old
`sheen()` OSL closure, as they are essentially the same.
Should remove `sheen()` in 6.0
Dispersion is defined by two parameters:
Abbe Number and Dispersion Scale.
This corresponds to OpenPBR v1.1.1
Pull Request: https://projects.blender.org/blender/blender/pulls/162041
Note: dispersion is temporarily disabled when using MNEE (shadow caustics) on oneAPI, due to a compiler bug. Waiting for a fix from the oneAPI side.

Co-authored-by: Sebastian Herholz <sebastian.herholz@gmail.com>
The ray differential widening for the texture cache blurred out caustics
both from procedural textures and image textures. The widening helps
reduce texture cache memory, but is also fundamentally biased.
The solution now is to tie this to the Filter Glossy setting that must
already be lowered from the default to make caustics work. When left at
the default the texture cache will be most efficient, when set to small
value or zero there will be little or no widening.
So this setting now becomes a more general bias control for shading
after a rough glossy or diffuse bounce, and the tooltip was updated.
Pull Request: https://projects.blender.org/blender/blender/pulls/162153
We are running out of bits. Split the original `ShaderDataFlag` into
`ShaderRuntimeFlag`, which is determined by closures in the shader and
set up during rendering, and `ShaderDataFlag`, which is the same for the
whole shader graph and hence determined during shader compilation
The size of `ShaderData` did not change due to padding
Pull Request: https://projects.blender.org/blender/blender/pulls/161856
The problem was that next event estimation and forward sampling did not
have the same ray differential, which caused differnet mip level
selection. For multiple importance both of these need to converge to the
exact same result. The visible seam is especially an issue when there
are some discontinuous decision in sampling like the tree.
For forward sampling, previously it used the roughness of the sampled
BSDF. Instead we now use a weighted average of roughness that depends on
the direction and can be computed for both cases, as well as path guiding.
It uses the MIS weights for the weighted average, to ensure a mix of
sharp specular BSDFs and diffuse BDFs still gives a fixed differential
per direction that is suitable for both. MIS weights for sharp BSDFs
will be high in the reflection directions, while in other directions
the diffuse BSDF roughness will dominate.
Computing this weighted average roughness has some overhead. It may be
beneficial for the texture cache, as each direction will now only need
a single mip level instead of multiple.
Pull Request: https://projects.blender.org/blender/blender/pulls/159824
Most of Cycles assumes that the ray direction is normalized. The ray
portal BSDF allowed users to overwrite the ray direction with a
non-normalized vector, which leads to rendering issues later on.
This commit fixes the issue by normalizing the portal BSDF ray
direction.
Pull Request: https://projects.blender.org/blender/blender/pulls/160472
Always do the table lookup for the albedo in case it's needed for the
glossy or transmissions pass even if not needed for the combined pass.
This also improves sampling quality for some rough glass, due to better
sampling weights.
Pull Request: https://projects.blender.org/blender/blender/pulls/158819
Removing the legacy scaling by 1/(4pi) of the radius scale for the randomwalk
subsurface scattering mode.
Adjusting the default subsurface radius from 0.05 to 0.005 to keep similar
default behavior.
Updating test cases for Cycles and EEVEE.
- Thin glass:
An infinitesimally thin sheet of dielectric, approximated by a reflected
lobe and a trasmitted lobe, the respective weights of both lobes are
analytically computed by summing up infinite geometric series that
account for internal reflections.
- Thin subsurface:
An infinitesimally thin sheet of dense scattering material, approximated
by a diffuse lobe and a translucent lobe, the respective weights of both
lobes are given by subsurface anisotropy, with specifies the relative
amount of backward and forward scattering.
Co-authored-by: Jesse Yurkovich <jesse.y@gmail.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/157469
Evaluate fresnel per channel, this means fewer variables have to be live
and less spilling is needed. This can be a bit slower to evaluate but
seems overall better for GPU kernels.
When this code was added it caused glowing SSS in #149904, which is
popping up again in #157469. This may or may not help, but regardless
it may be be a good defensive change.
Pull Request: https://projects.blender.org/blender/blender/pulls/158690
Either use no spaces (common for disabled arguments) `/*arg*/`
or spaces `/* Regular comment. */`
Cleanup lop-sided comments such as `/* X*/` or `/*X */`.
Also use doxy-sections for bmesh_structure.hh,
the ad-hoc section comments had become outdated.
Ref !158467
Following OpenPBR spec, subsurface anisotropy now goes from -1 to 1.
-1 is strong backward scatter, 1 is strong forward scatter, 0 is
isotropic.
For compatibility, previous Random Walk model is preserved and renamed
to Random Walk (Legacy).
Both Random Walk (Skin) and Random Walk (Legacy) preserve their looks.
For the new negative range and the new Random Walk model, Van de Hulst
mapping is used.
Ref: #145127
Pull Request: https://projects.blender.org/blender/blender/pulls/156355
This was previously only used by the Principled Hair BSDF, We'll need this
for image texture sampling, simplest is to make it unconditional as many
materials will have images.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
A series of commits which reduces the number of sign conversions
(int <-> uint) in the Cycles kernel.
While it is not expected that the conversion emits any instructions,
it is quite confusing to follow the code and choose proper type.
Additionally, from some development in !151540 it seemed that such
mismatch was responsible for the performance drop in HIP-RT.
Pull Request: https://projects.blender.org/blender/blender/pulls/152009
The addition of metallic thin film is causing some kind of problem for the
compiler, reorganize code to reduce register pressure.
Fix#150282: Performance regression with shader raytrace
Pull Request: https://projects.blender.org/blender/blender/pulls/150537
Shadow terminator tricks might access uninitialized value of `wo`,
which happens in cases BSDF evaluation returns LABEL_NONE (i.e. when
the incident ray is in the lower hemisphere).
It is likely harmless from the render result perspective, as the `eval`
does not seem to be used. Also, since the `eval` is initialized to 0
it would need to be aa case when undefined `wo` leads to NAN or INF
returned by the shadow terminator offsets. Still good to solve this issue
to allow use of debugging tools like Valgrind.
Pull Request: https://projects.blender.org/blender/blender/pulls/149932
Applies thin film iridescence to metals in Metallic BSDF and Principled BSDF.
To get the complex IOR values for each spectral band from F82 Tint colors,
the code uses the parametrization from "Artist Friendly Metallic Fresnel",
where the g parameter is set to F82. This IOR is used to find the phase shift,
but reflectance is still calculated with the F82 Tint formula after adjusting
F0 for the film's IOR.
Co-authored-by: Lukas Stockner <lukas@lukasstockner.de>
Co-authored-by: Weizhen Huang <weizhen@blender.org>
Co-authored-by: RobertMoerland <rmoerlandrj@gmail.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/141131
The distance sampling is mostly based on weighted delta tracking from
[Monte Carlo Methods for Volumetric Light Transport Simulation]
(http://iliyan.com/publications/VolumeSTAR/VolumeSTAR_EG2018.pdf).
The recursive Monte Carlo estimation of the Radiative Transfer Equation is
\[\langle L \rangle=\frac{\bar T(x\rightarrow y)}{\bar p(x\rightarrow
y)}(L_e+\sigma_s L_s + \sigma_n L).\]
where \(\bar T(x\rightarrow y) = e^{-\bar\sigma\Vert x-y\Vert}\) is the
majorant transmittance between points \(x\) and \(y\), \(p(x\rightarrow
y) = \bar\sigma e^{-\bar\sigma\Vert x-y\Vert}\) is the probability of
sampling point \(y\) from point \(x\) following exponential
distribution.
At each recursive step, we randomly pick one of the two events
proportional to their weights:
* If \(\xi < \frac{\sigma_s}{\sigma_s+\vert\sigma_n\vert}\), we sample
scatter event and evaluate \(L_s\).
* Otherwise, no real collision happens and we continue the recursive
process.
The emission \(L_e\) is evaluated at each step.
This also removes some unused volume settings from the UI:
* "Max Steps" is removed, because the step size is automatically specified
by the volume octree. There is a hard-coded threshold `VOLUME_MAX_STEPS`
to prevent numerical issues.
* "Homogeneous" is automatically detected during density evaluation
An option "Unbiased" is added to the UI. When enabled, densities above
the majorant are clamped.
The Principled BSDF has a ton of inputs, and the previous SVM code just always
allocated stack space for all of them. This results in a ton of additional
NODE_VALUE_x SVM nodes, which slow down execution.
However, this is not really needed for two reasons:
- First, many inputs are only used consitionally. For example, if the
subsurface weight is zero, none of the other subsurface inputs are used.
- Many of the inputs have a "usual" value that they will have in most
materials, so if they happen to have that value we can just indicate that
by not allocating space for them.
This is a bit similar to the standard "pack the fixed value and provide
a stack offset if there's a link" pattern, except that the fixed value
is a constant in the code and we allocate a NODE_VALUE_x if a different
fixed value is used.
Therefore, this PR re-implements the parameter packing in a more efficient way:
- If we can determine that a component is disabled, all conditional inputs are
disconnected (to avoid generating upstream nodes).
- If we can determine that a component is disabled, we skip allocating all
conditional inputs on the stack.
- The inputs for which a reasonable "usual" value exists are changed to
respect that, and to only be allocated if they differ.
- param1 and param2 (which are fixed-value-packed as on all BSDF nodes) are
used to store IOR and roughness, which have a decent chance to be fixed
values.
- The parameter packing is more aggressive about using uchar4, which allows
to get rid of two SVM nodes while still storing the same inputs.
The result is a considerable speedup in scenes that make heavy use of the
Principled BSDF:
| Scene | CPU speedup | OptiX speedup |
| --- | --- | --- |
| attic | 5% | 9% |
| bistro | 5% | 8% |
| junkshop | 5% | 10% |
| monster | 3% | 4% |
| spring | 1% | 6% |
Pull Request: https://projects.blender.org/blender/blender/pulls/143910
This PR adds a new `fresnel_conductor_polarized` function, which calculates reflectance and phase shift (if requested) for both parallel and perpendicular polarized light. This is needed for applying thin film iridescence to conductors (see !141131).
For consistency, this PR also makes `fresnel_conductor` call `fresnel_conductor_polarized` instead of using a fast approximation of the Fresnel equations that is inaccurate at lower n and k values. This will change the output of some Metallic BSDF renders using Physical Conductor and prevent discrepancies when enabling thin film iridescence.
I didn't do any rigorous performance testing, but from timing the functions outside of Blender, `fresnel_conductor_polarized` is significantly slower than the approximation, between 1.5-3x depending on the compiler. This makes sense because it has three square roots and the approximation has none. In some informal tests with metallic_multiggx_physical.blend modified to have more spheres, the new renders took around 1-2% longer on both CPU and GPU.
There are some avoidable inefficiencies in this approach of just calling `fresnel_conductor_polarized`:
- one of the three square roots could be saved since `fresnel_conductor` never needs the phase shift and there are simplifications possible when only calculating the reflectance
- there are several unnecessary multiplications by 1.0 since `fresnel_conductor` uses relative IOR and `fresnel_conductor_polarized` doesn't, though those could get optimized out if inlined
Pull Request: https://projects.blender.org/blender/blender/pulls/143903
Add new "Linear 3D Curves" option in the Curves panel in the render
properties. This renders curves as linear segments rather than smooth
curves, for faster render time at the cost of accuracy.
On NVIDIA Blackwell GPUs, this can give a 6x speedup compared to smooth
curves, due to hardware acceleration. On NVIDIA Ada there is still
a 3x speedup, and CPU and other GPU backends will also render this
faster.
A difference with smooth curves is that these have end caps, as this
was simpler to implement and they are usually helpful anyway.
In the future this functionality will also be used to properly support
the CURVE_TYPE_POLY on the new curves object.
Pull Request: https://projects.blender.org/blender/blender/pulls/139735
Previously, we used precomputed Gaussian fits to the XYZ CMFs, performed
the spectral integration in that space, and then converted the result
to the RGB working space.
That worked because we're only supporting dielectric base layers for
the thin film code, so the inputs to the spectral integration
(reflectivity and phase) are both constant w.r.t. wavelength.
However, this will no longer work for conductive base layers.
We could handle reflectivity by converting to XYZ, but that won't work
for phase since its effect on the output is nonlinear.
Therefore, it's time to do this properly by performing the spectral
integration directly in the RGB primaries. To do this, we need to:
- Compute the RGB CMFs from the XYZ CMFs and XYZ-to-RGB matrix
- Resample the RGB CMFs to be parametrized by frequency instead of wavelength
- Compute the FFT of the CMFs
- Store it as a LUT to be used by the kernel code
However, there's two optimizations we can make:
- Both the resampling and the FFT are linear operations, as is the
XYZ-to-RGB conversion. Therefore, we can resample and Fourier-transform
the XYZ CMFs once, store the result in a precomputed table, and then just
multiply the entries by the XYZ-to-RGB matrix at runtime.
- I've included the Python script used to compute the table under
`intern/cycles/doc/precompute`.
- The reference implementation by the paper authors [1] simply stores the
real and imaginary parts in the LUT, and then computes
`cos(shift)*real + sin(shift)*imag`. However, the real and imaginary parts
are oscillating, so the LUT with linear interpolation is not particularly
good at representing them. Instead, we can convert the table to
Magnitude/Phase representation, which is much smoother, and do
`mag * cos(phase - shift)` in the kernel.
- Phase needs to be unwrapped to handle the interpolation decently,
but that's easy.
- This requires an extra trig operation in the kernel in the dielectric case,
but for the conductive case we'll actually save three.
Rendered output is mostly the same, just slightly different because we're
no longer using the Gaussian approximation.
[1] "A Practical Extension to Microfacet Theory for the Modeling of
Varying Iridescence" by Laurent Belcour and Pascal Barla,
https://belcour.github.io/blog/research/publication/2017/05/01/brdf-thin-film.html
Pull Request: https://projects.blender.org/blender/blender/pulls/140944