We are running out of bits. Split the original `ShaderDataFlag` into
`ShaderRuntimeFlag`, which is determined by closures in the shader and
set up during rendering, and `ShaderDataFlag`, which is the same for the
whole shader graph and hence determined during shader compilation
The size of `ShaderData` did not change due to padding
Pull Request: https://projects.blender.org/blender/blender/pulls/161856
When the texture cache evaluates the background shader at a low mipmap level
for next event estimation, it can hit a bright pixel from e.g. very bright
blurred. But in the importance map generated at a higher resolution, the
corresponding pixels may not be bright at all and get a low importance
sampling probability. This leads to significant noise.
To avoid this, clamp the ray differentials so they are at least the ray
differentials used for generating the importance map. This both reduces
noise and mipmap bias. It comes at some extra memory cost, but the
importance map resolution is limited by default and overall it's worth it.
Pull Request: https://projects.blender.org/blender/blender/pulls/159824
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.
Pull Request: https://projects.blender.org/blender/blender/pulls/158728
The issue this change aims to solve is the fact that we are out of bits
in the PathRayFlag. There is a TODO in the code about the fact that we
might need a flag to record primary scatter for volumes. Additionally,
there are plans to introduce visibility flags for the raycast node
rays.
The choice of the underlying type for the PathRayVisibilityFlag keeps
the raycast visibility in mind.
There is now a dedicated type alias to pass visiiblity around in the
kernel. Currently it is 32 bit, keeping in mind possibility of shift
introduced by the shadow catcher. It is possible to limit it to 16,
but then we'd be limited to only 8 visibility bits (which would be
fine until we introduce other flag after the raycast visibility).
Although, it is not really clear it will brings an actual performance
improvement.
The main path state is using 16 to store path visibility, as per the
underlying type of the PathRayVisibilityFlag. It is possible to limit
it to 8 bit to keep state size small until even more ray visibility
bits be needed in the future (the unaligned node bit is not used by the
path visibility).
The shadow path state also needs path visibility to check which path
type it deals with.
Ref !157822
Implementation is pretty much the same as for actual light objects,
and follows the same limitations w.r.t MIS.
The option is in the World -> Settings -> Surface. The world volume
does not seem to affect scene lighting, and having option under the
surface panel makes the UI look better.
Pull Request: https://projects.blender.org/blender/blender/pulls/156586
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
Now in Cycles, the term "Sun Light" is used to denote the Sun Light in
Blender UI.
The term "Distant Light" refers to both sun light and background light,
or is occasionally used in places where there is only sun light, but we
would like to add support to background light in the future.
Pull Request: https://projects.blender.org/blender/blender/pulls/155133
This is more consistent with Blender and fixes various discrepancies
between Cycles and Blender with sharp edges and mixed flat and smooth
faces. This uses more memory, but is compensated by earlier octahedral
mapping and in the benchmark scenes memory is always the same or lower.
Fix#152812: Mix of flat and smooth faces has incorrect sharp edges
Fix#152496: Smooth faces that share one vertex have bad normals
Fix#100425: Issues with motion blur and autosmooth or edge splitting
Pull Request: https://projects.blender.org/blender/blender/pulls/153836
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.
Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.
A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.
The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.
LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.
The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.
Pull Request: https://projects.blender.org/blender/blender/pulls/152649
A series of commits which reduces the number of sign conversions
(int <-> uint) in the Cycles kernel.
While it is not expected that the conversion emits any instructions,
it is quite confusing to follow the code and choose proper type.
Additionally, from some development in !151540 it seemed that such
mismatch was responsible for the performance drop in HIP-RT.
Pull Request: https://projects.blender.org/blender/blender/pulls/152009
When `delta` is significantly larger than the difference between `tmin`
and `tmax`, the precision of `theta_a` and `theta_b` reduces, resulting
in banding artefacts.
For equiangular sampling, we use uniform sampling to fix this case.
For light tree, we use the equality
atan(a) + atan(b) = atan2(a + b, 1 - a*b).
Pull Request: https://projects.blender.org/blender/blender/pulls/142845
Dropping the inlining hint for `light_tree_pdf` and reverting to the
default inlining thresholds for DPC++ compiler gives a ~4% speedup on
classroom and other scenes on Arc B580.
Pull Request: https://projects.blender.org/blender/blender/pulls/135042
This is an intermediate steps towards making lights actual geometry.
Light is now a subclass of Geometry, which simplifies some code.
The geometry is not added to the BVH yet, which would be the next
step and improve light intersection performance with many lights.
This makes object attributes work on lights.
Co-authored-by: Lukas Stockner <lukas@lukasstockner.de>
Pull Request: https://projects.blender.org/blender/blender/pulls/134846
In a previous commit (1), adjustments to light tree traversal were made
to try and skip distant lights when deciding which light to sample
while within a world volume.
The skip wasn't implemented properly, and as a result distant lights
were still included in the light tree traversal in that sitaution.
And due to the way the skip was implemented, there were some
unintialized variables used in the processing of the distant lights
importance which caused artifacts on some platforms.
This commit fixes this issue by reverting the skip of distant lights
in that situation.
(1) blender/blender@6fbc958e89
Pull Request: https://projects.blender.org/blender/blender/pulls/132344
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.
Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
When a scene contains distant lights and local lights, the first step
of the light tree traversal is to compute the importance of
distant lights vs local lights and pick one based on a random number.
In the specific case of when there is only one distant light,
the line of code that had been changed in this commit
effectively reduced to:
`min_importance = fast_cosf(x) < cosf(x) ? 0.0 : compute_min_importance`
And depending on the hardware, compiler, and the specific value being
tested, different configurations could take different code paths.
This commit fixes this issue by turning the comparison into
`fast_cosf(x) < fast_cosf(x)`.
---
Why does `cos_theta_plus_theta_u < cosf(bcone.theta_e - bcone.theta_o)`
reduce to `fast_cos(x) < cos(x)` in this specific case?
- `cos_theta_plus_theta_u` is computed as
`cos_theta * cos_theta_u - sin_theta * sin_theta_u`
- `cos_theta` is always 1.0 in the case of a single distant light.
- `cos_theta_u` is computed earlier as `fast_cosf(theta_e)` in
`distant_light_tree_parameters()`
- `sin_theta` is zero, and so that side of the equation doesn't matter.
This reduces `cos_theta_plus_theta_u` to `fast_cosf(theta_e)`.
`cosf(bcone.theta_e - bcone.theta_o)` reduces to `cosf(bcone.theta_e)`
because for the case of a single distant light `theta_o` is always 0.
Pull Request: https://projects.blender.org/blender/blender/pulls/131932
In volume segment, the minimal angle formed by the emitter bounding cone
axis and the vector pointing from the cluster centroid to any point on
the ray is computed via `dot(bcone.axis, point_to_centroid)`, see Fig.8.
in paper.
For distant light this angle is 0, but due to numerical issues this is
not always true. Therefore explicitly assign `-bcone.axis` to
`point_to_centroid` in this case.
Pull Request: https://projects.blender.org/blender/blender/pulls/129489
Gitea would complain the apostrophe in one of the code comments in
tree.h was an ambiguous Unicode character. So fix it by swapping it
for a more common apostrophe type.
Use a bounding sphere instead of the corners of a bounding box to
compute the subtended angle of a light tree node.
Using the corners of the bounding box was an underestimate in some
scenes, causing some light tree nodes being incorrectly skipped.
Using the subtended angle of a bounding sphere is an overestimate, but
it covers the entire node and would not skip any valid contribution,
and no other reliable algorithm to compute the minimal enclosing angle
is known to us.
We expect some increase in noise due to overestimation, but this has
not been observed yet, in our benchmark scenes only a difference in
noise is visible.
Thanks to Weizhen for the suggestion to use the bounding sphere.
Pull Request: https://projects.blender.org/blender/blender/pulls/126625
The original paper only considers the minimal distance of the cluster to
the ray, not the interval length, resulting in low weight for distant
lights that have large influence over a long distance.
This commit modifies the measure by considering `theta_b - theta_a` for
local lights and the ray length `t` for distant lights.
Pull Request: https://projects.blender.org/blender/blender/pulls/123537
This fixes#69535 and #98930.
We use a equi-solid-angle sampling algorithm for rectangular area lights,
but it is not particularly robust for small area lights (either small
in general and/or small because it's being viewed from grazing angles).
The actual sampling part is fine since it just gets clamped into the
valid area anyways, and the difference isn't notable for small lights.
However, we also need to compute the solid angle to get the sampling PDF,
and that computation is quite sensitive to numerical issues for small
values.
Therefore, this commit adds a fallback path for small values, which instead
uses the classic equi-area sampling PDF term times the area-to-solid-angle
Jacobian term. This approximation assumes that all points on the light have
the same distance and angle to the sampling point, which is of course not
strictly the case, but it's close enough for small area lights and better
than failing altogether.
Pull Request: https://projects.blender.org/blender/blender/pulls/122323
Reformulates some terms in the equi-solid-angle rectangle sampling code to
handle small area lamps better, and allows for some rounding error in the
check whether the sampled position is inside the area light.
Pull Request: https://projects.blender.org/blender/blender/pulls/122323
In the original paper, the falloff inside `bcone.theta_e` is assumed to
be `pi/2`, which is too large for spot light and resulted in an
overestimation near the cone boundary.
To address this issue, attenuate the energy of a spot light using the
minimal possible angle formed by the light axis and the shading point
when traversing the light tree.
Ref: #122362
Pull Request: https://projects.blender.org/blender/blender/pulls/122667
The refactor in 97d9bbbc97 changed the way q is computed in the spherical triangle sampling code. While the new approach is more efficient and saves a few operations, it introduces numerical precision issues for skinny/small (spherical) triangles.
Therefore, this change moves the computation of q back to the method from the paper, while keeping the more efficient solid angle computation.
Pull Request: https://projects.blender.org/blender/blender/pulls/119224
it is difficult to keep in mind when MIS weight is needed, better to
handle this logic in the lower-level functions.
This reduces code duplication in many places.
it is not clear from which point the `cos_theta_u` should be computed in
volume segment, so the original implementation was mixing the closest point
and the point where the minimal angle is formed.
Use the closest point on segment as a conservative measure.
Pull Request: https://projects.blender.org/blender/blender/pulls/119965
in the test scene `all_light_types_in_volume.blend`, `theta - theta_o -
theta_u` is slightly above the threshold. Even if we do a strict check
with `acos` on the failing cases, it will go to the other branch and
deliver a result which is also 1.0f. Better to relax the threshold.