Dispersion is defined by two parameters:
Abbe Number and Dispersion Scale.
This corresponds to OpenPBR v1.1.1
Pull Request: https://projects.blender.org/blender/blender/pulls/162041
Note: dispersion is temporarily disabled when using MNEE (shadow caustics) on oneAPI, due to a compiler bug. Waiting for a fix from the oneAPI side.

Co-authored-by: Sebastian Herholz <sebastian.herholz@gmail.com>
Fix integer overflows that were causing this, in the kernel and in tx
generation. The overflow caused both wrong render results and slowness
due to too many tiles being loaded due to wrong derivatives.
Also fix an additional overflows in CPU image sampling without the cache.
Pull Request: https://projects.blender.org/blender/blender/pulls/162195
We are running out of bits. Split the original `ShaderDataFlag` into
`ShaderRuntimeFlag`, which is determined by closures in the shader and
set up during rendering, and `ShaderDataFlag`, which is the same for the
whole shader graph and hence determined during shader compilation
The size of `ShaderData` did not change due to padding
Pull Request: https://projects.blender.org/blender/blender/pulls/161856
Evict unused texture cache tiles between render tiles and multiviews, to
reduce memory usage. Tiles that were loaded but not accessed during the
previous render tile are freed before the next tile begins rendering.
A per-tile access state byte tracks three values (NONE, REQUESTED, USED)
and is used both to drive tile loading from cache misses and to decide
which loaded tiles can be evicted at tile boundaries.
Cache eviction is automatically enabled, but there is a debug option to
disable it. There is also an debug option to preserve a specified amount
of unused image texture memory to aoid avoid trashing unnecessarily in
simple cases. It increases peak memory usage to render faster, however
in tests on a Macbook M3 it did not seem worth it. But left in to
experiment on other devices and scenes.
Pull Request: https://projects.blender.org/blender/blender/pulls/157244
* Detect tx files associated with images
* Add tiled image loading in ImageCache and OIIOImageLoader
* Pad tiles for fast hardware texture filtering
* Kernel side tile mapping with stochastic mip selection
* Statistics support for mipmaps and tiles loaded
Ref #68917
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
This is an internal change preparing for the texture cache. Only implemented
for surfaces, and currently supports the following nodes:
- Geometry
- Tangent
- Mapping
- Attribute
- Texture Coordinate
- Environment Texture
- Image Texture
- Vector math
- UV Map
- Combine/Separate XYZ
- Bump
This has some impact on GPU rendering performance. Various changes were
made to optimize this, but rendering can still be a few % slower on some
GPUs. Some of the optimizations done:
* Use different node types enum for derivative nodes, consecutive to ensure the
jump table works.
* Template various SVM derivative nodes to separate them from the
non-derivative case, and avoid using dual types in those implementations.
* Use template on return type for stack_store and stack_load to make the above
easier to implement.
* Use template on return type of primitive attribute reading to make derivative
and non-derivative variations.
* Unify derivative and bump dx/dy nodes. Now it's a single derivative node
that handles both cases.
* Derivative nodes are disabled in volume shaders for now.
* Tweak inlining on a few functions.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/155706
Refactor ImageHandle to used pointers to a new ImageSlot istead of slot
indices. ImageSlot has subclasses ImageSingle and ImageUDIM, for
individual images and UDIMs. ImageUDIM has a list of handles for
ImageSingle.
UDIM tile lookup was moved out of SVM and OSL code, and is now shared.
Early loading of images for displacement and volumes was refactored.
Now a set of ImageSingle pointers is collected, which are then loaded
by the image manager.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.
Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
The regression is caused by c0ad7c16dc.
Seems to be compiler issue. Hinting to unroll the loop solves the problem,
but introduces performance regression. So instead revert the code to the old
state for the HIP platform.
Additionally, move the weighting per-row weighting into the outter loop.
This requires an extra variable, but less multiplications.
Pull Request: https://projects.blender.org/blender/blender/pulls/152321
by using a loop instead of unrolling all 64 evaluations.
Seems that two nested loops has the best performance.
Compilation time measured on Metal M2 Ultra:
| Kernel| Before| After|
| --| --| --|
| integrator_shade_volume|148.13s|114.71s|
|integrator_shade_volume_ray_marching| 44.30s| 14.27s|
| integrator_shade_shadow| 87.83s| 58.82s|
| shader_eval_volume_density| 32.69s| 6.63s|
Also added test file because we were not testing deterministic tricubic
interpolation before
Ref: #150119
Blender grid rendering interprets voxel transforms in such a way that the voxel
values are located at the center of a voxel. This is inconsistent with OpenVDB
where the values are located at the lower corners for the purpose or sampling
and related algorithms.
While it is possible to offset grids when communicating with the OpenVDB
library, this is also error-prone and does not add any major advantage.
Every time a grid is passed to OpenVDB we currently have to take care to
transform by half a voxel to ensure correct sampling weights are used that match
the density displayed by the viewport rendering.
This patch changes volume grid generation, conversion, and rendering code so
that grid transforms match the corner-located values in OpenVDB.
- The volume primitive cube node aligns the grid transform with the location of
the first value, which is now also the same as min/max bounds input of the
node.
- Mesh<->Grid conversion does no longer require offsetting grid transform and
mesh vertices respectively by 0.5 voxels.
- Texture space for viewport rendering is offset by half a voxel, so that it
covers the same area as before and voxel centers remain at the same texture
space locations.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/138449
It has ~1.2x speed-up on CPU and ~1.5x speed-up on GPU (tested on Metal
M2 Ultra).
Individual samples are noisier, but equal time renders are mostly
better.
Note that volume emission renders differently than before.
Pull Request: https://projects.blender.org/blender/blender/pulls/144451
Stochastically turn a tricubic filter into a trilinear one. This
reduces the number of taps from 64 to 8. It combines ideas from
the "Stochastic Texture Filtering" paper and our previous GPU
sampling of 3D textures.
This is currently only used in a few places where we know stochastic
interpolation is valid or close enough in practice.
* Principled volume density, color and temperature
* Motion blur velocity
On an Macbook Pro M3 with the openvdb_smoke.blend regression test
and cubic sampling, this gives a ~2x speedup for CPU and ~4x speedup
for GPU. However it also increases noise, usually only a little. Equal
time renders for this scene show a clear reduction in noise for both
CPU and GPU.
Note we can probably get a bigger speedup with acceptable noise trade-off
using full stochastic sampling, but will investigate that separately.
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.
Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
Along with the 4.1 libraries upgrade, we are bumping the clang-format
version from 8-12 to 17. This affects quite a few files.
If not already the case, you may consider pointing your IDE to the
clang-format binary bundled with the Blender precompiled libraries.
The NanoVDB headers are not compatible with Metal due to missing address
space qualifiers. We currently have a big patch for NanoVDB header
files, which is difficult to update for OpenVDB 11. Instead extract a
few hundred lines of code from NanoVDB to do just what we need.
Pull Request: https://projects.blender.org/blender/blender/pulls/115992
While the multiscattering GGX code is cool and solves the darkening problem at higher roughnesses, it's also currently buggy, hard to maintain and often impractical to use due to the higher noise and render time.
In practice, though, having the exact correct directional distribution is not that important as long as the overall albedo is correct and we a) don't get the darkening effect and b) do get the saturation effect at higher roughnesses.
This can simply be achieved by adding a second lobe (https://blog.selfshadow.com/publications/s2017-shading-course/imageworks/s2017_pbs_imageworks_slides_v2.pdf) or scaling the single-scattering GGX lobe (https://blog.selfshadow.com/publications/turquin/ms_comp_final.pdf). Both approaches require the same precomputation and produce outputs of comparable quality, so I went for the simple albedo scaling since it's easier to implement and more efficient.
Overall, the results are pretty good: All scenarios that I tested (Glossy BSDF, Glass BSDF, Principled BSDF with metallic or transmissive = 1) pass the white furnace test (a material with pure-white color in front of a pure-white background should be indistinguishable from the background if it preserves energy), and the overall albedo for non-white materials matches that produced by the real multi-scattering code (with the expected saturation increase as the roughness increases).
In order to produce the precomputed tables, the PR also includes a utility that computes them. This is not built by default, since there's no reason for a user to run it (it only makes sense for documentation/reproducibility purposes and when making changes to the microfacet models).
Pull Request: https://projects.blender.org/blender/blender/pulls/107958
* Store compact ray differentials in ShaderData and compute full differentials
on demand. This reduces register pressure on the GPU.
* Remove BSDF differential code that was effectively doing nothing as the
differential orientation was discarded when making it compact.
This gives a 1-5% speedup with RTX A6000 + OptiX in our benchmarks, with the
bigger speedups in simpler scenes.
Renders appear to be identical except for the Both displacement option that
does both displacement and bump.
Differential Revision: https://developer.blender.org/D15677
These replace float3 and packed_float3 in various places in the kernel where a
spectral color representation will be used in the future. That representation
will require more than 3 channels and conversion to from/RGB. The kernel code
was refactored to remove the assumption that Spectrum and RGB colors are the
same thing.
There are no functional changes, Spectrum is still a float3 and the conversion
functions are no-ops.
Differential Revision: https://developer.blender.org/D15535
This was tested in some places to check if code was being compiled for the
CPU, however this is only defined in the kernel. Checking __KERNEL_GPU__
always works.
* Rename "texture" to "data array". This has not used textures for a long time,
there are just global memory arrays now. (On old CUDA GPUs there was a cache
for textures but not global memory, so we used to put all data in textures.)
* For CUDA and HIP, put globals in KernelParams struct like other devices.
* Drop __ prefix for data array names, no possibility for naming conflict now that
these are in a struct.
When converting from XYZ to RGB it can happen, in some sky models, that the resulting RGB values are negative.
Atm, this is not considered and the returned values for the sky model can be negative.
This patch clamps the returned RGB values to be `= 0.f`
Reviewed By: brecht, sergey
Differential Revision: https://developer.blender.org/D14777
Keep the existing Rec.709 fit and convert to other colorspace if needed, it
seems accurate enough in practice, and keeps the same performance for the
default case.
* Replace license text in headers with SPDX identifiers.
* Remove specific license info from outdated readme.txt, instead leave details
to the source files.
* Add list of SPDX license identifiers used, and corresponding license texts.
* Update copyright dates while we're at it.
Ref D14069, T95597
saturate is depricated in favour of __saturatef this replaces saturate
with __saturatef on CUDA by createing a saturatef function which replaces
all instances of saturate and are hooked up to the correct function on all
platforms.
Reviewed By: brecht
Differential Revision: https://developer.blender.org/D13010
Remove prefix of filenames that is the same as the folder name. This used
to help when #includes were using individual files, but now they are always
relative to the cycles root directory and so the prefixes are redundant.
For patches and branches, git merge and rebase should be able to detect the
renames and move over code to the right file.
* Split render/ into scene/ and session/. The scene/ folder now contains the
scene and its nodes. The session/ folder contains the render session and
associated data structures like drivers and render buffers.
* Move top level kernel headers into new folders kernel/camera/, kernel/film/,
kernel/light/, kernel/sample/, kernel/util/
* Move integrator related kernel headers into kernel/integrator/
* Move OSL shaders from kernel/shaders/ to kernel/osl/shaders/
For patches and branches, git merge and rebase should be able to detect the
renames and move over code to the right file.