The new raycast visibility is not like the other visibility flags, which
cover all possible light transport rays and so PATH_RAY_VISIBILITY_ALL
was a complete mask of those. But raycast rays are separate from that,
and should not affect e.g. volume stack interseciton.
Now define a separate PATH_RAY_VISIBILITY_OBJECT_ALL that indicates all
the possible flags that can be set on an object.
Issue introduced in 6e7a0eee32.
Pull Request: https://projects.blender.org/blender/blender/pulls/164137
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.
Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.
For the Embree and hardware intersection it is implemented as a custom
primitive type.
The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.
There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/163103
This appears to have been caused by invalid use of assert() instead of
kernel_assert() in the kernel. I don't think it's actually hitting that
assert, but maybe it generated an unsupported instruction or something
along those lines?
The one that caused the actual problem is in svm/convert.h. I changed
more instances that are currently CPU only but risk becoming enabled on
the CPU with future code changes.
Thanks to Sahar A. Kashi for finding this.
Pull Request: https://projects.blender.org/blender/blender/pulls/163970
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.
This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.
Pull Request: https://projects.blender.org/blender/blender/pulls/162785
Previously the prim_offset was passed as a geometry user data, which
makes it impossible to implement custom primitives: custom geometry
primitive requires a geometry bounds function and this function has
only access to the geometry user data (despite of the custom pointer
being accepted by rtcSetGeometryBoundsFunction: this pointer is being
simply ignored).
This change makes it possible to set geometry user data that can be
accessed from the bounds function in a way that does not conflict
with the previous prim_offset logic.
Ref #159470
Pull Request: https://projects.blender.org/blender/blender/pulls/162738
Fix integer overflows that were causing this, in the kernel and in tx
generation. The overflow caused both wrong render results and slowness
due to too many tiles being loaded due to wrong derivatives.
Also fix an additional overflows in CPU image sampling without the cache.
Pull Request: https://projects.blender.org/blender/blender/pulls/162195
* Detect tx files associated with images
* Add tiled image loading in ImageCache and OIIOImageLoader
* Pad tiles for fast hardware texture filtering
* Kernel side tile mapping with stochastic mip selection
* Statistics support for mipmaps and tiles loaded
Ref #68917
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.
Pull Request: https://projects.blender.org/blender/blender/pulls/154913
This is an internal change preparing for the texture cache. Only implemented
for surfaces, and currently supports the following nodes:
- Geometry
- Tangent
- Mapping
- Attribute
- Texture Coordinate
- Environment Texture
- Image Texture
- Vector math
- UV Map
- Combine/Separate XYZ
- Bump
This has some impact on GPU rendering performance. Various changes were
made to optimize this, but rendering can still be a few % slower on some
GPUs. Some of the optimizations done:
* Use different node types enum for derivative nodes, consecutive to ensure the
jump table works.
* Template various SVM derivative nodes to separate them from the
non-derivative case, and avoid using dual types in those implementations.
* Use template on return type for stack_store and stack_load to make the above
easier to implement.
* Use template on return type of primitive attribute reading to make derivative
and non-derivative variations.
* Unify derivative and bump dx/dy nodes. Now it's a single derivative node
that handles both cases.
* Derivative nodes are disabled in volume shaders for now.
* Tweak inlining on a few functions.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/155706
Refactor ImageHandle to used pointers to a new ImageSlot istead of slot
indices. ImageSlot has subclasses ImageSingle and ImageUDIM, for
individual images and UDIMs. ImageUDIM has a list of handles for
ImageSingle.
UDIM tile lookup was moved out of SVM and OSL code, and is now shared.
Early loading of images for displacement and volumes was refactored.
Now a set of ImageSingle pointers is collected, which are then loaded
by the image manager.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
Make a distinction between an image texture for shading systems, and a
device image object. For full images this is the same, but for tiled
images we'll store the pixels across multiple device image objects.
Rename slot to image_info_id and image_texture_id to help distinguish
indexes into these.
Pull Request: https://projects.blender.org/blender/blender/pulls/154668
This PR fixes the compilation problems when building Cycles without
OSL support and some Windows-related compile problems caused by the
'_USE_MATH_DEFINES' define.
The PR:
- Adds the cycles_util dependencies to the cycles cpu module.
- Adds the `_USE_MATH_DEFINES` as a global compiler definition when
building under Windows.
- Removes unneed setting and definitions of `_USE_MATH_DEFINES` in the code
and CMake files
Pull Request: https://projects.blender.org/blender/blender/pulls/154209
* Switch to more accurate half and float conversion supporting denormals
to match native instructions. This makes CPU and GPU match more
closely in some tests and avoids clipping some low values.
* Inf and NaN are not supported still, as we already filter these out
and there is no reason to have the overhead.
* Change avx2 kernel to require f16c. For all physical CPUs avx2 implies
f16c, and it's only for emulation and virtual machines that this would
not be the case. So it's fine to fall back to the sse4.1 kernel then.
* Assume half instructions are available with ARM NEON. There is no
defined minimum architecture, but Blender assumes the same and ARMv8.2-A
is relatively old.
* For everything else there are SIMD optimized fallbacks.
* New unit tests were added, coverting both native instructions and
fallback implementations, and half/half3/half4.
Fix#152763: Half float image low values are clipped
Pull Request: https://projects.blender.org/blender/blender/pulls/154042
Before this change the API consisted of two parts: return value which
indicated whether an opaque surface was hit, and throughput which was
used to accumulate curves transparency.
This change makes it so throughput=0 indicates that an opaque hit was
found, allowing to simplify intersection payload and state tracking:
the payload for Embree and MetalRT now has one less boolean flag.
For OptiX there is no direct affect on the payload size, as there is
some explicit rule about Ray using registers p6 and p7 (while the
boolean flag was on p5). It does, however, avoids need to track extra
boolean flag regardless.
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.
Pull Request: https://projects.blender.org/blender/blender/pulls/152665
Split the device-specific logic into individual files that reside
in kernel/device/<device>. Should be no functional change, but it
should make it easier to work with individual devices easier.
Some minor changes compared to prior to this change:
- There is an explicit target cycles_kernel_cpu.
It makes it easier pass header dependencies to GPU backends, but
also makes it possible to only try compile CPU kernels when GPU
binaries are enabled.
- There is a CMake option to allow building different device backends
in parallel: WITH_CYCLES_PARALLEL_DEVICE_KERNEL_BUILD. It is set
to OFF by default, matching old behavior.
Setting it to ON helps in situations when memory is not a concern,
or when it is only a couple of GPU architectures enabled.
Pull Request: https://projects.blender.org/blender/blender/pulls/152241
A series of commits which reduces the number of sign conversions
(int <-> uint) in the Cycles kernel.
While it is not expected that the conversion emits any instructions,
it is quite confusing to follow the code and choose proper type.
Additionally, from some development in !151540 it seemed that such
mismatch was responsible for the performance drop in HIP-RT.
Pull Request: https://projects.blender.org/blender/blender/pulls/152009
Guide the probability to scatter in or transmit through the volume.
Only applied for primary rays.
Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Add new "Linear 3D Curves" option in the Curves panel in the render
properties. This renders curves as linear segments rather than smooth
curves, for faster render time at the cost of accuracy.
On NVIDIA Blackwell GPUs, this can give a 6x speedup compared to smooth
curves, due to hardware acceleration. On NVIDIA Ada there is still
a 3x speedup, and CPU and other GPU backends will also render this
faster.
A difference with smooth curves is that these have end caps, as this
was simpler to implement and they are usually helpful anyway.
In the future this functionality will also be used to properly support
the CURVE_TYPE_POLY on the new curves object.
Pull Request: https://projects.blender.org/blender/blender/pulls/139735
All GPU backends now support NanoVDB, using our own kernel side code
that is easily portable. This simplifies kernel and device code.
Volume bounds are now built from the NanoVDB grid instead of OpenVDB,
to avoid having to keep around the OpenVDB grid after loading.
While this reduces memory usage, it does have a performance impact,
particularly for the Cubic filter. That will be addressed by
another commit.
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
In detail:
- Direct accesses of state attributes are replaced with the INTEGRATOR_STATE and INTEGRATOR_STATE_WRITE macros.
- Unified the checks for the __PATH_GUIDING define to use # if defined (__PATH_GUIDING__).
- Even if __PATH_GUIDING__ is defined, we now check if the feature is enabled using if ((kernel_data.kernel_features & KERNEL_FEATURE_PATH_GUIDING)) {. This is important for later GPU ports.
- The kernel usage of the guiding field, surface, and volume sampling distributions is wrapped behind macros for each specific device (atm only CPU). This will make it easier for a GPU port later.
Embree 4.4 introduces an improvement in the Embree GPU
implementation by dropping shared memory usage in favor
of direct controllable memory transfers. This should allow
addressing several problems spotted in Blender regarding
multithreading and memory corruption when BVH and rendering
happen at the same time. However, to implement such
improvements, the API has changed for several functions, and
this commit adopts Blender code to these changes, making Blender
buildable and functional with all existing Embree 4.X
versions, before and after 4.4.
No functional changes in Blender behavior are expected if
using Embree versions below 4.4.
Pull Request: https://projects.blender.org/blender/blender/pulls/139061
In forward path tracing, when we pass volume bounding meshes, we
accumulate `volume_bounds_bounce`. We should match this behaviour in NEE
instead of accumulating `transparent_bounce`.
Pull Request: https://projects.blender.org/blender/blender/pulls/137556
The transparent bounce test was too optimistic in regards to the intersection
being considered. The check needs to happen after it has been validated that
it is not duplicate.
It was already the case for Metal and HIP-RT, but not for Embree and BVH2.
Tests updated by: Alaska <Alaskayou01@gmail.com>
Pull Request: https://projects.blender.org/blender/blender/pulls/136325
The reason for this to happen is because when spatial split is used
the same intersection could be recorded twice (via different BVH nodes).
This change introduces check for the intersection being already recoded,
similar to the check in the local BVH. The check is done during BVH
intersection which allows to properly ignore intersections even for the
maximum bounce number check. A faster approach would be to do such
filtering after sorting, but then we can not keep bounce check in the
BVH code consistent with and without spatial splits.
Intuitively it seems that it should be possible to merge the new loop
with the one that checks for which intersection to keep. But it is not
so trivial in practice: it doesn't run for all intersections, and also
it is formulated in a way that updates isect_index for the next record.
Pull Request: https://projects.blender.org/blender/blender/pulls/136251
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.
Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361