Commit graph

98 commits

Author SHA1 Message Date
Sergey Sharybin
6b15792508 Cycles: Switch to CUDA 13 and enable DLSS
The commit switches Cycles to use CUDA-13 by default, which makes it
easier to add support for CUDA on Windows arm64 platform.

This commit also enables Cycles DLSS denoising, make Blender ready to
utilize this technology as soon as an updated Nvidia driver is releases
with the required runtime.

The changes are coupled together because they required changes on the
buildbot system side: the SDK's needed to be installed, and the
information about them somehow needed to be passed to Blender. To make
similar deployments easier in the future this change makes it so the SDK
versions from

  build_files/config/pipeline_config.yaml

to

  build_files/buildbot/config/blender_version.cmake

Buildbot provides information about root directory where the specific
SDKs are installed, giving flexibility to the buildbot to move things
around if needed, but also making it more control to Blender developers
to tweak the logic.

Last but not least, the way how CUDA toolkit is selected for Cycles
kernels got refactored to make it easier to follow:

- There is an easy to follow table of per-architecture or family
  toolkits.
- If there is no architectural preference, all provided toolkits are
  probed. For SM kernels newer toolkits are tested first, and for
  COMPUTE the oldest toolkits are probed first.
- If there is no suitable toolkit with explicit major version the
  default one is used.

A driver version 580 and above is now required. From quick checks it
seems that on Windows it shouldn't be a problem since 582 driver is
available for sm_50 devices (the oldest architecture we compile).

Pull Request: https://projects.blender.org/blender/blender/pulls/164002
2026-09-18 19:30:10 +02:00
Sergey Sharybin
819d52f3d2 GSplat: Initial rendering support for Cycles
Based on the "Stochastic ray tracing of transparent 3D Gaussians" paper
by Xin Sun et. al. The basic idea: perform stochastic intersection with
the Gaussian splat based on its transparency.

Gaussian splats are implemented as a dedicated primitive type, but it
shares the same layout for position and radius as points, so a lot of
existing functions (positions, attributes, etc) work for both points
and splats.

For the Embree and hardware intersection it is implemented as a custom
primitive type.

The choice of using bounding spheres mainly comes from a balance between
performance and memory usage. More ideal would be to use OBB, but it is
not supported for custom primitive types in Embree and GPU HW-RT on all
backends.

There is a known limitation that comes from the fact that the datasets
are trained in sRGB space and Cycles work in Linear space: areas with
low opacity and high radiance render noticeably differently from the
ground-truth implementation.

Ref #159470

Pull Request: https://projects.blender.org/blender/blender/pulls/163103
2026-09-16 16:52:10 +02:00
Alex Fuller
3ff1f508e6 Cycles: Fallback value for missing image or attribute in shader
The implements internal support in Cycles for fallback values when an
attribute or image texture is missing in a shader node. It is not yet
exposed in Blender shader nodes.

This is useful to provide an appropriate default value when a shader is
used across objects with different attributes, or when UDIMs don't cover
the entire UV space.

Pull Request: https://projects.blender.org/blender/blender/pulls/162785
2026-09-15 19:45:36 +02:00
Patrick Mours
f97ee31df7 Cleanup: Cycles: Query OptiX hardware ray-tracing state from device
Commit 976ca9d244 fixed the "Hardware
Ray-Tracing" state message being printed to the log always showing off
for OptiX. This improves that further by querying whether the device
actually supports it, since OptiX has a software fallback on devices
without hardware ray tracing support.

Pull Request: https://projects.blender.org/blender/blender/pulls/163807
2026-09-14 15:50:17 +02:00
Sergey Sharybin
f0ec6f183b Cycles: Split shadow modules into own files for OptiX OSL
Both shadow kernels seems to be expensive enough and having them in the
main module leads to very long module creation times.

These new modules are always created (at least for now), but they are
created in parallel with other modules.

Overall it drastically reduces the time it takes to run OptiX OSL tests
for the first time.

Combined with the previous commits this reduces:
- cycles_bake_optix_osl from 1330.44 sec to 177.30 sec
- cycles_attributes_optix_osl from 2809.71 sec to 501.62 sec

Measured on i9-11900K, NVIDIA RTX 6000 Ada.

Pull Request: https://projects.blender.org/blender/blender/pulls/163706
2026-09-10 10:02:12 +02:00
Sergey Sharybin
e03a5166a8 Cycles: Refactor: Use enum for OptiX hit kind
There is limited number of bits available for the user-specified OptiX
hit kind, and currently we used of all those bits available.

The easiest is to use a custom enum for custom types as there is no
real need to couple it it PrimitiveType: the hit kind is only used to
check what kind of intersection it is to fetch extra information and
the actual primitive type for intersection is stored on the OptiX
payload.

Ref #159470

Pull Request: https://projects.blender.org/blender/blender/pulls/162647
2026-08-17 10:33:20 +02:00
Brecht Van Lommel
406577e0d6 Fix #160052: Cycles: Performance regression with CUDA + Blackwell + BVH2
Store positions as separate arrays again. This appears to give better
memory access patterns or layout. This was a regression from fc9917352b
and 0baa98866c.

This is a relatively simple change, rerouting the position attribute to
dedicated arrays, still including motion steps in the same array. The rest
of the refactor is preserved.

Pull Request: https://projects.blender.org/blender/blender/pulls/160110
2026-06-23 12:33:04 +02:00
Brecht Van Lommel
c51fcf73a7 Cycles: Split OptiX kernels into smaller modules to improve load time
This helps especially for OSL, where load time is very long. By
splitting off shader raytrace, MNEE and volumes, the first time
rendering is much faster.

The downside is that noinline functions will be duplicated. However OSL
startup performance is very bad currently and this seems the better
trade-off for now. There are ways to make this work if we do not mark
noinline functions as static, but this will require some bigger code
reorganization.

Without OSL, shader raytrace and MNEE were combined in a single module,
and that has been split into two as well.

Besides a better user experience, This will fix OptiX OSL test timeouts
on the buildbot, where some tests need to compute a few different
specializations and the 600s timeout is exceeded, with the longest test
run time being around 300s.

Pull Request: https://projects.blender.org/blender/blender/pulls/159499
2026-06-04 23:38:05 +02:00
Brecht Van Lommel
0baa98866c Refactor: Cycles: Mesh positions include motion, store kernel attribute
Motion is now part of the position attribute. On the kernel side, this
position is now stored as an attribute as well, replacing the previous
motion only attribute.

Pull Request: https://projects.blender.org/blender/blender/pulls/158728
2026-05-27 22:01:25 +02:00
Brecht Van Lommel
5fc85d1cb7 Refactor: Cycles: Move MNEE walk into separate kernel
The new intersect_mnee kernel runs before shade_surface, and
shade_surface_mnee is eliminated. That large kernel was causing problems
for some GPU compilers.

MNEE state is packed into a shadow path state to avoid significantly
increasing the path state size. This shadow state is then either turned
into an actual shadow ray state or discarded in shade_surface.

MNEE was re-enabled on HIP RDNA2 as it works again now. Texture cache
misses now also work correctly with MNEE.

This adds some extra code to the regular shade_surface kernel even when
MNEE is not used, to use the MNEE sampled point instead of sampling a
light. But there seems to be no significant performance impact.

Co-authored-by: Sergey Sharybin <sergey@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/158698
2026-05-27 20:34:06 +02:00
Sergey Sharybin
d61fa843f5 Refactor: Cycles: Rename path visibility flags
Prefix all bits that are mapped to the BVH visibility mask with
`PATH_RAY_VISIBILITY_`.

Should be no functional changes.

Ref !157822
2026-05-27 19:22:53 +02:00
Brecht Van Lommel
bac976d38f Refactor: Cycles: Add writeable global GPU data arrays
This will be used for tile requests in the texture cache.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Brecht Van Lommel
fa383aa511 Refactor: Cycles: Texture cache miss handling
This adds the various kernel changes needed to handle texture cache
misses on the GPU. For the CPU there is a blocking wait until the pixels
have been loaded from disk. But for GPU we have to restart the kernel.

Pull Request: https://projects.blender.org/blender/blender/pulls/154913
2026-03-27 16:07:06 +01:00
Campbell Barton
48c15c655a CMake: convert macros to functions where possible
Functions scope their variables, macros don't. Use functions
for better variable hygiene, with `PARENT_SCOPE` where output
variables need to propagate.

This avoids the need to `unset(...)` every local variable which is
easy to forget - leaking variables into the callers scope.

Macros that must remain (flag modification, `find_package` forwarding,
paired state, caller-scope `return()`) are annotated with the reason
in their doc-strings.
2026-03-23 07:41:59 +00:00
Sergey Sharybin
6c3a61abb3 Refactor: Cycles, allow granular intersection tests in filter function
Should be no functional changes, preparing for an upcoming refactor.
2026-02-13 17:18:56 +01:00
Sebastian Parborg
0243d0c2fa Deps: Fix fragile openexr dependency linking
When working on the 5.1 library updates, I noticed that quite a few places in Blender where OpenEXR (and libraries depending on OpenEXR) didn't really seem to pull in all of the required libraries properly.
While we are working around this currently, I felt like actually doing the proper thing and using the new cmake libraries interfaces would be the best way forward as we then can rely on them pulling in all needed dependencies by themselves.

To be able to use the .cmake files, two patches are needed:

1. `openexr_deflate_cmake.diff`:
This fixes the cmake install targets to not require libdeflate if it has been complied into openexr as a static library.
2. `osl_relative_inc_cmake.diff`:
This makes it so the library targets no longer uses absolute paths for the include directories. They now use relative paths so that they can be relocated properly.

Co-authored-by: Ray Molenkamp <github@lazydodo.com>
Co-authored-by: Jonas Holzman <jonas@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/153424
2026-02-05 23:10:02 +01:00
Sebastian
87f31c0632 Fix: Remove unused variable compiler warnings when build without PATH_GUIDING
To suppress unused variable compiler warnings when Cycles is build without
path guiding support the `ccl_attr_maybe_unused` macro is added to the
parameters of the affected function definitions.

Pull Request: https://projects.blender.org/blender/blender/pulls/153193
2026-02-03 17:34:42 +01:00
Sergey Sharybin
2b4c0730b3 Refactor: De-duplicate local intersection hit index calculation
Metal-RT backend required a bit of not-so-obvious tricks:

- The LocalIntersection pointer can not be stored in the payload, so
  the utility function has been templated and the Metal-RT payload has
  been modified to match the actual LocalIntersection better.

- Explicit address space is required, and the utility function which
  calculates the index to store the intersection at is used from
  multiple code-paths, some of them require LCG state to be private,
  others require it to be ray_data. Address space can not be a part
  of a template type, so a temporary copy of the lcg_state is done
  in the intersection function. Since the utility is foeceinline the
  hope is that the compiler eliminates it.

- Using bit-packing for the payload fields lead to some unexpected
  test failures: non-deterministic render results, as if something
  was not properly initialized. It is unclear why it started to
  happen now. Using regular int fields avoids sign-comparison warning
  and solves the test failures.

- There was actually one potentially uninitialized member in the
  payload: has_lcg_state was never set to false.

Even with all this extra complexity it does not seem that there is
negative performance impact measured on Apple M4 Max.

Pull Request: https://projects.blender.org/blender/blender/pulls/151540
2026-01-26 14:20:28 +01:00
Sergey Sharybin
f7407a1cb0 Refactor: De-duplicate volume intersection filtering in Cycles
Move the logic to a common place.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
5b5d9f2f0c Refactor: De-duplicate Cycles shadow_all HW-RT function
The goal is to re-use as much of non-trivial logic across devices as
possible, ideally including data-structures and their initialization.

The payload is a bit bigger than it was and it does not utilize OptiX
registers. It seems to be a bit hard to get reliable numbers as they
fluctuate quite a bit and depend on the order of renders, but the
worst performance impact is about 1.5% on the Bistro scene.

It feels there is a room for improvements, like relying on OptiX
itself to keep track on the maximum intersection distance (the code
does accept intersection, but it still tracks the maximum distance in
the filter function), and, maybe, putting payload back on the register.
However, it'll be a bit harder to replicate such behavior on other
platforms, so perhaps it is not a bad trade-off to have.
2026-01-26 14:20:24 +01:00
Sergey Sharybin
3a5f47b6af Refactor: Simplify transparent shadow API in Cycles
Before this change the API consisted of two parts: return value which
indicated whether an opaque surface was hit, and throughput which was
used to accumulate curves transparency.

This change makes it so throughput=0 indicates that an opaque hit was
found, allowing to simplify intersection payload and state tracking:
the payload for Embree and MetalRT now has one less boolean flag.
For OptiX there is no direct affect on the payload size, as there is
some explicit rule about Ray using registers p6 and p7 (while the
boolean flag was on p5). It does, however, avoids need to track extra
boolean flag regardless.
2026-01-26 14:20:24 +01:00
Brecht Van Lommel
4b34743b4e Cycles: Perform direct light shader eval in own kernel
This improves performance by 5-10% for various benchmark scenes and GPU
devices, while on others it's roughly the same. There is a performance
regression with Intel Arc A750 on Linux related to shadow queueing
overhead, that is planned to be fixed separately.

Another goal of this change is to sidestep GPU compiler bugs that seems
more likely to happen with bigger kernels, and to make it easier for the
texture cache to cancel and resume on cache miss.

A new shade_light_nee kernel was added, and shade_light was renamed to
shade_light_forward (following naming for MIS functions). The shade_light_nee
kernel is only used when the light does not have constant emission.

The shade_dedicate_light kernel no longer does any shading. A future
optimization may be to fold this into the intersect_dedicated_light kernel.

LightSample.uv was removed as shading no longer happens immediately. A new
LightPdf was added for the cases where only the pdf is needed, avoiding the
overhead of constructing a full LightSample. There may be more room to
shrink LightSample in future refactors.

The integrate state memory usage is increased by 1 float when not using the
light tree, for the light threshold. All other informating for shading is
reconstructed the shadow ray, including position, normal and uv.

Pull Request: https://projects.blender.org/blender/blender/pulls/152649
2026-01-20 20:34:16 +01:00
Brecht Van Lommel
527f9ea306 Refactor: Cycles: More consistent naming of image functions and structs
Previously there was a mix of "image" and "texture" to refer to the same
thing, use "image" when possible now. An exception is MEM_IMAGE_TEXTURE
to avoid conflicts with the MEM_IMAGE macro on Windows.

Pull Request: https://projects.blender.org/blender/blender/pulls/152665
2026-01-14 17:57:46 +01:00
Sergey Sharybin
be477e0a0c Fix #152417: Main no longer compiles with CUDA 13
Introduced by b2cee9709f: CUDA_VERSION was uninitialized in the
OptiX CMakeLists.txt.

Pull Request: https://projects.blender.org/blender/blender/pulls/152419
2026-01-05 17:03:25 +01:00
Sergey Sharybin
b2cee9709f Refactor: Split Cycles kernel CMakeLists.txt
Split the device-specific logic into individual files that reside
in kernel/device/<device>. Should be no functional change, but it
should make it easier to work with individual devices easier.

Some minor changes compared to prior to this change:

- There is an explicit target cycles_kernel_cpu.

  It makes it easier pass header dependencies to GPU backends, but
  also makes it possible to only try compile CPU kernels when GPU
  binaries are enabled.

- There is a CMake option to allow building different device backends
  in parallel: WITH_CYCLES_PARALLEL_DEVICE_KERNEL_BUILD. It is set
  to OFF by default, matching old behavior.

  Setting it to ON helps in situations when memory is not a concern,
  or when it is only a couple of GPU architectures enabled.

Pull Request: https://projects.blender.org/blender/blender/pulls/152241
2026-01-02 17:29:29 +01:00
Weizhen Huang
188294d8ad Fix #151249: Cycles recording non-volume triangles when intersecting volume
Should check the shader instead of just the object, since an object can
have multiple shaders

Pull Request: https://projects.blender.org/blender/blender/pulls/151499
2025-12-15 12:34:47 +01:00
Brecht Van Lommel
4ae19df885 Cleanup: Cycles: Fix building without some kernel features for debugging
Pull Request: https://projects.blender.org/blender/blender/pulls/151307
2025-12-08 14:09:26 +01:00
Weizhen Huang
2b0a1cae06 Cycles: Add an option to use ray marching for volume rendering
Null Scattering currently has performance and noise issues, and it will
take time to address them. For now add the previous Ray Marching back as
an option.

Co-authored-by: Brecht Van Lommel <brecht@blender.org>
Pull Request: https://projects.blender.org/blender/blender/pulls/146317
2025-09-26 12:14:45 +02:00
Sergey Sharybin
15fd8ad7a1 Fix: Cycles linear curves on Metal-RT
Metal-RT implementation for curve intersect has an additional self
intersection check happening in curve_ribbon_accept(). It is done
for all curve types that has PRIMITIVE_CURVE_RIBBON bit set on them,
including Thick Linear curves. However, the logic in the function is
hardcoded to handle flat ribbon curves with the Catmull Rom basis.

This change makes it so curve_ribbon_accept() is only called for the
ribbon curve type, not when type has ribbon bit set.

Additionally, other places where curve type was checked as a bitmask
were fixed.

Ref #146072

Pull Request: https://projects.blender.org/blender/blender/pulls/146140
2025-09-12 14:16:09 +02:00
Patrick Mours
1b42975e94 Cycles: Add support for building with CUDA 13.0 and OptiX 9.0
The compiler in the CUDA 13 toolkit dropped support for Maxwell, Pascal and Volta architectures (sm_5X, sm_6X and sm_70), which affects both CUDA and OptiX kernel compilation for Cycles. This patch makes it so building CUDA kernel binaries for those architectures are skipped when CUDA 13 is used, but it will still build them if there is a CUDA 11 toolkit available (e.g. on buildbot), like how things are handled for other architectures. The OptiX PTX kernel is compiled with the minimum architecture available (compute_75 with CUDA 13, compute_50 with previous CUDA versions).

In addition, loading the PTX kernel after initializing OptiX version 9.0 would fail with a OPTIX_ERROR_INVALID_FUNCTION_USE, due to the use of "optixTrace" within direct callables (as part of the AO and bevel SVM nodes). Starting with OptiX 9.0 this is no longer allowed, rather one has to use "optixTraverse" in those cases. This patch thus changes the affected intersection routines to use "optixTraverse". As a side effect it also simplifies the `scene_intersect_shadow` function, which no longer invokes the closest hit program, and can just quickly return hit status. The minimum OptiX version Cycles requires is already 8.0, which supports "optixTraverse", so it can just be applied always.

Finally, this patch also adds the `--split-compile=0` argument to nvcc when available, which tells the compiler to internally split the module into pieces that can be processed in parallel on multiple threads (the `=0` notes to use as many threads as there are CPU cores), which can greatly improving compile times, while not making compromises on performance.

Pull Request: https://projects.blender.org/blender/blender/pulls/145130
2025-08-27 14:28:01 +02:00
Weizhen Huang
b2b2d9a4f3 Cycles: Render volume by ray marching through octrees
One octree per volume per shader based on the density. In preparation
for the null scattering
2025-08-13 10:28:50 +02:00
Brecht Van Lommel
dce6269d1f Fix #143714: Cycles OptiX fails to render linear and ribbon curves together
This case was not accounted for previously, but is now possible when
the new curves object has curves with type poly.

Pull Request: https://projects.blender.org/blender/blender/pulls/144087
2025-08-11 19:36:26 +02:00
Patrick Mours
6487395fa5 Cycles: Add linear curve shape
Add new "Linear 3D Curves" option in the Curves panel in the render
properties. This renders curves as linear segments rather than smooth
curves, for faster render time at the cost of accuracy.

On NVIDIA Blackwell GPUs, this can give a 6x speedup compared to smooth
curves, due to hardware acceleration. On NVIDIA Ada there is still
a 3x speedup, and CPU and other GPU backends will also render this
faster.

A difference with smooth curves is that these have end caps, as this
was simpler to implement and they are usually helpful anyway.

In the future this functionality will also be used to properly support
the CURVE_TYPE_POLY on the new curves object.

Pull Request: https://projects.blender.org/blender/blender/pulls/139735
2025-07-29 17:05:01 +02:00
Brecht Van Lommel
b6c4233b28 Refactor: Cycles: Remove now unused 3D image texture support
Pull Request: https://projects.blender.org/blender/blender/pulls/132908
2025-07-09 21:04:38 +02:00
Weizhen Huang
2f7797dd4d Merge branch 'blender-v4.5-release' 2025-06-20 14:20:00 +02:00
weizhen
bf9836da65 Fix: Cycles not building with OptiX 9.0
As suggested by @pmoursnv

Was throwing errors like  `identifier "half" is undefined`.

Pull Request: https://projects.blender.org/blender/blender/pulls/140676
2025-06-20 14:19:43 +02:00
Brecht Van Lommel
7f380e0644 Revert "Fix: Cycles: Do not count volume bounds bounce as transparent"
This reverts commit 23c762e388 in the
blender-v4.5-release branch to work around HIP compiler issues. It will
remain in the main branch.

Ref blender/blender#139836
2025-06-11 15:47:07 +02:00
Lukas Stockner
39d7576844 Cycles: Switch OptiX OSL to use LLVM bitcode for shadeops
This is required to make ray differentials work correctly for OSL custom
cameras.

But it also lets us simplify the implementation, and makes the OSL
functionality more complete, such as implementing all noise types.

Pull Request: https://projects.blender.org/blender/blender/pulls/138161
2025-06-03 20:12:07 +02:00
Lukas Stockner
0dc4754da4 Cycles: Move OptiX OSL Camera kernel into its own PTX module
On the one hand, this improves initialization time since we don't need to
load/compile the full OSL module with all the shading logic if we're only
using a custom camera with SVM shading.

On the other hand, it also fixes a bug I noticed while preparing test scenes:
The AO and Bevel nodes don't work when using custom cameras with SVM on OptiX.

The issue there is that those two are handled by the SHADE_SURFACE_RAYTRACE
kernel, but since that one has intersection logic, we use the OptiX-specific
kernel even if OSL shading is disabled.
However, with the previous unified OSL module, this would mean loading
SHADE_SURFACE_RAYTRACE from kernel_osl.cu, which has `#undef __SVM__` and
therefore doesn't handle them correctly.

With this change, we'll use the kernels from kernel_shader_raytrace.cu in that
case, which do support SVM nodes just fine.

Disk usage of the new kernel_optix_osl_camera.ptx.zst file is 30KB, so this
also doesn't blow up the kernel disk size (and kernel_optix_osl.ptx.zst is
probably smaller by that amount now).

Since it seems that we can mix modules just fine, I'm suspecting that we could
split the modules properly (intersection, SVM shading with raytracing,
OSL shading, OSL camera), instead of the current approach where modules
essentially correspond to feature set tiers and each includes the previous
one's kernels as well - but that's a separate refactor.

Pull Request: https://projects.blender.org/blender/blender/pulls/138021
2025-04-28 12:49:35 +02:00
Lukas Stockner
bf412ed9dd Cycles: Support for custom OSL cameras
This allows users to implement arbitrary camera models using OSL by writing
shaders that take an image position as input and compute ray origin and
direction.

The obvious applications for this are e.g. panorama modes, lens distortion
models and realistic lens simulation, but the possibilities are endless.

Currently, this is only supported on devices with OSL support, so CPU and
OptiX. However, it is independent from the shading model used, so custom
cameras can be used without getting the performance hit of OSL shading.

A few samples are provided as Text Editor templates.

One notable current limitation (in addition to the limited device support)
is that inverse mapping is not supported, so Window texture coordinates and
the Vector pass will not work with custom cameras.

Pull Request: https://projects.blender.org/blender/blender/pulls/129495
2025-04-25 19:27:30 +02:00
Weizhen Huang
23c762e388 Fix: Cycles: Do not count volume bounds bounce as transparent
In forward path tracing, when we pass volume bounding meshes, we
accumulate `volume_bounds_bounce`. We should match this behaviour in NEE
instead of accumulating `transparent_bounce`.

Pull Request: https://projects.blender.org/blender/blender/pulls/137556
2025-04-24 13:10:33 +02:00
Lukas Stockner
8cb5e05c48 Cleanup: Cycles: Deduplicate kernel attribute code using templating
The attribute handling code in the kernel is currently highly duplicated since
it needs to handle five different data types and we couldn't use templates
back then.
We can now, so might as well make use of it and get rid of ~1000 lines.

There are also some small fixes for the GPU OSL code:
- Wrong derivative for .w component when converting float2/float3->float4
- Different conversion for float2->float (CPU averages, GPU used to take .x)
- Removed useless code for converting to float2, not used by OSL

Pull Request: https://projects.blender.org/blender/blender/pulls/134694
2025-02-20 19:28:45 +01:00
Brecht Van Lommel
57ff24cb99 Refactor: Cycles: Add const keyword to more function parameters
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:24 +01:00
Brecht Van Lommel
dd51c8660b Refactor: Cycles: Add const keyword where possible, using clang-tidy
Check was misc-const-correctness, combined with readability-isolate-declaration
as suggested by the docs.

Temporarily clang-format "QualifierAlignment: Left" was used to get consistency
with the prevailing order of keywords.

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:23:20 +01:00
Brecht Van Lommel
d0c2e68e5f Refactor: Cycles: Automated clang-tidy fixups in Cycles
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:55 +01:00
Brecht Van Lommel
5c46063607 Refactor: Cycles: Make kernel headers work by themselves
Shuffle around some code and add more includes so that individual
header files compile without errors.

Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:50 +01:00
Brecht Van Lommel
3c2a6fbb9c Refactor: Cycles: Use nullptr instead of NULL
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
2025-01-03 10:22:43 +01:00
Weizhen Huang
e2d7681fe6 Cleanup: Cycles: remove unused ccl_loop_no_unroll
Was added in 6121c28501 to ensure compiling
on OpenCL, now the definition is empty on all platforms

Pull Request: https://projects.blender.org/blender/blender/pulls/131100
2024-11-28 16:37:01 +01:00
Brecht Van Lommel
d72c4f0096 Fix: Cycles build issues when disabling various kernel features 2024-06-13 19:41:19 +02:00
Michael Jones
5508b41a40 Cycles: MetalRT optimisations (scene_intersect_shadow + random_walk)
This PR contains optimisations and a general tidy-up of the MetalRT backend.

- Currently `scene_intersect` is used for both normal and (opaque) shadow rays, however the usage patterns are different enough to warrant specialisation. Shadow intersection tests (flagged with `PATH_RAY_SHADOW_OPAQUE`) only need a bool result, but need a larger "self" payload in order to exclude hits against target lights. By specialising we can minimise the payload size in each case (which is helps performance) and avoid some dynamic branching. This PR introduces a new `scene_intersect_shadow` function which is specialised in Metal, and currently redirects to `scene_intersect` in the other backends.

- Currently `scene_intersect_local` is implemented for worst-case payload requirements as demanded by `subsurface_disk` (where `max_hits` is 4). The random_walk case only demands 1 hit result which we can retrieve directly from the intersector object (rather than stashing it in the payload). By specialising, we significantly reduce the payload size for random_walk queries, which has a big impact on performance. Additionally, we only need to use a custom intersection function for the first ray test in a random walk (for self-primitive filtering), so this PR forces faster `opaque` intersection testing for all but the first random walk test.

- Currently `scene_intersect_volume` has a lot of redundant code to handle non-triangle primitives despite volumes only being enclosed by trimeshes. This PR removes this code.

Additionally, this PR tidies up the convoluted intersection function linking code, removes some redundant intersection handlers, and uses more consistent naming of intersection functions.

On a M3 MacBook Pro, these changes give 2-3% performance increase on typical scenes with opaque trimesh materials (e.g. barbershop, classroom junkshop), but can give over 15% performance increase for certain scenes using random walk SSS (e.g. monster).

Pull Request: https://projects.blender.org/blender/blender/pulls/121397
2024-05-10 16:38:02 +02:00