Local partitioned shader sorting has been used with Metal and oneAPI for
a while now. Turns out CUDA/OptiX also benefit, so this change enables it
there too.
Note that there is no need to operate on shared memory with atomics in
CUDA, so the implementation of atomic_store_local/atomic_load_local is
kept extremely simple.
Pull Request: https://projects.blender.org/blender/blender/pulls/163436
Some renderers generate tx files with small tile sizes like 32x32,
which have more overhead than our 64x64 tiles. Now we treat such
files as having 64x64 tiles anyway by loading 4 tiles at once.
There is also an internal debug option to test even larger tile
sizes for future benchmarking.
Pull Request: https://projects.blender.org/blender/blender/pulls/161395
Introduce `MTLResidencySet` to explicitly manage GPU memory residency on macOS 15.0+ devices.
This provides the Metal/GPU driver with a clear list of resources to optimise memory handling, potentially improving performance by reducing overhead from hundreds of `useResource` calls.
Memory allocation and deallocation operations are now routed through new wrapper functions that conditionally add or remove resources from the residency set. Explicit `useResource` calls are bypassed when residency sets are active, as resource residency is now managed through the `MTLResidencySet` API. A debug flag is added to enable or disable this feature at runtime (via env var `CYCLES_METAL_RESIDENCY_SETS=0`).
Pull Request: https://projects.blender.org/blender/blender/pulls/158558
Evict unused texture cache tiles between render tiles and multiviews, to
reduce memory usage. Tiles that were loaded but not accessed during the
previous render tile are freed before the next tile begins rendering.
A per-tile access state byte tracks three values (NONE, REQUESTED, USED)
and is used both to drive tile loading from cache misses and to decide
which loaded tiles can be evicted at tile boundaries.
Cache eviction is automatically enabled, but there is a debug option to
disable it. There is also an debug option to preserve a specified amount
of unused image texture memory to aoid avoid trashing unnecessarily in
simple cases. It increases peak memory usage to render faster, however
in tests on a Macbook M3 it did not seem worth it. But left in to
experiment on other devices and scenes.
Pull Request: https://projects.blender.org/blender/blender/pulls/157244
Currently MetalRT interpolates transformation matrix on per-element basis
which leads to issues like #135659.
This change adds implementation of for decomposed (Scale/Rotate/Translate)
motion interpolation, matching behavior of BVH2 and other HW-RT.
This requires macOS 15 and Xcode 16 in order to use this interpolation.
On older platforms and compilers old interpolation is used.
Currently there is no changes on the user (by default) and it is only
available via CYCLES_METALRT_PCMI environment variable. This is because
there are some issues with complex motion paths that need to be looked
into. Having code available makes it easier to do further debugging.
Ref #135659
Authored by Emma Liu
Pull Request: https://projects.blender.org/blender/blender/pulls/136253
* Use .empty() and .data()
* Use nullptr instead of 0
* No else after return
* Simple class member initialization
* Add override for virtual methods
* Include C++ instead of C headers
* Remove some unused includes
* Use default constructors
* Always use braces
* Consistent names in definition and declaration
* Change typedef to using
Pull Request: https://projects.blender.org/blender/blender/pulls/132361
This commit updates all defines, compiler flags and cleans up some code for unused CPU capabilities.
There should be no functional change, unless it's run on a CPU that supports sse41 but not sse42. It will fallback to the SSE2 kernel in this case.
In preparation for the new SSE4.2 minimum in Blender 4.2.
Pull Request: https://projects.blender.org/blender/blender/pulls/118043
This patch removes a workaround for an issue that is now understood to be undefined behaviour (and fixed by #108176). It also adds two useful debug flags that we would like to be available in Blender 3.6.
Pull Request: https://projects.blender.org/blender/blender/pulls/108322
This patch adds two new kernels: SORT_BUCKET_PASS and SORT_WRITE_PASS. These replace PREFIX_SUM and SORTED_PATHS_ARRAY on supported devices (currently implemented on Metal, but will be trivial to enable on the other backends). The new kernels exploit sort partitioning (see D15331) by sorting each partition separately using local atomics. This can give an overall render speedup of 2-3% depending on architecture. As before, we fall back to the original non-partitioned sorting when the shader count is "too high".
Reviewed By: brecht
Differential Revision: https://developer.blender.org/D16909
While keeping SSE2, SSE4.1 and AVX2. This does not affect hardware support, it
only slightly reduces performance for some older CPUs.
To reduce maintenance cost and improve compile times.
Differential Revision: https://developer.blender.org/D16978
* Replace license text in headers with SPDX identifiers.
* Remove specific license info from outdated readme.txt, instead leave details
to the source files.
* Add list of SPDX license identifiers used, and corresponding license texts.
* Update copyright dates while we're at it.
Ref D14069, T95597
This patch contains many small leftover fixes and additions that are
required for Metal-enablement:
- Address space fixes and a few other small compile fixes
- Addition of missing functionality to the Metal adapter headers
- Addition of various scattered `__KERNEL_METAL__` blocks (e.g. for
atomic support & maths functions)
Ref T92212
Differential Revision: https://developer.blender.org/D13263
Instead of printing debug flags listing various CPU and GPU settings that
may or may not be used, print when we are using them. This include CPU
kernel types, OptiX debugging and CUDA and HIP adaptive compilation. BVH
type was already printed.
Remove prefix of filenames that is the same as the folder name. This used
to help when #includes were using individual files, but now they are always
relative to the cycles root directory and so the prefixes are redundant.
For patches and branches, git merge and rebase should be able to detect the
renames and move over code to the right file.
2021-10-26 15:37:04 +02:00
Renamed from intern/cycles/util/util_debug.cpp (Browse further)