blender/intern/cycles/kernel/device/gpu
Patrick Mours 9433b655c3 Cycles: GPU: Optimize parallel prefix sum kernel
Each thread in the launched block runs over a portion of the counter
array. First to calculate a sum of each portion, then summing those across
all threads to get the base offset of each portion, before each thread
does the prefix sum of its local portion again including that base offset.

This could be improved further, but it already solves negative performance
impact when increasing the number of states in the following commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/163930
2026-09-23 15:22:44 +02:00
..
block_sizes.h
image.h Cycles: Fallback value for missing image or attribute in shader 2026-09-15 19:45:36 +02:00
kernel.h Cycles: Enable local shader sorting for CUDA/OptiX 2026-09-15 12:39:35 +02:00
parallel_active_index.h Cycles: Enable local shader sorting for CUDA/OptiX 2026-09-15 12:39:35 +02:00
parallel_prefix_sum.h Cycles: GPU: Optimize parallel prefix sum kernel 2026-09-23 15:22:44 +02:00
parallel_sorted_index.h Cycles: Enable local shader sorting for CUDA/OptiX 2026-09-15 12:39:35 +02:00
work_stealing.h