blender/intern/cycles/kernel/device
Patrick Mours 9433b655c3 Cycles: GPU: Optimize parallel prefix sum kernel
Each thread in the launched block runs over a portion of the counter
array. First to calculate a sum of each portion, then summing those across
all threads to get the base offset of each portion, before each thread
does the prefix sum of its local portion again including that base offset.

This could be improved further, but it already solves negative performance
impact when increasing the number of states in the following commit.

Pull Request: https://projects.blender.org/blender/blender/pulls/163930
2026-09-23 15:22:44 +02:00
..
cpu Fix: Cycles: Objects visible only to raycast hit by other rayso 2026-09-19 23:14:17 +02:00
cuda Cycles: Switch to CUDA 13 and enable DLSS 2026-09-18 19:30:10 +02:00
gpu Cycles: GPU: Optimize parallel prefix sum kernel 2026-09-23 15:22:44 +02:00
hip Cleanup: spelling in comments (make check_spelling_*) 2026-04-01 12:26:50 +11:00
hiprt GSplat: Initial rendering support for Cycles 2026-09-16 16:52:10 +02:00
metal GSplat: Initial rendering support for Cycles 2026-09-16 16:52:10 +02:00
oneapi GSplat: Initial rendering support for Cycles 2026-09-16 16:52:10 +02:00
optix Cycles: Switch to CUDA 13 and enable DLSS 2026-09-18 19:30:10 +02:00