GH-116939: Rewrite binarysort() (#116940)

Rewrote binarysort() for clarity.

Also changed the signature to be more coherent (it was mixing sortslice with raw pointers).

No change in method or functionality. However, I left some experiments in, disabled for now
via `#if` tricks. Since this code was first written, some kinds of comparisons have gotten
enormously faster (like for lists of floats), which changes the tradeoffs.

For example, plain insertion sort's simpler innermost loop and highly predictable branches
leave it very competitive (even beating, by a bit) binary insertion when comparisons are
very cheap, despite that it can do many more compares. And it wins big on runs that
are already sorted (moving the next one in takes only 1 compare then).

So I left code for a plain insertion sort, to make future experimenting easier.

Also made the maximum value of minrun a `#define` (``MAX_MINRUN`) to make
experimenting with that easier too.

And another bit of `#if``-disabled code rewrites binary insertion's innermost loop to
remove its unpredictable branch. Surprisingly, this doesn't really seem to help
overall. I'm unclear on why not. It certainly adds more instructions, but they're very
simple, and it's hard to be believe they cost as much as a branch miss.
This commit is contained in:
Tim Peters 2024-03-21 22:27:25 -05:00 • committed by GitHub
parent 97ba910e47
commit 8383915031
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 142 additions and 67 deletions

View file

@ -270,9 +270,9 @@ result. This has two primary good effects:
Computing minrun
----------------
If N < 64, minrun is N. IOW, binary insertion sort is used for the whole
array then; it's hard to beat that given the overheads of trying something
fancier (see note BINSORT).
If N < MAX_MINRUN, minrun is N. IOW, binary insertion sort is used for the
whole array then; it's hard to beat that given the overheads of trying
something fancier (see note BINSORT).
When N is a power of 2, testing on random data showed that minrun values of
16, 32, 64 and 128 worked about equally well. At 256 the data-movement cost
@ -310,12 +310,13 @@ place, and r < minrun is small compared to N), or q a little larger than a
power of 2 regardless of r (then we've got a case similar to "2112", again
leaving too little work for the last merge to do).
Instead we pick a minrun in range(32, 65) such that N/minrun is exactly a
power of 2, or if that isn't possible, is close to, but strictly less than,
a power of 2. This is easier to do than it may sound: take the first 6
bits of N, and add 1 if any of the remaining bits are set. In fact, that
rule covers every case in this section, including small N and exact powers
of 2; merge_compute_minrun() is a deceptively simple function.
Instead we pick a minrun in range(MAX_MINRUN / 2, MAX_MINRUN + 1) such that
N/minrun is exactly a power of 2, or if that isn't possible, is close to, but
strictly less than, a power of 2. This is easier to do than it may sound:
take the first log2(MAX_MINRUN) bits of N, and add 1 if any of the remaining
bits are set. In fact, that rule covers every case in this section, including
small N and exact powers of 2; merge_compute_minrun() is a deceptively simple
function.
The Merge Pattern