Credits Script: Improve author name collation by decomposing them

Some authors have names containing non-English letters. Using the
default sort, those names are written at the end of the list.

This commit uses `unicodedata.normalize('NFD', ...)` during sorting.
This outputs the decomposed unicode variant of a string, such that
diacritized letters are found after their undiacritized version
instead of at the very bottom.

This does not solve collation for names written in a non-latin
alphabet, which would probably require external libraries.

Ref !160527
This commit is contained in:
Damien Picard 2026-06-22 16:09:16 +02:00 • committed by Campbell Barton
parent e2cdb879fe
commit 9ddb532e51

View file

@ -215,14 +215,21 @@ class Credits:
use_email: bool = False,
) -> None:
commit_word = "commit", "commits"
from unicodedata import normalize
# Normalize using decomposed strings for sorting.
# That makes letters with diacritics grouped with the undiacritized variant.
# Collation is not perfect, but better than having non-English letters at the very bottom of the list.
if sort == "commit":
sorted_authors = dict(sorted(
self.users.items(),
key=lambda item: (item[1].commit_total, item[0]),
key=lambda item: (item[1].commit_total, normalize('NFD', item[0])),
))
else:
sorted_authors = dict(sorted(self.users.items()))
sorted_authors = dict(sorted(
self.users.items(),
key=lambda item: (normalize('NFD', item[0]), item[1].commit_total),
))
fh.write("<h3>Individual Contributors</h3>\n\n")
for author, cu in sorted_authors.items():