Skip to content

fix(content): keep graphemes whole and excerpts fast - #5

Merged
m1ngsama merged 4 commits into
mainfrom
fix/excerpt-graphemes
Sep 26, 2026
Merged

m1ngsama merged 4 commits into
mainfrom
fix/excerpt-graphemes

Conversation

@m1ngsama

Copy link
Copy Markdown
Member

Excerpts and summaries split emoji, and 40663fe made search excerpts about 10× slower.

  • Excerpts and truncated summaries end on a grapheme boundary, so 👩‍💻 is no longer cut to 👩….
  • Excerpt cost is back to the pre-40663fe level: 50 KB 5.2 → 0.66 ms, 300 KB 34.7 → 3.6 ms. Text is normalized in chunks that split only before printable ASCII or CJK U+4E00–9FFF, which NFKC never joins to what precedes them.
  • Excerpts are identical to the current algorithm across 180,000 fuzzed queries mixing İ, fi, compatibility Jamo, ς/Σ, ZWJ and flag emoji, and across 990 excerpts from nbtca/documents.
  • The list-marker pattern lives in one place; prose comments are cut to one-line traps.

@m1ngsama
m1ngsama merged commit 93482c1 into main Sep 26, 2026
5 checks passed
@m1ngsama
m1ngsama deleted the fix/excerpt-graphemes branch September 26, 2026 06:31
@m1ngsama m1ngsama mentioned this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant