Skip to content

fix[deburr]: keep letters of non-Latin scripts as they are - #2124

Merged
raon0211 merged 3 commits into
mainfrom
fix/deburr-non-latin
Sep 30, 2026
Merged

raon0211 merged 3 commits into
mainfrom
fix/deburr-non-latin

Conversation

@raon0211

@raon0211 raon0211 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

deburr in es-toolkit normalizes the whole string with NFD and then strips combining marks. NFD decomposes every character it can, not just Latin letters, so text in other scripts is corrupted:

import { deburr } from 'es-toolkit';

deburr('한국어');
// expected: '한국어' (3 characters)
// actual:   '한국어' (8 decomposed Hangul jamo, NFD)

deburr('Мой край');
// expected: 'Мой край'
// actual:   'Мои краи'

deburr('Ελληνικά');
// expected: 'Ελληνικά'
// actual:   'Ελληνικα'

deburr('क़'); // Devanagari KA with nukta
// expected: length 1
// actual:   length 2

The docs already promise that Korean text survives (deburr('résumé-김철수.pdf') // 'resume-김철수.pdf'), but the returned string was decomposed jamo that only looked the same.

This is the es-toolkit counterpart of #2123, which fixes the same problem in es-toolkit/compat by porting Lodash's fixed table. es-toolkit keeps its broader behavior of deburring any Latin letter with diacritics, such as Tiếng Việt → Tieng Viet and Ǎ → A, which Lodash does not handle.

Changes

  • Deburr each character on its own instead of normalizing the whole string. A character is decomposed with NFD only when it is in the Latin-1 Supplement, Latin Extended-A, Latin Extended-B, or Latin Extended Additional block, or is the Kelvin or Angstrom sign. These are the only code points whose NFD base is a Latin letter, verified over U+0080–U+10FFFF. Everything else, including Hangul, Cyrillic, Greek, Indic, and Arabic letters, is returned unchanged without calling normalize().
  • Standalone combining marks are still removed, and the U+00C0–U+017F range still matches Lodash and es-toolkit/compat for every code point.
  • Symbols that decompose into non-letter ASCII, such as ≠ (= + U+0338) and the Greek question mark ;, are now kept as they are since they are not letters. This is the only intentional behavior change for Latin-range input.
  • Add tests that fail on main: Hangul syllable length, Cyrillic/Greek/Kana preservation, Indic/Arabic preservation, Latin letters outside the Lodash range, letters decomposing into Æ/Ø/ſ, surrogate pairs, and a full U+00C0–U+017F comparison against es-toolkit/compat.
  • Document in all four languages that only Latin letters are converted.

An exhaustive check of every code point in U+0080–U+10FFFF confirmed that whenever the NFD base is an ASCII letter, every remaining code unit is a combining mark in the removed ranges, so the per-character rule never leaves partial decompositions behind.

Benchmark results

Measured with Node.js 24 on Linux, total time for the given number of calls (lower is better). Non-Latin characters skip normalize() entirely, so strings made mostly of other scripts are faster than on main. Latin strings pay for one normalize() per character instead of one per string.

Case lodash main This branch
deburr('déjà vu') ×200k 316ms 60ms 72ms
'déjà vu'.repeat(1000) ×300 201ms 47ms 113ms
deburr('hello world foo bar') ×200k 126ms 83ms 47ms
deburr('한국어 테스트 문자열') ×200k 129ms 215ms 94ms
deburr('Café Москва 한국어 Tiếng Việt') ×100k 144ms 115ms 105ms

🤖 Generated with Claude Code

@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
es-toolkit Ready Ready Preview Sep 30, 2026 6:02am UTC

Request Review

T3 Code Service User and others added 3 commits September 30, 2026 05:55
`deburr` normalized the whole string with NFD and stripped combining marks,
which decomposed Hangul syllables into jamo, turned Cyrillic `й` into `и`,
and stripped accents from Greek, Indic, and Arabic letters.

Deburr each character on its own instead: a character is only decomposed
when its base is an ASCII Latin letter or one of the special Latin letters
in `deburrMap`. Everything else is returned unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Only characters in U+00C0-U+0233, U+1E00-U+1EF9, and U+212A-U+212B
decompose into a Latin base, so check the code point before calling
`normalize()` instead of inspecting the decomposed base afterwards.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@raon0211
raon0211 force-pushed the fix/deburr-non-latin branch from 4744572 to 4ec7c2a Compare September 30, 2026 05:57
@raon0211
raon0211 merged commit 43e1118 into main Sep 30, 2026
12 checks passed

This branch was successfully deployed

1 active deployment
Preview — 4ec7c2a4 Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant