Skip to content

Commit de26c57

Browse files
docs: the project-ordinal headlines leave the coverage record too (#737)
35 sites — 19 in docs/language-maturity.md, 16 in catalogue.tsv, which is hand-maintained rather than generated (gen-seed.py carries none of them). Same rule as the code sweep: the typology survives, the ordinal goes. "the fleet's FIRST Bodish/Tibetic branch" → "Bodish/Tibetic (Sino-Tibetan)", "the fleet's 4th creole (after ht/kea/pcm)" → "a creole". KEPT, because they are comparisons rather than headlines — each names its comparand or gives the number: - "one of the fleet's thinnest anchors (7 words vs Luo 17 / Madurese 35)" - "one of the fleet's LARGEST referees (63,024 human headwords)" - "like the fleet's other ejective languages Quechua/Georgian" - "the fleet's Iranian modules (fa/ps/ckb) don't cover" A reader can check every one of those; "the fleet's FIRST X" is unfalsifiable without the commit history. NOT TOUCHED: docs/investigations/*.md, which carry ~18 more. Those are the defence thesis — a bring-up log is exactly where "this was the first Mayan language" belongs, and the export deletes the directory anyway. Verified the doc→tool contract survives: masking-coefficient.ts still parses the verdict column out of language-maturity.md, so the table structure is intact. 227 files / 3129 tests unchanged. Claude-Session: https://claude.ai/code/session_012Lc3WnUgogC7okV7n53vjr Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 75993e5 commit de26c57

3 files changed

Lines changed: 50 additions & 44 deletions

File tree

README.md

Lines changed: 15 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -20,19 +20,27 @@ import { phonemizeAsync } from "vernacula-phonemizer";
2020
// Worst input, best output. A language written in two scripts yields ONE canonical IPA either way (Tashelhit in
2121
// Latin/Tifinagh, Fula's Adlam, Bambara's N'Ko, Sundanese's Aksara Sunda, Zhuang's Sawndip).
2222
await phonemizeAsync("I read a book", "en"); // Latin → aᶦ ɹˈɛd ə bˈʊk
23-
await phonemizeAsync("Taclḥit", "shi"); // Berber Latin → taʃlħit
23+
await phonemizeAsync("世界", "cmn"); // Han → ʂʐ̩˥˩ t͡ɕiɛ˥˩
24+
await phonemizeAsync("भारत", "hi"); // Devanagari → bʱˈaːɾət̪
25+
await phonemizeAsync("العربية", "ar"); // Arabic → alʕarabˈijːa
26+
await phonemizeAsync("Україна", "uk"); // Cyrillic → ukrajina
27+
await phonemizeAsync("বাংলাদেশ", "bn"); // Bengali → baŋlad̪eʃ
28+
await phonemizeAsync("ひらがな", "ja"); // Hiragana → çiɾäɡäꜜnä
29+
await phonemizeAsync("カタカナ", "ja"); // Katakana → kätäkänä
30+
await phonemizeAsync("日本語", "ja"); // Kanji → niho̞ŋɡo̞
31+
2432
await phonemizeAsync("ⵜⴰⵛⵍⵃⵉⵜ", "shi"); // Tifinagh → taʃlħit
2533
await phonemizeAsync("𞤆𞤵𞤤𞤢𞥄𞤪", "ff"); // Adlam → pˈulaːɾ
2634
await phonemizeAsync("ߓߊߡߊߣߊ߲", "bm"); // N'Ko → bamanã
2735
await phonemizeAsync("Ελληνικά", "el"); // Greek → elinika
28-
await phonemizeAsync("Україна", "uk"); // Cyrillic → ukrajina
36+
2937
await phonemizeAsync("Հայերեն", "hy"); // Armenian → hɑjeɾen
3038
await phonemizeAsync("ქართული", "ka"); // Georgian → kʰaɾtʰuli
3139
await phonemizeAsync("עברית", "he"); // Hebrew → ʔivʁit
32-
await phonemizeAsync("भारत", "hi"); // Devanagari → bʱˈaːɾət̪
40+
3341
await phonemizeAsync("ਪੰਜਾਬੀ", "pa"); // Gurmukhi → pˈə̃ɲd͡ʒaːbiː
3442
await phonemizeAsync("ગુજરાતી", "gu"); // Gujarati → ɡˈud͡ʒɾat̪i
35-
await phonemizeAsync("বাংলাদেশ", "bn"); // Bengali → baŋlad̪eʃ
43+
3644
await phonemizeAsync("ꠍꠤꠟꠐꠤ", "syl"); // Syloti Nagri → silʈi
3745
await phonemizeAsync("ଓଡ଼ିଆ", "or"); // Odia → ˈoɽia
3846
await phonemizeAsync("தமிழ்", "ta"); // Tamil → t̪ˈɐmɪɻ
@@ -41,17 +49,15 @@ await phonemizeAsync("ಕನ್ನಡ", "kn"); // Kannada → kˈanːaɖa
4149
await phonemizeAsync("മലയാളം", "ml"); // Malayalam → mˈalajaːɭam
4250
await phonemizeAsync("සිංහල", "si"); // Sinhala → sˈiŋhələ
4351
await phonemizeAsync("ᱥᱟᱱᱛᱟᱲᱤ", "sat"); // Ol Chiki → santaɽi
44-
await phonemizeAsync("العربية", "ar"); // Arabic → alʕarabˈijːa
52+
4553
await phonemizeAsync("فارسی", "fa"); // Perso-Arabic → faːɾsˈiː
4654
await phonemizeAsync("አማርኛ", "am"); // Geʽez → amaɾɲa
4755
await phonemizeAsync("မြန်မာ", "my"); // Myanmar → mja˨ɴma˨
4856
await phonemizeAsync("လိၵ်ႈတႆး", "shn"); // Shan → lik̚˧˧˨taj˥
4957
await phonemizeAsync("ខ្មែរ", "km"); // Khmer → kʰmae
5058
await phonemizeAsync("ภาษาไทย", "th"); // Thai → pʰˈaː˧saː˩˩˦tʰˌa˧j
51-
await phonemizeAsync("世界", "cmn"); // Han → ʂʐ̩˥˩ t͡ɕiɛ˥˩
52-
await phonemizeAsync("ひらがな", "ja"); // Hiragana → çiɾäɡäꜜnä
53-
await phonemizeAsync("カタカナ", "ja"); // Katakana → kätäkänä
54-
await phonemizeAsync("日本語", "ja"); // Kanji → niho̞ŋɡo̞
59+
60+
5561
await phonemizeAsync("한국어", "ko"); // Hangul → hˈɐnɡuɡɘ
5662
await phonemizeAsync("ꦗꦮ", "jv"); // Javanese → d͡ʒˈɔwɔ
5763
await phonemizeAsync("ᮞᮥᮔ᮪ᮓ", "su"); // Aksara Sunda → sˈunda

0 commit comments

Comments
 (0)