data(hanja): a lone consonant is a reading too

Every Korean keyboard's 한자 key answers a lone consonant with the KS X
1001 symbol palette, and has since 한글 워드프로세서: ㅁ for ※ ○ △ ㈜, ㄴ
for the brackets, ㄹ for the units, ㅇ for the circled numbers.  strans
sends that key to the same search as a syllable -- Khanja is Ctrl+H at
strans.c:830, and startsearch seeds the query with whatever ko.c left
pending -- but every one of hanja.dict's 187286 readings is a syllable, so
the popup came up with a query in it and nothing to pick:

	ㅁ: 0 candidates
	ㄴ: 0 candidates
	ㄹ: 0 candidates
	한: 99 candidates 韓 漢 寒 限 閑 恨 旱 汗 翰 邯 罕 悍 澣 閒 瀚

libhangul ships that palette beside the Hanja table already imported here:
data/hanja/mssymbol.txt, same commit, same author, same BSD-3 terms, same
key:value:comment format -- and keyed by the compatibility jamo ko.c
already holds, U+3141 for ㅁ.  So the engine does not change at all; the
same dictlookup on the same trie now finds something:

	ㅁ: 75 candidates # & * @ § ※ ☆ ★ ○ ● ◎ ◇ ◆ □ ■ △ ▲ ▽ ▼
	ㄴ: 23 candidates " ( ) [ ] { } ‘ ’ “ ” 〔 〕 〈 〉 《 》 「 」
	ㄹ: 94 candidates $ % ₩ F ′ ″ ℃ Å ¢ £ ¥ ¤ ℉ ‰ € ㎕ ㎖ ㎗ ℓ
	한: 99 candidates 韓 漢 寒 限 閑 恨 旱 汗 翰 邯 罕 悍 澣 閒 瀚

Both scripts widen by one rule -- a syllable reading gives Hanja, a jamo
reading gives a symbol -- and hanja.src regenerates byte for byte as it
was, because upstream's own non-syllable readings are words like ㄱ자집
whose values were never Hanja and still fall out.  985 of mssymbol.txt's
987 rows survive: its ideographic space and its soft hyphen do not, since
a candidate the popup cannot draw is not a candidate, and the row format
separates candidates with a space besides.

The two keyspaces cannot collide -- one is syllables, one is single jamo --
so the 187286 existing rows are unchanged, byte for byte, and 18 rows join
them.  mkhanja takes a source list as mkemoji already does, and keeps each
upstream header, which is why the licence text now appears twice.

89 unit, check-live, check-stress and valgrind all clean.  The five new
assertions were checked by breaking the change five ways: dropping
mssymbol.src from SOURCES, letting issymbol keep a formatting character,
letting a jamo reading keep Hanja, widening isjamo to the vowels, and
making mkhanja reject jamo readings.  Each fails only the tests that exist
for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-18 01:56:23 +09:00
parent 82078a60e7
commit 62ddc2177b
8 changed files with 1211 additions and 55 deletions

View File

@@ -1,8 +1,9 @@
# License inventory # License inventory
- `BSD-3-Clause-libhangul-hanja.txt` contains the terms for the - `BSD-3-Clause-libhangul-hanja.txt` contains the terms for the Hanja data
single-character Hanja data in `map/hanja.src` and generated in `map/hanja.src` and the symbol data in `map/mssymbol.src`, and for the
`map/hanja.dict`, derived from libhangul's `data/hanja/hanja.txt`. generated `map/hanja.dict` they share, derived from libhangul's
`data/hanja/hanja.txt` and `data/hanja/mssymbol.txt`.
- `GPL-2.0-or-later.txt` contains GNU GPL version 2. The header of - `GPL-2.0-or-later.txt` contains GNU GPL version 2. The header of
`map/kanji.dict` grants the option to use GPL version 2 or any later `map/kanji.dict` grants the option to use GPL version 2 or any later
version. version.

View File

@@ -16,7 +16,7 @@ and a GTK 3 module.
| `Ctrl+T` | English | | `Ctrl+T` | English |
| `Ctrl+V` | Vietnamese Telex | | `Ctrl+V` | Vietnamese Telex |
| `Ctrl+E` | Emoji and symbol search | | `Ctrl+E` | Emoji and symbol search |
| `Ctrl+H` | One-shot Hanja search | | `Ctrl+H` | One-shot Hanja and symbol search |
Only a plain `Ctrl` chord is a strans key: with `Shift`, `Alt` or `Super` Only a plain `Ctrl` chord is a strans key: with `Shift`, `Alt` or `Super`
held it commits what is pending and goes to the application, so held it commits what is pending and goes to the application, so
@@ -55,8 +55,11 @@ emoji's CLDR name and keywords in English, Korean and Japanese, and against
ASCII aliases such as `->` and `<=`. `Ctrl+H` takes the syllable being ASCII aliases such as `->` and `<=`. `Ctrl+H` takes the syllable being
composed as its query and composes on from it, converting a word as well as composed as its query and composes on from it, converting a word as well as
a syllable — 한자 gives 漢字, 대한민국 gives 大韓民國 — and `Esc` gives the a syllable — 한자 gives 漢字, 대한민국 gives 大韓民國 — and `Esc` gives the
syllable back. Text already committed belongs to the application and syllable back. A lone consonant is a reading too, and answers with the
cannot be converted. symbol table a Korean keyboard's 한자 key has always offered: ㅁ gives ※ ○
△ ㈜, ㄴ the brackets 「」『』, ㄹ the units ℃ ㎏ , ㅇ the circled numbers
①②③. Text already committed belongs to the application and cannot be
converted.
Dead keys and Compose sequences are composed by strans itself for the Dead keys and Compose sequences are composed by strans itself for the
Wayland, XIM and IBus frontends, from `XCOMPOSEFILE` or the locale; the Wayland, XIM and IBus frontends, from `XCOMPOSEFILE` or the locale; the

View File

@@ -1,42 +1,53 @@
# Dictionary data # Dictionary data
## Korean Hanja data ## Korean Hanja and symbol data
`hanja.src` is the tracked, reviewable source for Korean Hanja conversion. `hanja.src` and `mssymbol.src` are the tracked, reviewable sources for what
Every data row is Hanja, a tab, and its modern Hangul reading, a character the Hanja search converts. Every data row is a result, a tab, and its
or a word: Hangul reading. A Hanja reads as a syllable or a word; a symbol reads as
the lone consonant a Korean keyboard's 한자 key offers it under:
``` ```
漢 한 漢 한
漢字 한자 漢字 한자
※ ㅁ
``` ```
`mkhanja` validates that contract and groups rows by their reading to `mkhanja` validates that contract and groups the rows of both sources by
produce the runtime dictionary. Candidate order follows source order. The their reading to produce the runtime dictionary. The two keyspaces do not
source keeps all retained pairs for review; the generated dictionary stores meet: a reading is either syllables or one jamo. Candidate order follows
the first 128 candidates per reading because that is the engine's lookup source order. The sources keep all retained pairs for review; the
limit. generated dictionary stores the first 128 candidates per reading because
that is the engine's lookup limit.
The source is derived from libhangul release tag `libhangul-0.2.0`. The The sources are derived from libhangul release tag `libhangul-0.2.0`. The
annotated tag object is `20afc38922e3595ee3ed5b186f2ea05afe663763`, its annotated tag object is `20afc38922e3595ee3ed5b186f2ea05afe663763`, its
peeled commit is `41c702f5d3581325b646ef6249f1f641b0427ae0`, and peeled commit is `41c702f5d3581325b646ef6249f1f641b0427ae0`, and
`data/hanja/hanja.txt` has blob ID `data/hanja/hanja.txt` and `data/hanja/mssymbol.txt` have blob IDs
`199cfd70c4b306257ac2e018714a1285f7cf0ed3`. Import and regenerate it from `199cfd70c4b306257ac2e018714a1285f7cf0ed3` and
the repository root with: `31c4e63d74293b6a759d4a2005cd1e7333746631`. Import and regenerate them
from the repository root with:
```sh ```sh
curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/hanja.txt -o upstream curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/hanja.txt -o upstream
printf '%s %s\n' dd44dcc856cf542b1022d0f39c2e9b9f8805fdcc5923be80f04849ed97ce0996 upstream | sha256sum -c - printf '%s %s\n' dd44dcc856cf542b1022d0f39c2e9b9f8805fdcc5923be80f04849ed97ce0996 upstream | sha256sum -c -
map/libhangul2hanja upstream >map/hanja.src map/libhangul2hanja upstream >map/hanja.src
curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/mssymbol.txt -o upstream
printf '%s %s\n' b685a4ebe2716b25eb29c42f6da6493716ecd78b604b44a33420987d21948e99 upstream | sha256sum -c -
map/libhangul2hanja upstream >map/mssymbol.src
map/mkhanja >map/hanja.dict map/mkhanja >map/hanja.dict
``` ```
`libhangul2hanja` retains only rows whose reading is modern Hangul syllables `libhangul2hanja` keeps a syllable reading only when its value is Hanja the
throughout and whose value is Hanja the popup can draw: U+3400U+4DBF, popup can draw: U+3400U+4DBF, U+4E00U+9FFF, or U+F900U+FAFF. Thus mixed
U+4E00U+9FFF, or U+F900U+FAFF. Thus jamo readings, mixed values, and values and supplementary-plane ideographs are omitted. It keeps a jamo
supplementary-plane ideographs are omitted. The import preserves the upstream license header and row order. reading, U+3131U+314E, only when its value is one rune the popup can draw
The retained data is BSD 3-Clause licensed by Choe Hwanjin; the complete text and a candidate row can carry — so `mssymbol.txt`'s ideographic space and
is in `LICENSES/BSD-3-Clause-libhangul-hanja.txt`. soft hyphen are omitted, leaving 985 of its 987 rows, and a Hanja under a
jamo reading is omitted as noise. The import preserves each upstream
license header and row order. The retained data is BSD 3-Clause licensed
by Choe Hwanjin; the complete text is in
`LICENSES/BSD-3-Clause-libhangul-hanja.txt`.
## Emoji and symbol data ## Emoji and symbol data

View File

@@ -24,6 +24,32 @@
;; CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ;; CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
;; ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE ;; ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
;; POSSIBILITY OF SUCH DAMAGE. ;; POSSIBILITY OF SUCH DAMAGE.
;; Copyright (c) 2005-2014 Choe Hwanjin
;; All rights reserved.
;;
;; Redistribution and use in source and binary forms, with or without
;; modification, are permitted provided that the following conditions are met:
;;
;; 1. Redistributions of source code must retain the above copyright notice,
;; this list of conditions and the following disclaimer.
;; 2. Redistributions in binary form must reproduce the above copyright notice,
;; this list of conditions and the following disclaimer in the documentation
;; and/or other materials provided with the distribution.
;; 3. Neither the name of the author nor the names of its contributors
;; may be used to endorse or promote products derived from this software
;; without specific prior written permission.
;;
;; THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
;; AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
;; IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE
;; ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE
;; LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR
;; CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF
;; SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS
;; INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN
;; CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE)
;; ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE
;; POSSIBILITY OF SUCH DAMAGE.
가 可 家 加 歌 價 街 假 佳 暇 架 价 賈 伽 柯 迦 軻 嘉 駕 嫁 稼 哥 呵 苛 袈 訶 枷 珂 茄 痂 跏 斝 葭 舸 笳 哿 坷 檟 謌 珈 砢 佉 呿 榎 仮 傢 咖 瘕 耞 宊 椵 叚 価 猳 髂 瘸 鴐 䴥 拁 㤎 榢 毠 溊 牁 豭 槚 诃 贾 轲 镓 驾 鲄 婽 尜 岢 幏 彁 徦 愘 戓 戨 斚 樖 泇 渮 滒 炣 牫 牱 犌 玍 癿 笴 糘 胢 腵 蚵 貑 跒 酠 鉫 鉲 鎵 鎶 魺 鴚 麚 㗎 㚙 㞹 㢦 㤉 㧝 㪃 㪼 㱒 㹢 䂟 䈔 䑝 䔅 䕒 䪪 䯊 䶗 嘏 가 可 家 加 歌 價 街 假 佳 暇 架 价 賈 伽 柯 迦 軻 嘉 駕 嫁 稼 哥 呵 苛 袈 訶 枷 珂 茄 痂 跏 斝 葭 舸 笳 哿 坷 檟 謌 珈 砢 佉 呿 榎 仮 傢 咖 瘕 耞 宊 椵 叚 価 猳 髂 瘸 鴐 䴥 拁 㤎 榢 毠 溊 牁 豭 槚 诃 贾 轲 镓 驾 鲄 婽 尜 岢 幏 彁 徦 愘 戓 戨 斚 樖 泇 渮 滒 炣 牫 牱 犌 玍 癿 笴 糘 胢 腵 蚵 貑 跒 酠 鉫 鉲 鎵 鎶 魺 鴚 麚 㗎 㚙 㞹 㢦 㤉 㧝 㪃 㪼 㱒 㹢 䂟 䈔 䑝 䔅 䕒 䪪 䯊 䶗 嘏
가가 假家 可呵 可嘉 家家 呵呵 가가 假家 可呵 可嘉 家家 呵呵
@@ -187311,3 +187337,21 @@
힐책 詰責 힐책 詰責
힐척 詰斥 힐척 詰斥
힐항 詰抗 힐항 詰抗
_  ̄ 、 。 · ‥ … ¨ 〃 ― ∥ ´ ˇ ˘ ˝ ˚ ˙ ¸ ˛ ¡ ¿ ː
ㄲ Æ Ð Ħ IJ Ŀ Ł Ø Œ Þ Ŧ Ŋ æ đ ð ħ ı ij ĸ ŀ ł ø œ ß þ ŧ ŋ ʼn
“ ” 〈 〉 《 》 「 」 『 』 【 】
± × ÷ ≠ ≤ ≥ ∞ ∴ ♂ ♀ ∠ ⊥ ⌒ ∂ ∇ ≡ ≒ ≪ ≫ √ ∽ ∝ ∵ ∫ ∬ ∈ ∋ ⊆ ⊇ ⊂ ⊃ ∩ ∧ ¬ ⇒ ⇔ ∀ ∃ ∮ ∑ ∏
ㄸ ぁ あ ぃ い ぅ う ぇ え ぉ お か が き ぎ く ぐ け げ こ ご さ ざ し じ す ず せ ぜ そ ぞ た だ ち ぢ っ つ づ て で と ど な に ぬ ね の は ば ぱ ひ び ぴ ふ ぶ ぷ へ べ ぺ ほ ぼ ぽ ま み む め も ゃ や ゅ ゆ ょ よ ら り る れ ろ ゎ わ ゐ ゑ を ん
″ ℃ Å ¢ £ ¥ ¤ ℉ ‰ € ㎕ ㎖ ㎗ ㎘ ㏄ ㎣ ㎤ ㎥ ㎦ ㎙ ㎚ ㎛ ㎜ ㎝ ㎞ ㎟ ㎠ ㎡ ㎢ ㏊ ㎍ ㎎ ㎏ ㏏ ㎈ ㎉ ㏈ ㎧ ㎨ ㎰ ㎱ ㎲ ㎳ ㎴ ㎵ ㎶ ㎷ ㎸ ㎹ ㎀ ㎁ ㎂ ㎃ ㎄ ㎺ ㎻ ㎼ ㎽ ㎾ ㎿ ㎐ ㎑ ㎒ ㎓ ㎔ Ω ㏀ ㏁ ㎊ ㎋ ㎌ ㏖ ㏅ ㎭ ㎮ ㎯ ㏛ ㎩ ㎪ ㎫ ㎬ ㏝ ㏐ ㏓ ㏃ ㏉ ㏜ ㏆
§ ※ ☆ ★ ○ ● ◎ ◇ ◆ □ ■ △ ▲ ▽ ▼ → ← ↑ ↓ ↔ 〓 ◁ ◀ ▷ ▶ ♤ ♠ ♡ ♥ ♧ ♣ ⊙ ◈ ▣ ◐ ◑ ▒ ▤ ▥ ▨ ▧ ▦ ▩ ♨ ☏ ☎ ☜ ☞ ¶ † ‡ ↕ ↗ ↙ ↖ ↘ ♭ ♩ ♪ ♬ ㉿ ㈜ № ㏇ ™ ㏂ ㏘ ℡ ® ª º
ㅂ ─ │ ┌ ┐ ┘ └ ├ ┬ ┤ ┴ ┼ ━ ┃ ┏ ┓ ┛ ┗ ┣ ┳ ┫ ┻ ╋ ┠ ┯ ┨ ┷ ┿ ┝ ┰ ┥ ┸ ╂ ┒ ┑ ┚ ┙ ┖ ┕ ┎ ┍ ┞ ┟ ┡ ┢ ┦ ┧ ┩ ┪ ┭ ┮ ┱ ┲ ┵ ┶ ┹ ┺ ┽ ┾ ╀ ╁ ╃ ╄ ╅ ╆ ╇ ╈ ╉ ╊
ㅃ ァ ア ィ イ ゥ ウ ェ エ ォ オ カ ガ キ ギ ク グ ケ ゲ コ ゴ サ ザ シ ジ ス ズ セ ゼ ソ ゾ タ ダ チ ヂ ッ ツ ヅ テ デ ト ド ナ ニ ヌ ネ ハ バ パ ヒ ビ ピ フ ブ プ ヘ ベ ペ ホ ボ ポ マ ミ ム メ モ ャ ヤ ュ ユ ョ ヨ ラ リ ル レ ロ ヮ ワ ヰ ヱ ヲ ン ヴ ヵ ヶ
ㅅ ㉠ ㉡ ㉢ ㉣ ㉤ ㉥ ㉦ ㉧ ㉨ ㉩ ㉪ ㉫ ㉬ ㉭ ㉮ ㉯ ㉰ ㉱ ㉲ ㉳ ㉴ ㉵ ㉶ ㉷ ㉸ ㉹ ㉺ ㉻ ㈀ ㈁ ㈂ ㈃ ㈄ ㈅ ㈆ ㈇ ㈈ ㈉ ㈊ ㈋ ㈌ ㈍ ㈎ ㈏ ㈐ ㈑ ㈒ ㈓ ㈔ ㈕ ㈖ ㈗ ㈘ ㈙ ㈚ ㈛
А Б В Г Д Е Ё Ж З И Й К Л М Н О П Р С Т У Ф Х Ц Ч Ш Щ Ъ Ы Ь Э Ю Я а б в г д е ё ж з и й к л м н о п р с т у ф х ц ч ш щ ъ ы ь э ю я
ㅇ ⓐ ⓑ ⓒ ⓓ ⓔ ⓕ ⓖ ⓗ ⓘ ⓙ ⓚ ⓛ ⓜ ⓝ ⓞ ⓟ ⓠ ⓡ ⓢ ⓣ ⓤ ⓥ ⓦ ⓧ ⓨ ⓩ ① ② ③ ④ ⑤ ⑥ ⑦ ⑧ ⑨ ⑩ ⑪ ⑫ ⑬ ⑭ ⑮ ⒜ ⒝ ⒞ ⒟ ⒠ ⒡ ⒢ ⒣ ⒤ ⒥ ⒦ ⒧ ⒨ ⒩ ⒪ ⒫ ⒬ ⒭ ⒮ ⒯ ⒰ ⒱ ⒲ ⒳ ⒴ ⒵ ⑴ ⑵ ⑶ ⑷ ⑸ ⑹ ⑺ ⑻ ⑼ ⑽ ⑾ ⑿ ⒀ ⒁ ⒂
ⅱ ⅲ ⅳ ⅵ ⅶ ⅷ ⅸ Ⅱ Ⅲ Ⅳ Ⅵ Ⅶ Ⅷ Ⅸ
ㅊ ½ ⅓ ⅔ ¼ ¾ ⅛ ⅜ ⅝ ⅞ ¹ ² ³ ⁴ ⁿ ₁ ₂ ₃ ₄
ㅋ ㄱ ㄲ ㄳ ㄴ ㄵ ㄶ ㄷ ㄸ ㄹ ㄺ ㄻ ㄼ ㄽ ㄾ ㄿ ㅀ ㅁ ㅂ ㅃ ㅄ ㅅ ㅆ ㅇ ㅈ ㅉ ㅊ ㅋ ㅌ ㅍ ㅎ ㅏ ㅐ ㅑ ㅒ ㅓ ㅔ ㅕ ㅖ ㅗ ㅘ ㅙ ㅚ ㅛ ㅜ ㅝ ㅞ ㅟ ㅠ ㅡ ㅢ ㅣ
ㅌ ㅥ ㅦ ㅧ ㅨ ㅩ ㅪ ㅫ ㅬ ㅭ ㅮ ㅯ ㅰ ㅱ ㅲ ㅳ ㅴ ㅵ ㅶ ㅷ ㅸ ㅹ ㅺ ㅻ ㅼ ㅽ ㅾ ㅿ ㆀ ㆁ ㆂ ㆃ ㆄ ㆅ ㆆ ㆇ ㆈ ㆉ ㆊ ㆋ ㆌ ㆍ ㆎ
Α Β Γ Δ Ε Ζ Η Θ Ι Κ Λ Μ Ν Ξ Ο Π Ρ Σ Τ Υ Φ Χ Ψ Ω α β γ δ ε ζ η θ ι κ λ μ ν ξ ο π ρ σ τ υ φ χ ψ ω

View File

@@ -1,7 +1,8 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Extract Hanja readings from libhangul data.""" """Extract Hanja and symbol readings from libhangul data."""
import sys import sys
import unicodedata
from pathlib import Path from pathlib import Path
@@ -9,12 +10,31 @@ def ishangul(s):
return s != "" and all(0xAC00 <= ord(c) <= 0xD7A3 for c in s) return s != "" and all(0xAC00 <= ord(c) <= 0xD7A3 for c in s)
def isjamo(s):
return len(s) == 1 and 0x3131 <= ord(s) <= 0x314E
def ishanja(s): def ishanja(s):
return s != "" and all(0x3400 <= ord(c) <= 0x4DBF return s != "" and all(0x3400 <= ord(c) <= 0x4DBF
or 0x4E00 <= ord(c) <= 0x9FFF or 0x4E00 <= ord(c) <= 0x9FFF
or 0xF900 <= ord(c) <= 0xFAFF for c in s) or 0xF900 <= ord(c) <= 0xFAFF for c in s)
def issymbol(s):
"""One rune the popup can draw and a candidate row can carry: not a
space, which the row separates candidates with, and not a formatting
character, which would leave a blank candidate to pick."""
return (len(s) == 1 and not s.isspace() and not ishanja(s)
and not unicodedata.category(s).startswith("C"))
def keeps(reading, value):
"""A syllable reading gives Hanja; a jamo reading gives a symbol."""
if ishangul(reading):
return ishanja(value)
return isjamo(reading) and issymbol(value)
def extract(src, name): def extract(src, name):
comments = [] comments = []
entries = [] entries = []
@@ -32,16 +52,16 @@ def extract(src, name):
fields = line.split(":") fields = line.split(":")
if len(fields) != 3: if len(fields) != 3:
raise ValueError(f"{name}:{lineno}: need key:value:comment") raise ValueError(f"{name}:{lineno}: need key:value:comment")
reading, hanja, _ = fields reading, value, _ = fields
if not ishangul(reading) or not ishanja(hanja): if not keeps(reading, value):
continue continue
pair = (hanja, reading) pair = (value, reading)
if pair in seen: if pair in seen:
raise ValueError(f"{name}:{lineno}: duplicate Hanja reading") raise ValueError(f"{name}:{lineno}: duplicate reading")
seen.add(pair) seen.add(pair)
entries.append(pair) entries.append(pair)
if not entries: if not entries:
raise ValueError(f"{name}: no Hanja readings") raise ValueError(f"{name}: no readings")
return comments, entries return comments, entries
@@ -63,8 +83,8 @@ def main():
print(line) print(line)
if comments: if comments:
print() print()
for hanja, reading in entries: for value, reading in entries:
print(f"{hanja}\t{reading}") print(f"{value}\t{reading}")
except (OSError, UnicodeError, ValueError) as error: except (OSError, UnicodeError, ValueError) as error:
print(error, file=sys.stderr) print(error, file=sys.stderr)
return 1 return 1

View File

@@ -1,11 +1,13 @@
#!/usr/bin/env python3 #!/usr/bin/env python3
"""Write hanja.dict from a Hanja-per-row UTF-8 source.""" """Write hanja.dict from result-per-row UTF-8 sources."""
import sys import sys
import unicodedata
from pathlib import Path from pathlib import Path
SOURCE = Path(__file__).with_name("hanja.src") SOURCES = [Path(__file__).with_name(name)
for name in ("hanja.src", "mssymbol.src")]
MAXCANDIDATES = 128 MAXCANDIDATES = 128
@@ -13,17 +15,29 @@ def ishangul(s):
return s != "" and all(0xAC00 <= ord(c) <= 0xD7A3 for c in s) return s != "" and all(0xAC00 <= ord(c) <= 0xD7A3 for c in s)
def isjamo(s):
return len(s) == 1 and 0x3131 <= ord(s) <= 0x314E
def ishanja(s): def ishanja(s):
return s != "" and all(0x3400 <= ord(c) <= 0x4DBF return s != "" and all(0x3400 <= ord(c) <= 0x4DBF
or 0x4E00 <= ord(c) <= 0x9FFF or 0x4E00 <= ord(c) <= 0x9FFF
or 0xF900 <= ord(c) <= 0xFAFF for c in s) or 0xF900 <= ord(c) <= 0xFAFF for c in s)
def read(path): def issymbol(s):
"""One rune the popup can draw and a candidate row can carry: not a
space, which the row separates candidates with, and not a formatting
character, which would leave a blank candidate to pick."""
return (len(s) == 1 and not s.isspace() and not ishanja(s)
and not unicodedata.category(s).startswith("C"))
def read(path, table, seen):
"""Adds path's rows to table, keyed by reading, and returns its header."""
comments = [] comments = []
table = {}
seen = set()
leading = True leading = True
found = False
with path.open(encoding="utf-8") as src: with path.open(encoding="utf-8") as src:
for lineno, raw in enumerate(src, 1): for lineno, raw in enumerate(src, 1):
line = raw.rstrip("\r\n") line = raw.rstrip("\r\n")
@@ -36,31 +50,36 @@ def read(path):
leading = False leading = False
fields = line.split("\t") fields = line.split("\t")
if len(fields) != 2: if len(fields) != 2:
raise ValueError(f"{path}:{lineno}: need Hanja<TAB>reading") raise ValueError(f"{path}:{lineno}: need result<TAB>reading")
hanja, reading = fields value, reading = fields
if not ishanja(hanja): if ishangul(reading):
if not ishanja(value):
raise ValueError(f"{path}:{lineno}: need BMP Hanja") raise ValueError(f"{path}:{lineno}: need BMP Hanja")
if not ishangul(reading): elif isjamo(reading):
raise ValueError(f"{path}:{lineno}: need Hangul syllables") if not issymbol(value):
pair = (hanja, reading) raise ValueError(f"{path}:{lineno}: need one symbol rune")
else:
raise ValueError(f"{path}:{lineno}: need a syllable or jamo "
"reading")
pair = (value, reading)
if pair in seen: if pair in seen:
raise ValueError(f"{path}:{lineno}: duplicate Hanja reading") raise ValueError(f"{path}:{lineno}: duplicate reading")
seen.add(pair) seen.add(pair)
table.setdefault(reading, []).append(hanja) found = True
if not seen: table.setdefault(reading, []).append(value)
raise ValueError(f"{path}: no Hanja readings") if not found:
return comments, table raise ValueError(f"{path}: no readings")
return comments
def main(): def main():
sys.stdout.reconfigure(encoding="utf-8") sys.stdout.reconfigure(encoding="utf-8")
sys.stderr.reconfigure(encoding="utf-8") sys.stderr.reconfigure(encoding="utf-8")
if len(sys.argv) > 2: paths = [Path(arg) for arg in sys.argv[1:]] or SOURCES
print(f"usage: {sys.argv[0]} [hanja.src]", file=sys.stderr) table = {}
return 2 seen = set()
path = Path(sys.argv[1]) if len(sys.argv) == 2 else SOURCE
try: try:
comments, table = read(path) comments = [line for path in paths for line in read(path, table, seen)]
for line in comments: for line in comments:
print(line) print(line)
if comments: if comments:

1012
map/mssymbol.src Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -64,6 +64,24 @@ class MkhanjaTest(unittest.TestCase):
"\u91d1\t\uae40", "\u91d1\t\uae40",
"\u6f22\u5b57\t\ud55c\uc790"]) "\u6f22\u5b57\t\ud55c\uc790"])
def test_imports_symbol_rows_under_a_jamo(self):
result = run(
IMPORT,
self.source(
"\u3141:\u203b:reference mark\n"
"\u3134:\u300c:bracket\n"
"\u3131:\u3000:ideographic space\n"
"\u3131:\u00ad:soft hyphen\n"
"\u3131:\u52a0:Hanja under a jamo\n"
"\u3141:\u203b\u203b:two runes\n"
"\u3141\u3134:\u203b:two jamo\n"
"\u314f:\u203b:a vowel\n"
),
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stdout.splitlines(),
["\u203b\t\u3141", "\u300c\t\u3134"])
def test_import_rejects_malformed_and_duplicate_rows(self): def test_import_rejects_malformed_and_duplicate_rows(self):
bad = [ bad = [
"한 漢\n", "한 漢\n",
@@ -110,6 +128,19 @@ class MkhanjaTest(unittest.TestCase):
], ],
) )
def test_groups_symbols_under_their_jamo(self):
result = run(
GENERATE,
self.source(
"\t\n"
"\t\n"
"\t\n"
),
)
self.assertEqual(result.returncode, 0, result.stderr)
self.assertEqual(result.stdout.splitlines(),
["\t※ ○", "\t"])
def test_generator_rejects_bad_rows(self): def test_generator_rejects_bad_rows(self):
bad = [ bad = [
"漢a\t\n", "漢a\t\n",
@@ -117,6 +148,12 @@ class MkhanjaTest(unittest.TestCase):
"𠀀\t\n", "𠀀\t\n",
"\t\n\t\n", "\t\n\t\n",
"漢 한\n", "漢 한\n",
"\t\n",
"\t\n",
"\tㅁㄴ\n",
"※※\t\n",
"\u3000\t\n",
"\u00ad\t\n",
] ]
for text in bad: for text in bad:
with self.subTest(text=repr(text)): with self.subTest(text=repr(text)):
@@ -131,6 +168,15 @@ class MkhanjaTest(unittest.TestCase):
self.assertEqual(result.returncode, 0, result.stderr) self.assertEqual(result.returncode, 0, result.stderr)
self.assertIn("\t", result.stdout) self.assertIn("\t", result.stdout)
def test_generator_reads_both_sources(self):
result = run(GENERATE)
self.assertEqual(result.returncode, 0, result.stderr)
lines = result.stdout.splitlines()
self.assertEqual(lines.count(";; All rights reserved."), 2)
rows = dict(line.split("\t") for line in lines if "\t" in line)
self.assertIn("", rows[""].split())
self.assertIn("", rows[""].split())
def test_generator_keeps_runtime_candidate_limit(self): def test_generator_keeps_runtime_candidate_limit(self):
rows = "".join(f"{chr(0x4E00 + n)}\t\n" for n in range(129)) rows = "".join(f"{chr(0x4E00 + n)}\t\n" for n in range(129))
result = run(GENERATE, self.source(rows)) result = run(GENERATE, self.source(rows))