The import kept only readings of one syllable, so the Hanja search could convert 한 but never 한자, 학교, or 대한민국 — the conversion every other Korean input method offers. libhangul's table has 187k readings; both scripts now keep them all, and the search finds a word as readily as a syllable. The daemon pays for it: 24 MB instead of 12, and 170 ms to start instead of 30.
120 lines
5.5 KiB
Plaintext
120 lines
5.5 KiB
Plaintext
# Dictionary data
|
||
|
||
## Korean Hanja data
|
||
|
||
`hanja.src` is the tracked, reviewable source for Korean Hanja conversion.
|
||
Every data row is Hanja, a tab, and its modern Hangul reading, a character
|
||
or a word:
|
||
|
||
```
|
||
漢 한
|
||
漢字 한자
|
||
```
|
||
|
||
`mkhanja` validates that contract and groups rows by their reading to
|
||
produce the runtime dictionary. Candidate order follows source order. The
|
||
source keeps all retained pairs for review; the generated dictionary stores
|
||
the first 128 candidates per reading because that is the engine's lookup
|
||
limit.
|
||
|
||
The source is derived from libhangul release tag `libhangul-0.2.0`. The
|
||
annotated tag object is `20afc38922e3595ee3ed5b186f2ea05afe663763`, its
|
||
peeled commit is `41c702f5d3581325b646ef6249f1f641b0427ae0`, and
|
||
`data/hanja/hanja.txt` has blob ID
|
||
`199cfd70c4b306257ac2e018714a1285f7cf0ed3`. Import and regenerate it from
|
||
the repository root with:
|
||
|
||
```sh
|
||
curl -L https://raw.githubusercontent.com/libhangul/libhangul/41c702f5d3581325b646ef6249f1f641b0427ae0/data/hanja/hanja.txt -o upstream
|
||
printf '%s %s\n' dd44dcc856cf542b1022d0f39c2e9b9f8805fdcc5923be80f04849ed97ce0996 upstream | sha256sum -c -
|
||
map/libhangul2hanja upstream >map/hanja.src
|
||
map/mkhanja >map/hanja.dict
|
||
```
|
||
|
||
`libhangul2hanja` retains only rows whose reading is modern Hangul syllables
|
||
throughout and whose value is Hanja the popup can draw: U+3400–U+4DBF,
|
||
U+4E00–U+9FFF, or U+F900–U+FAFF. Thus jamo readings, mixed values, and
|
||
supplementary-plane ideographs are omitted. The import preserves the upstream license header and row order.
|
||
The retained data is BSD 3-Clause licensed by Choe Hwanjin; the complete text
|
||
is in `LICENSES/BSD-3-Clause-libhangul-hanja.txt`.
|
||
|
||
## Emoji and symbol data
|
||
|
||
`symbol.src` is hand-written: a symbol, a tab, and its ASCII aliases, kept
|
||
so that a bare `1`-`9` picks the matching superscript or subscript from a
|
||
`^` or `_` prefix search. `emoji.src` is generated: every fully-qualified
|
||
emoji of Unicode's `emoji-test.txt` except the skin-tone variants, with
|
||
its CLDR names and keywords in English, Korean, and Japanese. `mkemoji`
|
||
folds the aliases, adds each Japanese one in the other kana so a query
|
||
typed in either mode finds it, and writes one row per alias to
|
||
`emoji.dict`; the engine searches that dictionary by prefix, so the rows
|
||
carry no prefixes.
|
||
|
||
The sources are Emoji 17.0 and CLDR 48.2.0. Regenerate them from the
|
||
repository root with:
|
||
|
||
```sh
|
||
curl -L https://www.unicode.org/Public/17.0.0/emoji/emoji-test.txt -o emoji-test.txt
|
||
printf '%s %s\n' 1d8a944f88d7952f7ef7c5167fef3c67995bcae24543949710231b03a201acda emoji-test.txt | sha256sum -c -
|
||
for l in en ko ja; do
|
||
curl -L https://raw.githubusercontent.com/unicode-org/cldr-json/48.2.0/cldr-json/cldr-annotations-full/annotations/$l/annotations.json -o $l.json
|
||
curl -L https://raw.githubusercontent.com/unicode-org/cldr-json/48.2.0/cldr-json/cldr-annotations-derived-full/annotationsDerived/$l/annotations.json -o $l-derived.json
|
||
done
|
||
map/cldr2emoji emoji-test.txt en.json en-derived.json ko.json ko-derived.json ja.json ja-derived.json >map/emoji.src
|
||
map/mkemoji >map/emoji.dict
|
||
```
|
||
|
||
The data is under the Unicode License v3; the complete text is in
|
||
`LICENSES/Unicode-3.0.txt`.
|
||
|
||
## Japanese romaji
|
||
|
||
`hira.map` and `kata.map` are Mozc's default romaji table, one row per key,
|
||
in gojūon order; `kata.map` is the same table in Katakana, and
|
||
`make verify-map` checks that the two have the same keys. A doubled
|
||
consonant, or `tch`, maps to っ and keeps the consonant pending; the engine
|
||
knows that rule, the table only lists the rows.
|
||
|
||
## Japanese dictionary data
|
||
|
||
`kanji.dict` is the historical dictionary bundled with strans. Its own header
|
||
identifies it as the SKK Medium dictionary, version 8.1 of May 24, 1995,
|
||
rearranged for Plan 9 ktrans by Kenji Okamoto on February 17, 2000. It was
|
||
already present when this repository was created in commit
|
||
`dcd1147638f908eca7e4b7616815e215022cc99d` on December 23, 2025. No original
|
||
SKK input file or upstream revision was committed, so the current file cannot
|
||
honestly be regenerated byte-for-byte from repository artifacts. The license
|
||
notice in its header permits redistribution and modification under GPL version
|
||
2 or later.
|
||
|
||
The repository copy keeps that historical candidate order. Duplicate keys
|
||
were merged at their first occurrence, later unseen candidates were appended
|
||
in source order, duplicate candidates were removed, and the empty
|
||
`ようたつ` row was removed.
|
||
|
||
## Importing a current SKK dictionary
|
||
|
||
From the repository root, fetch a known upstream revision and convert it with:
|
||
|
||
```
|
||
git clone https://github.com/skk-dev/dict.git map/skkdicts
|
||
git -C map/skkdicts checkout --detach REVISION
|
||
map/skk2ktrans map/skkdicts/SKK-JISYO.M >map/kanji.dict.new
|
||
python3 map/verifymap.py map/kanji.dict.new
|
||
```
|
||
|
||
Use a full skk-dev/dict commit ID for `REVISION` and record it when replacing
|
||
the bundled data. Git is needed only to fetch upstream data; it is not part of
|
||
the normal build image.
|
||
|
||
`skk2ktrans` accepts one or more EUC-JP SKK files (or standard input), writes
|
||
UTF-8 tab-separated rows, and merges input in command-line and source order.
|
||
It strips annotations, deduplicates candidates, and omits candidates containing
|
||
whitespace. Rows containing expressions or escaped candidate delimiters are
|
||
rejected; rewrite or omit those rows before import.
|
||
|
||
`verifymap.py` checks UTF-8, row structure, unique keys, 64-rune keys and
|
||
values, canonical candidate spacing, and duplicate dictionary candidates.
|
||
Pass it the exact `.map` and `.dict` files used for runtime or validation.
|
||
The normal build does not install map data.
|