# NOTICE — Data Sources & Attribution

OpenJLPT is a **derived work**. It is assembled from the open sources below, plus original
content written for the project (the grammar dataset and the corrections file). Because it
incorporates data licensed under **CC BY-SA 4.0** (a share-alike license), the entire
OpenJLPT dataset is released under **CC BY-SA 4.0** (see `LICENSE`).

If you use OpenJLPT, you must:

1. Give appropriate credit to OpenJLPT **and** the upstream sources listed here.
2. Provide a link to the CC BY-SA 4.0 license.
3. Distribute any derivative dataset under CC BY-SA 4.0 (ShareAlike).

A line like this is enough:

> Contains data from [OpenJLPT](https://github.com/evanclan/OpenJLPT) (CC BY-SA 4.0), which uses
> JMdict and KANJIDIC2 (EDRDG), Jonathan Waller's JLPT lists, and Tatoeba.

## Sources

| Source | Used for | License | Link |
|---|---|---|---|
| **Jonathan Waller's JLPT Resources** (tanos.co.uk) | The **N5–N1 level assignments** for vocabulary and kanji, plus the English glosses for vocabulary. A verbatim snapshot is in `sources/waller/`. | CC BY | https://www.tanos.co.uk/jlpt/ |
| **JMdict** — Electronic Dictionary Research and Development Group (EDRDG) | Vocabulary `jmdict_id` and `pos`, verification and repair of readings, spellings and truncated glosses | CC BY-SA 4.0 | https://www.edrdg.org/jmdict/j_jmdict.html |
| **KANJIDIC2** — EDRDG | Kanji readings, meanings, stroke counts, grade, frequency, radical and name readings | CC BY-SA 4.0 | https://www.edrdg.org/wiki/KANJIDIC_Project.html |
| **Tatoeba** | Vocabulary example sentences (Japanese + English), found via Tatoeba's Japanese word index (the Tanaka Corpus "B lines") | CC BY 2.0 FR | https://tatoeba.org |
| **KanjiVG** (website only) | Stroke-order diagrams shown on the website's kanji pages; not included in the dataset | CC BY-SA 3.0 | https://kanjivg.tagaini.net/ |
| **OpenJLPT contributors** | Grammar points (explanations and example sentences), `sources/corrections/` | CC BY-SA 4.0 | this repository |

The EDRDG files (JMdict, KANJIDIC2) are the property of the Electronic Dictionary Research and
Development Group, and are used in conformance with the Group's
[licence](https://www.edrdg.org/edrdg/licence.html).

Example sentences from [Tatoeba](https://tatoeba.org) are licensed
[CC BY 2.0 FR](https://creativecommons.org/licenses/by/2.0/fr/). Each carries its Tatoeba
sentence ID (`tatoeba_id`), which links to the sentence and its contributors at
`https://tatoeba.org/sentences/show/<id>`. Combining CC BY material into this CC BY-SA 4.0
dataset is permitted, and attribution to Tatoeba is given here.

## Important note on JLPT levels

The Japan Foundation / JLPT organisation does **not** publish official N5–N1 vocabulary, kanji
or grammar lists. Vocabulary and kanji levels come from **Jonathan Waller's community-standard
lists**, which are widely used and reliable, but they are *unofficial approximations* of the
real (undisclosed) test content. KANJIDIC2's own `jlpt` field refers to the **pre-2010
four-level system (1–4)** and is intentionally **not** used for level assignment. Grammar
levels follow the consensus of common JLPT preparation materials.

## Data freshness

Per the EDRDG license, projects redistributing this data should keep it reasonably current.
OpenJLPT rebuilds from upstream monthly (`.github/workflows/update-data.yml`), and
`data/json/meta.json` records the upstream versions used.

## 日和日語 build 10 的教材修改（2026-10-11）

原有 OpenJLPT 衍生詞條 7,814 筆保留原 ID、N5～N1 參考分級、來源例句與修訂釋義。另直接從 EDRDG 的 [JMdict English](https://www.edrdg.org/pub/Nihongo/JMdict_e.gz) 選入 12,186 筆詞條，共 20,000 筆；以詞面加假名讀音去重。新增詞條歸入「延伸」，不另行推定 JLPT 等級，也未新增例句。

JMdict 的詞面、讀音、詞性及英文釋義由 Electronic Dictionary Research and Development Group 提供，依 [EDRDG 授權](https://www.edrdg.org/edrdg/licence.html) 使用。新增繁體中文釋義使用 Apple Translation 在 Mac 本機翻譯，MP3 一筆補寫為「MP3 音訊格式」，其餘未逐條人工校訂。這些選編、翻譯與結構轉換屬衍生教材，仍採 [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)。

原始 JMdict 壓縮檔的 SHA-256：`4855280c53bfd52c41194af2da37b7190a7b8c170f379d643555820dc9f5d10e`。目前完整教材保存於 `NihongoSteps/Resources/content.json`；可重建的選編腳本為 `scripts/expand_japanese_vocabulary.py`，計數與雜湊見 `docs/VOCABULARY_20000_AUDIT.json`。此授權聲明適用於教材資料。
