Workspace · 03 · Unseen-word generalisation
Does morphology help when the lemma is unseen?
Three transparent inflection systems, five languages, random versus lemma-disjoint splits, measured on pinned UniMorph data.
Language
Evaluation split
Training examples
Scoring
Strict exact match: the prediction must equal the stored form of the test record.
Atomic-tag rules
Baseline · feature bundle treated as one opaque label
Paradigm memory
Surface · reuses seen forms of the same lemma
Feature-aware rules
Morphology-aware · decomposes bundles into features
Neural, atomic tag
PyTorch Transformer with copy · bundle is one input token
Neural, features
Same network · one input token per feature
Task: Morphological inflection. Given a lemma and a UniMorph feature bundle, predict the inflected form. · mean ± s.d. over 5 seeds · test set of 500 items per seed (at most 500; 25% of the sampled records when that is smaller) · data: 570,420 UniMorph triples.
Accuracy by evaluation split · n = 1000
MEASURED · 5 SEEDS · STRICTBars: mean accuracy; whiskers: ±1 s.d. across seeds. The selected split is shown at full opacity.
What the numbers say · n = 1000
STRICTParadigm memory vs. atomic baseline
random
+6.2pts
[+4.6, +7.8] · p < 0.001
lemma-disjoint
±0.0pts
[±0.0, ±0.0] · n.s.
Feature-aware vs. atomic baseline
random
+0.4pts
[+0.1, +0.6] · p = 0.001
lemma-disjoint
+0.4pts
[+0.2, +0.7] · p < 0.001
Neural: features vs. atomic tag
random
+61.4pts
[+59.5, +63.4] · p < 0.001
lemma-disjoint
+47.7pts
[+45.6, +49.6] · p < 0.001
Neural features vs. feature-aware rules
random
+22.9pts
[+20.6, +25.2] · p < 0.001
lemma-disjoint
+7.4pts
[+5.1, +9.7] · p < 0.001
- Test lemmas seen in training (random)
- 97%
- Test bundles unseen (lemma-disjoint)
- 15%
- Test cells with several stored forms
- 0%
Brackets: 95% paired-bootstrap interval over 2,500 pooled test items; grey = not significant. Paradigm memory can only help when a test lemma was seen, so its advantage is confined to the random split.
Where the differences come from · random · n = 1000 · strict
MEASURED| System | Seen bundle | Unseen bundle (16% of items) | Seen lemma (97%) | Unseen lemma | Errors that copy the lemma | Mean edit distance |
|---|---|---|---|---|---|---|
| Atomic-tag rules | 49.7% | 0.0% | 41.9% | 42.3% | 28.5% | 2.63 |
| Paradigm memory | 57.0% | 0.0% | 48.3% | 42.3% | 31.9% | 2.44 |
| Feature-aware rules | 49.8% | 1.8% | 42.3% | 43.6% | 0.1% | 1.91 |
| Neural, atomic tag | 4.5% | 0.3% | 3.8% | 3.8% | 0.7% | 5.10 |
| Neural, features | 76.7% | 3.3% | 65.5% | 55.1% | 2.1% | 1.82 |
Accuracy on test items whose feature bundle (or lemma) does / does not occur in the training sample, pooled over seeds (2,500 items). A lemma copy means the output is the lemma unchanged; for the rule systems this happens when no rule applies.
Learning curve · random
MEASURED · STRICTReal predictions · random · n = 500
MODEL OUTPUT · SEED 1| Lemma · bundle | Stored form | Atomic-tag rules | Paradigm memory | Feature-aware rules | Neural, atomic tag | Neural, features |
|---|---|---|---|---|---|---|
| veterinerN;GEN;SG;PSS2Sonly feature-aware correct | veterinerinin | veteriner | veteriner | veterinerinin | veterinerlerinizde | veterinerinin |
| komikleşmekV;IND;PRS;HAB;3;PL;NEG;DECLonly feature-aware correctunseen bundle | komikleşmezler | komikleşmek | komikleşmek | komikleşmezler | komikleşecek miydin | komikleşecektim |
| el sıkışmakV;IND;PRS;PROG;3;PL;INTR;POSonly feature-aware correctunseen bundle | el sıkışıyorlar mı | el sıkışmak | el sıkışmak | el sıkışıyorlar mı | el sıkışıyor olmayacak mıymış | el sıkışır olmayacaklar mıymız |
| idadiN;ABL;SGonly memory correct | idadiden | idadi | idadiden | idadi | idadide | idadinden |
| kekN;ACC;SGonly memory correct | keki | keğı | keki | keğı | kekleriniz | kekini |
| konveks dörtgenN;NOM;PLonly memory correct | konveks dörtgenler | konveks dörtgenlar | konveks dörtgenler | konveks dörtgenlar | konveks dörtgenlerin | konveks dörtgenler |
| portN;NOM;PL;PSS3Sall systems wrong | portları | portleri | portleri | portleri | portları | portları |
| kızanN;ABL;SG;PSS2Pall systems wrong | kızanınızdan | kızannizden | kızannizden | kızannizden | kızanımızda | kızanınızdan |
Selected automatically from the test set to show disagreements; not a random sample, so it says nothing about overall accuracy.
Findings across languages
| Dataset | n | Baseline · LD | Feature-aware · LD | Δ feature-aware [95% CI] | Neural atomic · LD | Neural features · LD | Δ neural features vs. atomic [95% CI] | Edit distance · LD | Δ memory · random [95% CI] | Lemma overlap | Multi-form cells |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Turkish | 500 | 37.3 ± 6.5 | 37.4 ± 6.5 | +0.2 [±0.0, +0.4]n.s. | 1.7 ± 0.6 | 27.4 ± 2.8 | +25.6 [+24.0, +27.4]p < 0.001 | 2.85 → 2.23 | +2.4 [+1.4, +3.5]p < 0.001 | 80% | 0% |
| Urdu | 500 | 65.2 ± 5.3 | 70.0 ± 4.4 | +4.8 [+3.8, +5.9]p < 0.001 | 36.3 ± 5.7 | 59.1 ± 9.7 | +22.8 [+20.5, +25.0]p < 0.001 | 1.30 → 0.47 | +3.7 [+2.2, +5.3]p < 0.001 | 94% | 0% |
| Evenki | 500 | 42.8 ± 4.1 | 42.6 ± 4.4 | −0.2 [−0.6, +0.2]n.s. | 16.2 ± 5.1 | 29.0 ± 5.7 | +12.8 [+11.2, +14.3]p < 0.001 | 2.02 → 1.81 | +0.2 [±0.0, +0.4]n.s. | 34% | 16% |
| Chukchi | 100 | 56.6 ± 3.3 | 58.3 ± 5.4 | +1.7 [−0.3, +3.8]n.s. | 22.8 ± 8.2 | 25.2 ± 5.3 | +2.4 [−1.7, +6.2]n.s. | 1.74 → 1.56 | ±0.0 [±0.0, ±0.0]n.s. | 17% | 4% |
| Romanian (verbs) | 500 | 79.6 ± 4.9 | 79.6 ± 4.9 | ±0.0 [±0.0, +0.1]n.s. | 21.6 ± 5.5 | 54.8 ± 2.6 | +33.2 [+31.2, +35.2]p < 0.001 | 0.44 → 0.43 | −6.4 [−7.6, −5.2]p < 0.001 | 80% | 0% |
| Romanian, all records (raw)unreliable noun/adjective tags; not a morphological claim | 500 | 55.1 ± 3.9 | 55.5 ± 4.0 | +0.4 [+0.2, +0.7]p = 0.001 | n/a | n/a | n/a | 1.51 → 1.47 | +0.2 [−1.5, +1.8]n.s. | 79% | 40% |
LD = lemma-disjoint. Accuracy in % (mean ± s.d. over seeds). Δ = difference in accuracy points from a paired bootstrap over all test items pooled across seeds (2,000 resamples); “n.s.” = the 95% interval includes zero. Edit distance = mean Levenshtein distance to the stored form (baseline → feature-aware; lower is better). The neural columns are the same PyTorch network given the bundle as one token (atomic) or one token per feature, trained on the same items. Multi-form cells = share of test items whose lemma and bundle have more than one stored form. Romanian (verbs) excludes nouns and adjectives because their Number and Gender tags are unreliable; the raw run on all records is shown separately. Chukchi (n = 100): of 241 UniMorph records, 234 are usable after excluding 7 with non-schema tags, and each test set holds 58 items per seed; it is shown for completeness only.
How these numbers were produced
One universe, two splits.
- Per seed, sample lemmas at random and keep up to 10 cells each until ~3,000 triples.
- Hold out a test set of up to 500 triples (25% for small languages): by item for the random split, by lemma for the lemma-disjoint split.
- Train each system on the first n ∈ {50, 100, 250, 500, 1000} triples of the remaining pool (sizes the pool cannot support are skipped).
- Score exact-match accuracy; repeat for seeds 1, 2, 3, 4, 5 and report mean ± s.d.
Reproduce: npm run data:fetch && npm run experiment · script: scripts/run-experiment.mjs
Atomic-tag rules learn lemma→form edit rules (prefix and suffix rewrites around the longest shared substring) keyed by the whole feature bundle, and apply the rule of the training lemma with the longest shared ending, in the spirit of the CoNLL-SIGMORPHON 2017 non-neural baseline.
Paradigm memory first looks for other forms of the same lemma in training and reinflects from them; otherwise it falls back to the baseline. It can only profit from lemma overlap.
Feature-aware rules decompose bundles into UniMorph features. For an unseen bundle they inflect into a seen bundle A and apply a rule for the feature difference A→T learned from any training paradigm (e.g. NOM→ABL), falling back to the most similar bundle.
These are deliberately simple, inspectable systems, not multilingual neural encoders. They establish the evaluation pipeline that XLM-R-style models would be run through next.
Why lemma-disjoint?
A surface form can be new while the word is not.
Suppose a model sees walk, walking and walked during training and walks during testing. The surface form is technically unseen, but the lexical item is familiar.
A lemma-disjoint split removes every form of the test lemma from training, creating a stronger test of morphological generalisation. Lemma overlap has been shown to inflate reported inflection accuracy (Goldman, Guriel & Tsarfaty, 2022).
Random split (lemma D leaks into training)
TRAIN
TEST
Lemma-disjoint split (lemma D is never seen)
TRAIN
TEST

