Workspace · 03 · Unseen-word generalisation

Does morphology help when the lemma is unseen?

Three transparent inflection systems, five languages, random versus lemma-disjoint splits, measured on pinned UniMorph data.

Language

Evaluation split

Training examples

Scoring

Strict exact match: the prediction must equal the stored form of the test record.

Atomic-tag rules

Baseline · feature bundle treated as one opaque label

Paradigm memory

Surface · reuses seen forms of the same lemma

Feature-aware rules

Morphology-aware · decomposes bundles into features

Neural, atomic tag

PyTorch Transformer with copy · bundle is one input token

Neural, features

Same network · one input token per feature

Task: Morphological inflection. Given a lemma and a UniMorph feature bundle, predict the inflected form. · mean ± s.d. over 5 seeds · test set of 500 items per seed (at most 500; 25% of the sampled records when that is smaller) · data: 570,420 UniMorph triples.

Accuracy by evaluation split · n = 1000

MEASURED · 5 SEEDS · STRICT

Bars: mean accuracy; whiskers: ±1 s.d. across seeds. The selected split is shown at full opacity.

What the numbers say · n = 1000

STRICT

Paradigm memory vs. atomic baseline

random

+6.2pts

[+4.6, +7.8] · p < 0.001

lemma-disjoint

±0.0pts

[±0.0, ±0.0] · n.s.

Feature-aware vs. atomic baseline

random

+0.4pts

[+0.1, +0.6] · p = 0.001

lemma-disjoint

+0.4pts

[+0.2, +0.7] · p < 0.001

Neural: features vs. atomic tag

random

+61.4pts

[+59.5, +63.4] · p < 0.001

lemma-disjoint

+47.7pts

[+45.6, +49.6] · p < 0.001

Neural features vs. feature-aware rules

random

+22.9pts

[+20.6, +25.2] · p < 0.001

lemma-disjoint

+7.4pts

[+5.1, +9.7] · p < 0.001

Test lemmas seen in training (random)
97%
Test bundles unseen (lemma-disjoint)
15%
Test cells with several stored forms
0%

Brackets: 95% paired-bootstrap interval over 2,500 pooled test items; grey = not significant. Paradigm memory can only help when a test lemma was seen, so its advantage is confined to the random split.

Where the differences come from · random · n = 1000 · strict

MEASURED
SystemSeen bundleUnseen bundle (16% of items)Seen lemma (97%)Unseen lemmaErrors that copy the lemmaMean edit distance
Atomic-tag rules49.7%0.0%41.9%42.3%28.5%2.63
Paradigm memory57.0%0.0%48.3%42.3%31.9%2.44
Feature-aware rules49.8%1.8%42.3%43.6%0.1%1.91
Neural, atomic tag4.5%0.3%3.8%3.8%0.7%5.10
Neural, features76.7%3.3%65.5%55.1%2.1%1.82

Accuracy on test items whose feature bundle (or lemma) does / does not occur in the training sample, pooled over seeds (2,500 items). A lemma copy means the output is the lemma unchanged; for the rule systems this happens when no rule applies.

Learning curve · random

MEASURED · STRICT

Real predictions · random · n = 500

MODEL OUTPUT · SEED 1
Lemma · bundleStored formAtomic-tag rulesParadigm memoryFeature-aware rulesNeural, atomic tagNeural, features
veterinerN;GEN;SG;PSS2Sonly feature-aware correctveterinerininveterinerveterinerveterinerininveterinerlerinizdeveterinerinin
komikleşmekV;IND;PRS;HAB;3;PL;NEG;DECLonly feature-aware correctunseen bundlekomikleşmezlerkomikleşmekkomikleşmekkomikleşmezlerkomikleşecek miydinkomikleşecektim
el sıkışmakV;IND;PRS;PROG;3;PL;INTR;POSonly feature-aware correctunseen bundleel sıkışıyorlar mıel sıkışmakel sıkışmakel sıkışıyorlar mıel sıkışıyor olmayacak mıymışel sıkışır olmayacaklar mıymız
idadiN;ABL;SGonly memory correctidadidenidadiidadidenidadiidadideidadinden
kekN;ACC;SGonly memory correctkekikeğıkekikeğıkeklerinizkekini
konveks dörtgenN;NOM;PLonly memory correctkonveks dörtgenlerkonveks dörtgenlarkonveks dörtgenlerkonveks dörtgenlarkonveks dörtgenlerinkonveks dörtgenler
portN;NOM;PL;PSS3Sall systems wrongportlarıportleriportleriportleriportlarıportları
kızanN;ABL;SG;PSS2Pall systems wrongkızanınızdankızannizdenkızannizdenkızannizdenkızanımızdakızanınızdan

Selected automatically from the test set to show disagreements; not a random sample, so it says nothing about overall accuracy.

Findings across languages

MEASURED · PAIRED BOOTSTRAP
DatasetnBaseline · LDFeature-aware · LDΔ feature-aware [95% CI]Neural atomic · LDNeural features · LDΔ neural features vs. atomic [95% CI]Edit distance · LDΔ memory · random [95% CI]Lemma overlapMulti-form cells
Turkish50037.3 ± 6.537.4 ± 6.5+0.2 [±0.0, +0.4]n.s.1.7 ± 0.627.4 ± 2.8+25.6 [+24.0, +27.4]p < 0.0012.85 → 2.23+2.4 [+1.4, +3.5]p < 0.00180%0%
Urdu50065.2 ± 5.370.0 ± 4.4+4.8 [+3.8, +5.9]p < 0.00136.3 ± 5.759.1 ± 9.7+22.8 [+20.5, +25.0]p < 0.0011.30 → 0.47+3.7 [+2.2, +5.3]p < 0.00194%0%
Evenki50042.8 ± 4.142.6 ± 4.4−0.2 [−0.6, +0.2]n.s.16.2 ± 5.129.0 ± 5.7+12.8 [+11.2, +14.3]p < 0.0012.02 → 1.81+0.2 [±0.0, +0.4]n.s.34%16%
Chukchi10056.6 ± 3.358.3 ± 5.4+1.7 [−0.3, +3.8]n.s.22.8 ± 8.225.2 ± 5.3+2.4 [−1.7, +6.2]n.s.1.74 → 1.56±0.0 [±0.0, ±0.0]n.s.17%4%
Romanian (verbs)50079.6 ± 4.979.6 ± 4.9±0.0 [±0.0, +0.1]n.s.21.6 ± 5.554.8 ± 2.6+33.2 [+31.2, +35.2]p < 0.0010.44 → 0.43−6.4 [−7.6, −5.2]p < 0.00180%0%
Romanian, all records (raw)unreliable noun/adjective tags; not a morphological claim50055.1 ± 3.955.5 ± 4.0+0.4 [+0.2, +0.7]p = 0.001n/an/an/a1.51 → 1.47+0.2 [−1.5, +1.8]n.s.79%40%

LD = lemma-disjoint. Accuracy in % (mean ± s.d. over seeds). Δ = difference in accuracy points from a paired bootstrap over all test items pooled across seeds (2,000 resamples); “n.s.” = the 95% interval includes zero. Edit distance = mean Levenshtein distance to the stored form (baseline → feature-aware; lower is better). The neural columns are the same PyTorch network given the bundle as one token (atomic) or one token per feature, trained on the same items. Multi-form cells = share of test items whose lemma and bundle have more than one stored form. Romanian (verbs) excludes nouns and adjectives because their Number and Gender tags are unreliable; the raw run on all records is shown separately. Chukchi (n = 100): of 241 UniMorph records, 234 are usable after excluding 7 with non-schema tags, and each test set holds 58 items per seed; it is shown for completeness only.

How these numbers were produced

One universe, two splits.

  1. Per seed, sample lemmas at random and keep up to 10 cells each until ~3,000 triples.
  2. Hold out a test set of up to 500 triples (25% for small languages): by item for the random split, by lemma for the lemma-disjoint split.
  3. Train each system on the first n ∈ {50, 100, 250, 500, 1000} triples of the remaining pool (sizes the pool cannot support are skipped).
  4. Score exact-match accuracy; repeat for seeds 1, 2, 3, 4, 5 and report mean ± s.d.

Reproduce: npm run data:fetch && npm run experiment · script: scripts/run-experiment.mjs

Atomic-tag rules learn lemma→form edit rules (prefix and suffix rewrites around the longest shared substring) keyed by the whole feature bundle, and apply the rule of the training lemma with the longest shared ending, in the spirit of the CoNLL-SIGMORPHON 2017 non-neural baseline.

Paradigm memory first looks for other forms of the same lemma in training and reinflects from them; otherwise it falls back to the baseline. It can only profit from lemma overlap.

Feature-aware rules decompose bundles into UniMorph features. For an unseen bundle they inflect into a seen bundle A and apply a rule for the feature difference A→T learned from any training paradigm (e.g. NOM→ABL), falling back to the most similar bundle.

These are deliberately simple, inspectable systems, not multilingual neural encoders. They establish the evaluation pipeline that XLM-R-style models would be run through next.

Why lemma-disjoint?

A surface form can be new while the word is not.

Suppose a model sees walk, walking and walked during training and walks during testing. The surface form is technically unseen, but the lexical item is familiar.

A lemma-disjoint split removes every form of the test lemma from training, creating a stronger test of morphological generalisation. Lemma overlap has been shown to inflate reported inflection accuracy (Goldman, Guriel & Tsarfaty, 2022).

Random split (lemma D leaks into training)

TRAIN

A1A2A3B1B2C1C2C3D1D2

TEST

A4B3C4D3

Lemma-disjoint split (lemma D is never seen)

TRAIN

A1A2A3A4B1B2B3C1C2C3C4

TEST

D1D2D3