Open research prototype · low-resource morphology
MORPHOLENS
Can multilingual models understand morphology,or do they mostly memorise word forms?
MorphoLens is an interactive research workspace for exploring morphological structure across high- and low-resource languages, built on UniMorph data and a reproducible lemma-overlap experiment.
Explore how words change across Turkish, Urdu, Evenki, Chukchi and Romanian. Every form is traceable to its source record.
TURKISH
evlerimizden
‘from our houses’
Languages in the workspace
Five languages, very uneven resources.
From 570k UniMorph triples for Turkish to 241 for Chukchi (234 usable in the experiment). The gap is part of the research question, so the interface shows it rather than hiding it.
TURKISH
Türkçe
- Family
- Turkic
- Resources
- Higher-resource comparison
- UniMorph triples
- 570,420
- Lemmas
- 3,579
URDU
اردو
- Family
- Indo-Aryan
- Resources
- Low-resource NLP
- UniMorph triples
- 12,572
- Lemmas
- 182
EVENKI
- Family
- Tungusic
- Resources
- Low-resource
- UniMorph triples
- 11,371
- Lemmas
- 4,495
CHUKCHI
- Family
- Chukotko-Kamchatkan
- Resources
- Low-resource
- UniMorph triples
- 241
- Lemmas
- 196
ROMANIAN
Română
- Family
- Romance (Indo-European)
- Resources
- Higher-resource comparison
- UniMorph triples
- 80,262
- Lemmas
- 4,405
The problem
Familiar lemmas can make a model look smarter than it is.
Random train/test splits often let a model see closely related forms of the same lemma during training.
MorphoLens contrasts this with lemma-disjoint evaluation: can a model generalise when the lexical item itself was never observed?
Measured, not simulated
What the first experiment shows.
Turkish · n = 1000 · paradigm memory vs. baseline
+6.2
points on the random split, where 97% of test lemmas were seen (95% CI +4.6 to +7.8). Once lemmas are disjoint the gain is +0.0.
Urdu · n = 500 · feature-aware vs. baseline, lemma-disjoint
+4.8
points on unseen lemmas from decomposing feature bundles (95% CI +3.8 to +5.9, p < 0.001).
Setup
674,866
UniMorph triples, pinned commits, 5 seeds, three transparent rule-based systems and two PyTorch neural models.

