”

Open research prototype · low-resource morphology

MORPHOLENS

Can multilingual models understand morphology,or do they mostly memorise word forms?

MorphoLens is an interactive research workspace for exploring morphological structure across high- and low-resource languages, built on UniMorph data and a reproducible lemma-overlap experiment.

Explore how words change across Turkish, Urdu, Evenki, Chukchi and Romanian. Every form is traceable to its source record.

Specimen 01 / 05 REFERENCE

TURKISH

evlerimizden

evHOUSE
lerPL
imiz1PL.POSS
denABL

‘from our houses’

The problem

TRAINING DATA
walkwalkedwalking
MODELUNSEEN FORMwalkability?

Familiar lemmas can make a model look smarter than it is.

Random train/test splits often let a model see closely related forms of the same lemma during training.

MorphoLens contrasts this with lemma-disjoint evaluation: can a model generalise when the lexical item itself was never observed?

Measured, not simulated

What the first experiment shows.

Full results

Turkish · n = 1000 · paradigm memory vs. baseline

+6.2

points on the random split, where 97% of test lemmas were seen (95% CI +4.6 to +7.8). Once lemmas are disjoint the gain is +0.0.

Urdu · n = 500 · feature-aware vs. baseline, lemma-disjoint

+4.8

points on unseen lemmas from decomposing feature bundles (95% CI +3.8 to +5.9, p < 0.001).

Setup

674,866

UniMorph triples, pinned commits, 5 seeds, three transparent rule-based systems and two PyTorch neural models.