Methodology · data · limitations
How to read this prototype.
MorphoLens does not claim state-of-the-art morphological analysis. It is an inspectable pipeline: pinned public data, simple systems, measured results with uncertainty, and provenance for every form it shows.
Research question
Can explicit morphological structure help models generalise to unseen word forms and unseen lemmas in low-resource languages?
In morphologically rich languages a single lemma yields many surface forms, most of them rare. If test lemmas also occur in training, a system can succeed by recalling the lemma’s paradigm instead of composing morphology. MorphoLens separates the two by comparing random and lemma-disjoint evaluation.
Data
All forms and all experiment inputs come from UniMorph repositories, downloaded at fixed commits. Facts below are taken from each repository’s README.
| Language | Source | Triples | Lemmas | Bundles | Notes |
|---|---|---|---|---|---|
| Turkish | UniMorph tur @ 6c179ac CC BY-SA 3.0 | 570,420 | 3,579 | 883 | Annotators: Omer Goldman & Duygu Ataman. Verbs semi-automatically generated and partially verified by a native speaker; nouns and adjectives from Wiktionary, unverified (per repository README). |
| Urdu | UniMorph urd @ 17b2fb3 | 12,572 | 182 | 217 | No README in the repository; provenance beyond the UniMorph project is not documented there. |
| Evenki | UniMorph evn @ cdbe7b4 | 11,371 | 4,495 | 1,021 | The repository README lists the source as TBA, names Elena Klyachko as annotator and cites Kazakevich & Klyachko (2013) under references. Pimentel et al. (2021) state that the data were converted from a corpus of oral Evenki texts (Kazakevich & Klyachko, 2013), which uses IPA. Used in SIGMORPHON 2020 and 2021 (README). |
| Chukchi | UniMorph ckt @ facbd48 | 241 | 196 | 99 | Corpus-derived records from transcriptions of spoken Chukchi, Amguema variant (Amguema corpus, Chuklang), annotated in CoNLL-U by Tyers & Mishchenkova (2020) and converted to UniMorph for SIGMORPHON 2021 (Pimentel et al. 2021). Annotators: Karina Sheifer, Maria Ryskina (README). |
| Romanian | UniMorph ron @ 0910e78 CC BY-SA 3.0 | 80,262 | 4,405 | 59 | Romanian paradigms (nouns, verbs, adjectives). Source: Wikipedia; licence CC BY-SA 3.0 (repository README). |
Counts are after removing duplicate lines. A small set of hand-annotated entries (Turkish, Urdu and one Chukchi perfect) adds morpheme segmentation, which UniMorph does not provide. These were annotated by the MorphoLens author following the cited sources and have not been expert-reviewed; each is labelled by whether UniMorph stores the form.
How search works
The explorer searches the full UniMorph file for the selected language on the server. Source records are retrieved, never generated. Everything MorphoLens adds to a record (surface alignment, ambiguity, record audit, difficulty tags) is derived from the stored records and shown separately, labelled as interpretation.
- Exact match: the same spelling, ignoring only letter case, Unicode normalisation and keyboard variants that never distinguish words (Romanian ş/ș and ţ/ț, Arabic-keyboard ي/ك vs. Urdu ی/ک, apostrophe variants, zero-width characters).
- Lemma match: if the query is a citation form, its full paradigm is shown.
- Loose match: only if nothing matches exactly: diacritics and transcription details are ignored (ı/i, ă/a, ə/e, ː, Chukchi ԓ/л …). These results are labelled, because they can be different words (Romanian casă ≠ casa).
- Not attested: the page says so, reports how many triples were searched, and lists the closest attested forms (edit distance ≤ 2). Absence from UniMorph is not evidence that a word does not exist.
A form realising several bundles (syncretism) is shown once per bundle. Segmentations of UniMorph forms are automatic lemma-form alignments and are labelled as such.
Reading an analysis
Every UniMorph result is split into two cards so it is always clear what came from the resource and what MorphoLens inferred.
- Source record: the verbatim line (lemma, form, bundle), dataset and commit, the upstream source named in the repository README, the licence, and the status UniMorph record: stored in UniMorph, not independently verified by MorphoLens.
- MorphoLens interpretation: the automatic surface alignment and its confidence, any stem alternation, ambiguity with other records of the lemma, and the record audit.
Surface alignment is not morphological segmentation. UniMorph stores feature bundles, never morpheme boundaries. The alignment removes known citation endings (Turkish -mek/-mak, Urdu -nā and masculine -ā, Romanian infinitive vowels) and then reports one of four confidence levels:
- High: the citation stem occurs unchanged (Turkish kedi + ler, Romanian copil + ului).
- Medium: one stem segment alternates and the stem variant recurs in other forms of the paradigm (Romanian case + lor, casă ~ case-, also in casei and casele), or a final ending is replaced (cas + a).
- Low: the alternation is not supported by other forms, or only a longest shared stretch could be aligned (Chukchi гэ + тэйкы + ԓин).
- No split: suppletive forms that share nothing with the lemma (Romanian fă of face).
Hand segmentations exist only for the hand-annotated entries and carry the label Hand segmentation. They were made by the MorphoLens author following the cited sources and have not been expert-reviewed.
Multiple analyses is the generic case of one string stored more than once. MorphoLens separates it into syncretism (the same lemma and part of speech has the same form in different paradigm cells) and cross-category ambiguity (the same string under a different part of speech of the lemma, or under a different lemma). The distinction matters for corpus-derived data such as Evenki, where participles, converbs and finite forms of one lemma can coincide.
Difficulty tags say where they come from. Data is read from stored records (syncretism, cross-category ambiguity, variant forms, multi-word forms). Heuristic is inferred by MorphoLens: allomorphy means a stem variant with one alternating segment (casă ~ case-), stem change means lemma material is replaced or removed without a single consistent alternation. Hand comes from hand annotation, and experiment tags (unseen lemma, rare feature) only make sense relative to a train/test split.
Tag validation. Every tag atom is checked against the union of the two published UniMorph feature lists (unimorph-schema-json and um-canonicalize). Atoms missing from both are reported per dataset: frequent ones are dataset conventions and kept; rare ones that fit no template are anomalies, stored verbatim, flagged and excluded from the experiment.
| Language | Tag | Records | Treatment |
|---|---|---|---|
| Turkish | INFR | 183,456 | Convention: used systematically; kept |
| Turkish | LOC | 21,802 | Convention: used systematically; kept |
| Evenki | PSSRS | 451 | Convention: used systematically; kept |
| Evenki | PSSRP | 248 | Convention: used systematically; kept |
| Chukchi | ARBEB1S | 1 | Anomaly: fits no template; flagged, excluded from the experiment |
| Chukchi | ARBEB1P | 1 | Anomaly: fits no template; flagged, excluded from the experiment |
| Chukchi | ARBAB1S | 1 | Anomaly: fits no template; flagged, excluded from the experiment |
| Chukchi | ARBAB3S | 4 | Anomaly: fits no template; flagged, excluded from the experiment |
Several analyses for one form are all shown. None is marked primary, because MorphoLens has no corpus frequencies; they are listed by how many records in the full file carry each bundle, so Romanian caselor lists GEN/DAT · PL · DEF before VOC · PL.
Record audit. For Romanian nouns and adjectives, the Number tag is compared with unambiguous cues: indefinite articles (o, un, unei, unui are singular; niște, unor are plural) and definite endings (-lui, -ul, genitive/dative -ei and feminine -a are singular; -lor and nominative -ii are plural). For adjectives, the Gender tag is compared with endings that are unambiguous for gender (-ă, definite -a and genitive/dative -ei are feminine singular; definite -ul(ui) is masculine or neuter singular; definite -ii is masculine plural; neuter adjectives agree like masculines in the singular, so -ă is never neuter). A cue only counts when the ending is not already part of the lemma. Number: 10,171 of 22,943 cue-bearing records conflict (44%). Adjective Gender: 2,822 of 4,398 (64%). Conflicting records are flagged in the explorer and paradigm tables, shown exactly as stored, and kept out of comparisons. Records without a cue are not checked, so an unflagged record is not thereby confirmed, and other dimensions are not audited.
Experiment design
The task is morphological inflection: predict the form for a (lemma, feature bundle) pair. Metrics: exact-match accuracy and mean Levenshtein distance to the stored form. For each of 5 seeds a universe of up to ~3,000 triples is sampled (10 cells per lemma at most). The random split holds out items; the lemma-disjoint split holds out whole lemmas. Test sets hold up to 500 items; training sizes are 50, 100, 250, 500, 1000. Both splits use the same universe.
Differences between systems are tested with a paired bootstrap (2,000 resamples) over the test items pooled across seeds; we report the difference in accuracy points, its 95% interval and a two-sided p-value.
Every result is reported with two exact-match scores. Strict requires the stored form of the test record. Variant-awareaccepts any form stored for the same lemma and bundle, because oral Evenki in particular has many dialectal and transcription variants per cell: Pimentel et al. (2021) note Evenki outputs that are practically correct but belong to a different dialect.
Data preparation: records with anomalous non-schema tags are excluded (7 Chukchi records). Romanian is evaluated on verbs only, because its noun and adjective Number and Gender tags are unreliable; the same design on all Romanian records is reported separately as a raw run and is not used for claims.
Systems
- Atomic-tag rules: edit rules keyed by the whole bundle; nearest-ending analogy; unseen bundle ⇒ copy the lemma. In the spirit of the CoNLL-SIGMORPHON 2017 non-neural baseline.
- Paradigm memory: reinflects from seen forms of the same lemma (cell-to-cell rules); otherwise the baseline.
- Feature-aware rules: bundles decomposed into features; composes lemma→A with a feature-difference rule A→T learned across paradigms; otherwise nearest bundle.
- Neural, atomic tag: a small character-level Transformer encoder-decoder in PyTorch, the design of the SIGMORPHON 2020 baseline (Wu, Cotterell & Hulden, 2021) scaled down for CPU, with a copy head over the lemma characters in the style of the pointer-generator (See et al., 2017) used for low-resource inflection by Sharma et al. (2018). Tag tokens precede the lemma characters, as in Kann & Schütze (2016). Here the whole bundle is one input token; a bundle never seen in training becomes an unknown token.
- Neural, features: the same network, schedule and seeds, but each feature of the bundle is its own input token, so an unseen combination of seen features can still be represented.
All systems are trained from scratch in every run on exactly the same training items and scored on the same test items. The rule systems are transparent. The two neural systems differ only in how the bundle is presented, so their difference is the neural counterpart of the atomic vs. feature-aware rule comparison. They are small, trained on CPU with a schedule fixed in advance (no development set, no tuning, because the rule systems get no tuning data either): baselines for the evaluation design, not state of the art.
Results
Generated from the results file at the largest training size ≤ 500 per language; the Experiment page has every configuration.
- Urdu (n = 500, lemma-disjoint): feature-aware rules are +4.8 points above the atomic-tag baseline (95% CI +3.8 to +5.9, p < 0.001).
- No significant feature-aware difference on the lemma-disjoint split for Turkish (+0.2, 95% CI ±0.0 to +0.4); Evenki (−0.2, 95% CI −0.6 to +0.2); Chukchi (+1.7, 95% CI −0.3 to +3.8); Romanian (verbs) (±0.0, 95% CI ±0.0 to +0.1).
- On all Romanian records, including nouns and adjectives with unreliable Number and Gender tags, the difference is +0.4 (95% CI +0.2 to +0.7); on verbs alone it is ±0.0, so the raw figure should not be read as a morphological effect.
- Evenki: 16% of lemma-disjoint test items have more than one stored form. Variant-aware scoring raises baseline accuracy from 42.8 to 46.2; the feature-aware difference is −0.1 (95% CI −0.5 to +0.3).
- Paradigm memory beats the baseline on the random split for Turkish (+2.4, 80% of test lemmas seen in training) and Urdu (+3.7, 94% of test lemmas seen in training); on the lemma-disjoint split it equals the baseline by construction.
- Paradigm memory is significantly below the baseline on the random split for Romanian (verbs) (−6.4, 95% CI −7.6 to −5.2): there, reinflecting from another stored form of the same lemma is less accurate than inflecting from the lemma.
- It shows no significant random-split gain for Evenki (34% overlap), Chukchi (17% overlap).
- Neural models, lemma-disjoint: one input token per feature instead of one per bundle helps for Turkish (+25.6, 95% CI +24.0 to +27.4); Urdu (+22.8, 95% CI +20.5 to +25.0); Evenki (+12.8, 95% CI +11.2 to +14.3); Romanian (verbs) (+33.2, 95% CI +31.2 to +35.2).
- Neural models, lemma-disjoint: no significant difference between feature and atomic tag tokens for Chukchi (+2.4, 95% CI −1.7 to +6.2).
- Neural feature model vs. feature-aware rules on the lemma-disjoint split: Turkish −10.1, Urdu −10.9, Evenki −13.6, Chukchi −33.1, Romanian (verbs) −24.8 (positive = the network is more accurate).
- Moving from the random to the lemma-disjoint split changes the neural feature model's accuracy by −16.2 for Turkish, −14.8 for Urdu, +1.2 for Evenki, +1.4 for Chukchi, −1.9 for Romanian (verbs); the atomic-tag rules change by +3.1, −0.5, +3.2, −2.7, +0.5 (the two splits have different test items, so this is descriptive, not a paired test).
The feature-aware gains come almost entirely from test items whose feature bundle never occurred in training (see “Where the differences come from” on the Experiment page); on seen bundles the systems are nearly identical.
Provenance & data quality
Every form in the interface carries one of two source labels and opens a verbatim evidence record:
- UniMorph record: the exact triple is stored in the pinned UniMorph file. MorphoLens has not independently verified it.
- Reference pattern: a hand-annotated textbook form that UniMorph does not contain (e.g. Turkish ev, kitap).
Being stored in UniMorph is not expert validation. Problems observed in the data are reported, not silently corrected:
Romanian · Potential source-data inconsistency (Number, adjective Gender)
Several records assign Singular to plural forms and Plural to singular forms (casa and casei tagged PL; unor case and niște case tagged SG), and adjective Gender labels are also inconsistent (abandonabilă tagged NEUT). A rule-based audit of the full file finds 10,171 of 22,943 cue-bearing noun and adjective records in conflict with their Number tag, and 2,822 of 4,398 cue-bearing adjective records in conflict with their Gender tag. Other dimensions have not been audited and may also contain errors. Records are shown as stored and flagged; flagged records are kept out of comparisons, and the Romanian experiment uses verbs only.
casei → source N;GEN/DAT;PL;DEF · expected GEN/DAT · SG · DEF
Chukchi · Small, corpus-derived sample
241 records over 196 lemmas from spoken Chukchi (Amguema variant); 168 lemmas have a single record and 128 records are citation forms. The experiment uses 234 of them (7 with non-schema tags are excluded), with test sets of 58 items per seed and at most 100 training items, so Chukchi estimates have high variance.
Chukchi · Non-schema source tags
7 records use ARBEB1S, ARBEB1P, ARBAB1S, ARBAB3S, which fit no UniMorph argument-marking template (ARG + case + person/number, e.g. ARGAB3S). They are stored verbatim and flagged as likely annotation or typographical issues, never silently corrected, and excluded from the experiment.
Evenki · Dialectal and transcription variants
24% of records share their lemma and bundle with at least one other stored form, reflecting the oral IPA corpus the data was converted from (for example ə, V.PTCP;SG: 28 forms). Strict exact match counts such variants as errors, so the experiment also reports variant-aware accuracy.
Evenki · Sparse paradigms
About 2.5 forms per lemma on average (corpus-derived), so random splits rarely share lemmas between train and test.
Limitations
- The neural models are small, untuned CPU baselines with a fixed training schedule and no development set; transformers, data hallucination or pretrained models could change the picture.
- Strict exact match treats stored variants (several forms for one cell) as errors; variant-aware accuracy and edit distance are reported alongside it.
- Bootstrap intervals treat pooled test items as independent; items from different seeds can overlap.
- Surface alignments for UniMorph forms are automatic and handle one alternating stem segment at most; they are not morpheme boundaries, and exponents are never attributed to individual features.
- The record audit covers Romanian Number and adjective Gender only, and only records with an unambiguous cue; other dimensions are not audited.
- Turkish nouns and adjectives in UniMorph are Wiktionary-derived and unverified (per its README); Romanian noun labels show systematic issues.
- Evenki data is in IPA transcription from an oral corpus, sparse per lemma and rich in variants; Chukchi has 241 corpus-derived records in UniMorph, mostly citation forms; 234 are usable in the experiment after excluding 7 with non-schema tags, its test sets hold 58 items per seed, and training stops at n = 100.
- Urdu romanisation in hand-annotated entries is simplified; Nastaliq rendering depends on the font.
Downloads & citation
Experiment results (JSON)
All cells, per-seed runs, bootstrap tests, breakdowns, examples
Experiment results (CSV)
One row per language × split × n × system
Explorer sample (JSON)
Featured lemmas and comparison examples, verbatim UniMorph triples
@misc{morpholens,
title = {MorphoLens: Exploring Morphological Generalisation to Unseen Lemmas},
howpublished = {\url{https://morpholens-app.vercel.app}},
note = {Version 0.6. Data: UniMorph tur, urd, evn, ckt, ron at pinned commits},
year = {2026}
}Please also cite UniMorph (Batsuren et al., 2022) and the per-language sources listed above.
References
- Batsuren, K. et al. (2022). UniMorph 4.0: Universal Morphology. Proceedings of LREC 2022.
- Cotterell, R. et al. (2017). CoNLL-SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection in 52 Languages.
- Goldman, O., Guriel, D. & Tsarfaty, R. (2022). (Un)solving Morphological Inflection: Lemma Overlap Artificially Inflates Models’ Performance. Proceedings of ACL 2022.
- Pimentel, T., Ryskina, M. et al. (2021). SIGMORPHON 2021 Shared Task on Morphological Reinflection: Generalization Across Languages.
- McCarthy, A. D. et al. (2018). Marrying Universal Dependencies and Universal Morphology. UDW 2018.
- Kazakevich, O. A. & Klyachko, E. L. (2013). Создание мультимедийного аннотированного корпуса текстов как исследовательская процедура [Creating a multimedia annotated text corpus as a research procedure].
- Dunn, M. (1999). A Grammar of Chukchi. PhD thesis, Australian National University.
- Vylomova, E. et al. (2020). SIGMORPHON 2020 Shared Task 0: Typologically Diverse Morphological Inflection.
- Tyers, F. & Mishchenkova, K. (2020). Dependency annotation of noun incorporation in polysynthetic languages. UDW 2020.
- Kann, K. & Schütze, H. (2016). Single-Model Encoder-Decoder with Explicit Morphological Representation for Reinflection. Proceedings of ACL 2016.
- Wu, S., Cotterell, R. & Hulden, M. (2021). Applying the Transformer to Character-level Transduction. Proceedings of EACL 2021.
- See, A., Liu, P. J. & Manning, C. D. (2017). Get To The Point: Summarization with Pointer-Generator Networks. Proceedings of ACL 2017.
- Sharma, A., Katrapati, G. & Sharma, D. M. (2018). IIT(BHU)-IIITH at CoNLL-SIGMORPHON 2018 Shared Task on Universal Morphological Reinflection.
- Göksel, A. & Kerslake, C. (2005). Turkish: A Comprehensive Grammar. Routledge.
- Schmidt, R. L. (1999). Urdu: An Essential Grammar. Routledge.
Reproduce
npm install npm run data:fetch # UniMorph files at pinned commits → unimorph-data/ npm run data:build # explorer + comparison sample + record audit npm run experiment:splits # exact train/test splits → unimorph-data/splits/ npm run experiment:neural # PyTorch (CPU), 2 neural systems → experiments/neural/ npm run experiment # 5 languages × 2 splits × 5 sizes × 5 seeds (+ bootstrap) npm test # regression checks: alignment, audit, matching npm run dev

