Morphotag Reconciliation Invariants
Status: Current Last updated: 2026-08-30 19:35 EDT
This page documents the 1-to-1 invariant that the morphotag pipeline relies on, the three stages that together make it hold deterministically, the two legitimate modes that intentionally skip it, and the typed outcome model that replaces the old silent-skip pattern.
The invariant
For every CHAT utterance the pipeline visits:
Post-mapping, the number of
%moritems equals the number of Mor-alignable words on the main tier.
“Mor-alignable” is defined by
counts_for_tier(word, TierDomain::Mor)
in talkbank-model. It is the authoritative CHAT policy: regular words
count, replacement words count, tag-marker separators (comma ,,
tag „, vocative ‡) count; fillers (&-hmm), nonwords (&~uh),
phonological fragments (&+le), untranscribed material (xxx, yyy,
www), omissions, retrace content, and utterance terminators do not.
The canonical count is available as
Utterance::mor_alignable_word_count() on the
talkbank_model::model::Utterance type. Any pipeline stage that
validates “dependent tier count matches main tier content” must call
this method. Duplicating the walk locally risks drift from CHAT
policy and from sibling implementations: two copies of a word-counting
rule are two places the rule can silently drift out of sync.
Why this can hold by construction
Three independent pipeline stages cooperate to make the invariant
deterministic. When each does its job, |mors| == N without any
post-hoc alignment.
flowchart TD
U["Utterance U<br/>(main tier)"] --> E
E["Stage 1: Extract<br/>collect_utterance_content(Mor)<br/>(../chatter/crates/talkbank-transform/src/extract.rs)"] --> Nlabel
Nlabel{{"N Mor-alignable words<br/>(counts_for_tier rule)"}}
Nlabel -->|"N == 0"| NA["MorOutcome::NotApplicable<br/>no %mor produced (correct)"]
Nlabel -->|"N > 0"| D
D["Stage 2: Dispatch<br/>tok_ctx.original_words = word_lists<br/>(_realignment_applied in batchalign/inference/morphosyntax.py)"] --> S
S["Stanza neural tokenizer<br/>re-tokenizes text, realigning<br/>to word_lists boundaries"] --> P
P["Stage 3: Project<br/>map_ud_sentence + MWT Range reassembly<br/>(crates/batchalign-transform/src/morphosyntax/sentence_mapping.rs)"] --> M
M{{"|mors| =?= N"}}
M -->|"yes"| OK["MorOutcome::Aligned<br/>inject %mor + %gra"]
M -->|"no"| BUG["MorOutcome::MisalignmentBug<br/>typed diagnostic; loud, investigable"]
Stage 1 produces N from the CHAT side. Stage 2 tells Stanza “use
these N word boundaries when you re-tokenize the combined text.” Stage 3
takes Stanza’s UD output and reassembles MWT ranges (don't → do + n't)
back into one %mor chunk per CHAT word.
Each stage’s correctness is independently testable:
- Stage 1: the parity test at
crates/batchalign/tests/chat_ops_mor_count_parity_reference_corpus.rsasserts thatUtterance::mor_alignable_word_count()andextract::collect_utterance_content(..., Mor, ...).len()agree on every utterance in the 98-file reference corpus. - Stage 2: the contract test at
batchalign/tests/inference/test_morphosyntax_realignment_contract.pyasserts thattok_ctx.original_words = word_listshappens before everynlp()call in normal mode, and is empty under--retokenize. - Stage 3: unit tests in
nlp/mapping/mod.rsexercise MWT reassembly across French (du → de + le), English (don't → do + n't), German (im → in + dem), Italian, Portuguese, Dutch, and a comma regression test that locks the mid-utterance-comma handling.
When a count mismatch nevertheless surfaces at the injection boundary, it is always a bug in one of those three stages: never an expected divergence class to be resolved by after-the-fact alignment.
The two legitimate non-realignment modes
Two documented modes intentionally skip the realignment step because they want Stanza to own tokenization:
| Mode | Trigger | Why skip realignment |
|---|---|---|
| CJK retokenize | Mandarin (zho/cmn) with retokenize=True; use_retok_pipeline=True in morphosyntax.py | Chinese has no whitespace word boundaries; Stanza’s neural segmenter produces the correct word boundaries for Chinese, and CHAT’s main tier is rewritten to match. |
| Generic retokenize | Any language with --retokenize CLI flag; req.retokenize=True | The user has asked the pipeline to expand MWTs (gonna → going + to) and rewrite the CHAT main tier accordingly. Stanza must own tokenization for that. |
In both modes, the 1-to-1 invariant is not violated, it simply
operates at the Stanza-token level instead of the CHAT-word level.
The main tier gets rewritten to match Stanza’s output, and the
resulting %mor count equals the rewritten word count by construction.
(The current implementation of this rewrite lives in
crates/batchalign-transform/src/retokenize.rs plus the
crates/batchalign-transform/src/retokenize/ sub-modules.)
MorOutcome: the typed outcome vocabulary
Every utterance the pipeline visits produces exactly one
MorOutcome
with one of three kinds. This replaces the previous silent-skip
behavior that previously let an upstream regression mask itself as
silent %mor loss.
| Kind | Meaning | Surfaces as |
|---|---|---|
NotApplicable { reason } | The utterance had zero Mor-alignable words. No %mor is produced, and that is correct. Reasons: FillerOnly, FragmentOnly, NonwordOnly, UntranscribedOnly, AllRetraced, MixedNonLinguistic, Empty. | Typed morphosyntax:not_applicable record. |
Aligned { n_words } | N CHAT words, N %mor items; happy path. | No anomaly record. |
MisalignmentBug(diag) | ` | mors |
No review_level value writes %xalign or %xrev. The typed outcomes exist
and anomaly records are traced, but morphotag does not yet persist the collected
records in a per-file evidence sidecar. See
Decision Evidence for the current sink
matrix and the Review Tiers guide for the
CHAT policy.
The MisalignmentClass classifier (best-effort) points developers at
the most likely failing stage:
RealignmentSkipped: Stanza’s tokenizer-realignment context wasNonefor the dispatch language; Stanza ran without boundary hints.MwtReassemblyBug: the UD→Mor projection consumed the wrong number of tokens during MWT Range expansion.TerminatorFilterBug:is_terminator_punctdropped too many or too few PUNCT tokens, e.g. dropping mid-utterancecm|cmseparators that CHAT counts as alignable.LanguageDispatchIssue: per-language chunk of a code-switched utterance disagreed with the CHAT main-tier count.Unknown: diagnostic alone is insufficient; developer must inspect.
What this architecture explicitly does not do
- It does not add a DP alignment layer between Stanza output and
CHAT words. That would paper over bugs in the three stages rather
than fix them. Batchalign2 did this (character-level DP inside
Stanza’s
tokenize_postprocessor) and still silently skipped residual mismatches; the new architecture replaces both halves with a built-in Stanza realignment + typed outcomes. - It does not change the CHAT Mor-alignable policy. If a future
CHAT-manual-approved decision includes fillers in
%mor, that is a change tocounts_for_tierintalkbank-model, propagated throughmor_alignable_word_count()to every caller. Pipeline-internal heuristics are banned. - It does not reduce observability on the happy path.
Alignedoutcomes do not produce decision tiers by default, morphotag users see exactly the same output as before for every successful utterance.
See also
crates/batchalign-transform/src/morphosyntax/outcome.rs: outcome typescrates/batchalign-transform/src/inject.rs: invariant check & outcome emissioncrates/batchalign-transform/src/morphosyntax/payload.rs: NotApplicable classificationbatchalign/inference/morphosyntax.py: realignment stagebatchalign/tests/inference/test_morphosyntax_realignment_contract.py: Stage 2 testscrates/batchalign/tests/chat_ops_mor_count_parity_reference_corpus.rs: Stage 1 teststalkbank-tools/../chatter/crates/talkbank-model/src/alignment/helpers/rules.rs:counts_for_tier
This page last changed: 2026-09-16 (commit 197c81e6). The whole book last changed: 2026-09-16 (commit 34d249d8).