Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Overlapping Speech in CHAT

Status: Current Last updated: 2026-05-21 08:39 EDT

Two Encodings for Overlapping Speech

When two speakers talk at the same time, CHAT supports two ways to represent this in the transcript. Both are valid; which one you use depends on your transcription conventions and analysis needs.

&*: Embedded Overlap Marker

The &* marker embeds one speaker’s words inside another speaker’s utterance. The syntax is &*SPEAKER:word or &*SPEAKER:word_word (underscores join compound expressions because &* only allows a single token).

*PAR:	I went to the store &*INV:mhm and bought some milk . 0_6000

Here, INV said “mhm” while PAR was talking. The &*INV:mhm is placed at the approximate position in PAR’s text where the overlap occurred.

Properties:

  • INV’s backchannel has no timing of its own: it is subsumed by PAR’s bullet.
  • INV’s backchannel has no %mor, %gra, or %wor: it is invisible to all dependent tiers and alignment.
  • INV’s backchannel cannot be counted as an independent utterance by analysis tools (FREQ, MLU, etc.).
  • Multi-word overlaps use underscores: &*INV:oh_okay_yeah.

Corpus scale: ~35,000 &* markers across ~2,200 files in 8 corpora.

Each speaker’s words go on their own line. The +< (lazy overlap) linker marks that the utterance started before the previous one finished:

*PAR:	I went to the store and bought some milk . 0_6000
*INV:	+< mhm . 3500_4000

Properties:

  • INV’s backchannel gets its own timing from the aligner.
  • INV’s backchannel can receive its own %mor and %wor tiers.
  • INV’s backchannel is a separate utterance, countable by analysis tools.
  • PAR’s utterance stays intact, the participant’s thought is one unit.
  • Cross-speaker overlap is valid CHAT (E701 only requires non-decreasing start times).

Corpus scale: ~327,000 +< utterances across ~15,600 files in 14 corpora.

Which Should I Use?

For new transcription: Use +< with separate utterances. Each speaker’s words belong on their own tier. This gives backchannels their own timing, their own dependent tiers, and makes them countable by analysis tools. The two-pass overlap-aware alignment strategy is available via --utr-strategy two-pass; the default --utr-strategy auto currently falls back to single-pass GlobalUtr until the two-pass algorithm is validated against more operator corpora (per crates/batchalign/src/runner/dispatch/utr.rs:94-100).

For existing files with &*: They work fine as-is. The aligner already handles &* correctly (it is invisible to the DP alignment). No migration is required. However, backchannels encoded as &* will never get independent timing, they are invisible to the aligner by design.

Both encodings are valid CHAT. The aligner supports both. Files with &* and files with +< can coexist in the same corpus.

Summary of tradeoffs:

&* encoding+< separate utterances
Backchannel timingNone (invisible to aligner)Automatic (two-pass recovery)
Backchannel %mor/%worNoneYes (own tiers)
Countable as utteranceNoYes
Main speaker alignmentUnaffectedUnaffected
Readability at densityPoor (3+ &* per line)Clean
Requires migrationNo (existing files work)New convention for new transcripts

How the Aligner Handles Each Encoding

&* encoding

The content walker skips OtherSpokenEvent (&*) nodes entirely. They do not participate in UTR word extraction, forced alignment, or %wor generation. The backchannel words are invisible to the DP reference sequence.

Result: The main speaker’s alignment is unaffected. The backchannel gets no independent timing.

+< encoding

A two-pass UTR strategy is available for +< files. The algorithm:

  1. Pass 1: Build the global alignment reference from non-+< utterances only. Main-speaker words align correctly without backchannel interference.
  2. Pass 2: For each +< utterance, search the previous utterance’s audio window with adaptive widening to recover the backchannel’s timing.
  3. Fallback: If two-pass timed fewer utterances than the standard global algorithm would have, the global results are used instead. This ensures the strategy is never worse than the original algorithm, important for languages where ASR quality is lower.

Today the strategy is opt-in. Per crates/batchalign/src/runner/dispatch/utr.rs:94-100, the default --utr-strategy auto always resolves to GlobalUtr (“Auto always uses GlobalUtr until the two-pass algorithm is validated on an operator’s problem files”); the two-pass path is only constructed when the operator passes --utr-strategy two-pass. Files without +< use the original single-pass algorithm regardless.

# Default: auto (currently resolves to global single-pass)
batchalign3 align corpus/ -o output/

# Explicit global single-pass
batchalign3 align corpus/ -o output/ --utr-strategy global

# Opt in to the two-pass overlap-aware strategy
batchalign3 align corpus/ -o output/ --utr-strategy two-pass

Multi-Backchannel Example

The &* encoding becomes hard to read with multiple backchannels:

*PAR:	but I grew up in Princeton &*INV:oh_okay_yeah and came to
	graduate school &*INV:mhm at Chapel_Hill &*INV:oh in ninety
	one &*INV:mhm or maybe ninety two . 104745_118254

The +< encoding is cleaner:

*PAR:	but I grew up in Princeton and came to graduate school at
	Chapel_Hill and in ninety one or maybe ninety two . 104745_118254
*INV:	+< oh okay yeah .
*INV:	+< mhm .
*INV:	+< oh .
*INV:	+< mhm .

Each backchannel is a separate conversational act on its own line. PAR’s narrative is one unbroken utterance.

CHAT Validation Rules

Cross-speaker overlapping bullets are valid:

  • E701 (global timeline): Start times must be non-decreasing. PAR at 104745 ≤ INV at any time after that. Passes.
  • E704 (same-speaker self-overlap): Only prohibits the same speaker overlapping themselves beyond 500ms. Different speakers can overlap freely.
  • E362 (monotonicity): Same as E701. Passes.

References


This page last changed: 2026-06-19 (commit c82a6d03). The whole book last changed: 2026-09-16 (commit 34d249d8).