Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

whisper_hub ASR engine

Status: Current Last updated: 2026-09-15 12:12 EDT

What it is

whisper_hub is an ASR engine variant that loads a community Whisper fine-tune from the Hugging Face Hub by model_id. It exists because stock openai/whisper-* checkpoints produce unusable output on some languages where Rev.AI also fails, and a per-language fine-tune is the only path to coherent transcription.

whisper_hub is parallel to the other ASR engine variants:

EngineWhat it loadsWhen to use
rev (default)Rev.AI cloud APILanguages where Rev.AI quality is good (English, Spanish, most European).
whisperStock openai/whisper-large-v3 via HF transformersLanguages where stock Whisper handles the acoustic + language combo well.
whisper_rswhisper.cpp in-process (Rust-native)Local transcription with no Python worker.
whisperx, whisper_oaiNothing: not implementedNever. Both are accepted names with no implementation and are refused at submission.
whisper_hubHF community fine-tune by model_idLanguages where both Rev.AI and stock Whisper fail.
tencent, aliyun, funaudioCantonese providersChinese variants only.

Quick start

# Uses the per-language default model_id resolved from the Rust manifest,
# crates/batchalign/src/model_manifest.rs. For Malayalam that's
# thennal/whisper-medium-ml, at the commit the manifest pins.
batchalign3 transcribe input/ output/ --lang mal --asr-engine whisper_hub

To override the model for a language that already has a default, or to pick a model for a language we haven’t seeded yet:

batchalign3 transcribe input/ output/ \
  --lang mal --asr-engine whisper_hub \
  --asr-engine whisper_hub --engine-overrides '{"model_id": "other/mal-model"}'

Per-language defaults

The per-language default model_id table lives in crates/batchalign/src/model_manifest.rs::WHISPER_HUB_DEFAULTS, together with the exact hub commit each default is pinned to. It is intentionally small and seeded reactively from empirical evaluation, a language only gets a default after we’ve confirmed the chosen fine-tune produces coherent output.

Language (ISO-639-3)Default HF model_idNotes
mal (Malayalam)thennal/whisper-medium-mlSee “Evaluation below.”

Rust owns that table because an id has to be known BEFORE a load in order to pin its revision, and a second copy on the worker side could only disagree with it. A planned job therefore resolves from the manifest, and an entry added only to batchalign/models/resolve.py leaves the job refused: planning fails with ModelPlanError::WhisperHubHasNoDefaultModel, naming the language, rather than loading something unpinned. The Python table and its WhisperHubModelNotFoundError remain for direct callers that have no control plane.

Any other language requires passing --asr-engine whisper_hub --engine-overrides '{"model_id":"..."}', instead of falling back to a stock Whisper checkpoint that would silently produce garbage.

What the run records about identity

An id the manifest pins is loaded at that exact commit, and the transcript’s stamp names the model with its revision. An id the manifest does not know, which is any model_id passed through --engine-overrides that is not a seeded default, is carried as a FLOATING identity: it loads, and the revision recorded for it is whatever the worker reports for the weights it actually resolved, so an override never leaves the transcript naming the engine alone. A floating identity can never build a cache key.

Why a per-language table and not auto-discovery?

HuggingFace lists dozens of Whisper fine-tunes per language. Their advertised WER is self-reported and wildly inconsistent, their test sets vary, and some checkpoints (e.g., DrishtiSharma/whisper-large-v2-malayalam) have a broken generation_config that refuses to load via HF transformers without a consumer-side workaround. Auto-picking by download count or name match would ship the first plausible-looking thing to users with no quality signal.

The table is hand-curated so every default is traceable to an actual empirical comparison. The escalation path to automated probing lives in revai-language-quality-strategy.md.

Evaluation behind the Malayalam seed

A 73-second Malayalam sample was transcribed by four candidates:

ModelMalayalam-script charsRepetitionUsable
thennal/whisper-medium-ml100.0 %0.01Yes
kavyamanohar/whisper-small-malayalam100.0 %0.04Yes, noisier
openai/whisper-large-v327.5 %0.27No (hallucinates “Thank you for watching.”)
openai/whisper-medium13.4 %0.73No (Khmer + Gurmukhi character loops)

Rev.AI on the same file returned 55 tokens of Hangul + Gurmukhi + Latin

  • U+FFFD, zero Malayalam script. That result drove the deny-list entry in revai/preflight.rs::REVAI_KNOWN_BROKEN, which now recommends whisper_hub for Malayalam specifically.

Artifacts live in an operational workspace outside this public repo.

Caveats:

  • Single audio file. A native Malayalam reader should compare word-level accuracy before shipping this default for a large corpus.
  • CPU inference (no GPU numbers). thennal/whisper-medium-ml took 353 seconds on a development machine’s CPU for 73 seconds of audio (4.8× real-time slower). GPU should be ~3-5× faster than real-time.

Fine-tune gotchas

HF Whisper fine-tunes differ from stock OpenAI checkpoints in one critical way: they bake language and task into their own generation_config. Passing them again via generate_kwargs produces gibberish, the model applies two competing prompts.

whisper_hub handles this by passing language="auto" through to the shared load_whisper_asr(), which makes the handle’s gen_kwargs("auto") branch fire and omit those overrides. Do not replicate the language=<concrete> path used by stock Whisper when adding a new fine-tune loader.

When this engine stops being a fit

If the per-language table grows past ~10 entries and we’re re-testing defaults often, it’s time to escalate to a probe harness, see Option C in revai-language-quality-strategy.md.

Cross-references


This page last changed: 2026-09-16 (commit 197c81e6). The whole book last changed: 2026-09-16 (commit 34d249d8).