Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Memory Safety: Preventing Kernel OOM Crashes

Status: Current Last updated: 2026-09-05 03:20 EDT

The Problem

Each Python ML worker loads 2-15 GB of models (Whisper, Stanza, etc.). When multiple workers spawn concurrently, from parallel test binaries, per-job pre-scaling, or job dispatch, they collectively exceed physical RAM and trigger a kernel-level OOM panic that crashes the entire machine. This is not a process-level OOM kill; it is a Jetsam-triggered kernel panic that requires a hard reboot.

Sample observed failure modes:

  • 5 python3.12 workers × ~13-15 GB each = ~70 GB on a 64 GB machine
  • A multi-file transcription run with the default auto-tuner exhausted the machine’s 64 GB

Architecture

flowchart TD
    A[Test / Job / Pre-scale<br>wants capacity] --> B{Host memory coordinator}
    B -->|startup lease| C[One worker/model startup window]
    B -->|job execution lease| D[Clamp per-job file parallelism]
    B -->|ML test lock| E[Allow one real-model test owner]
    C --> F{Local spawn guard}
    F -->|permit + local RAM check| G[Spawn Python worker]
    G --> H[Worker loads ML models<br>2-15 GB]
    H --> I[Drop startup lease<br>keep only job/runtime leases]

    style B fill:#ff9,stroke:#cc0,color:#000
    style G fill:#9f9,stroke:#0c0,color:#000
    style E fill:#f99,stroke:#c00,color:#000

Defense Layers

Layer 0: Live admission gates at the worker-pool spawn seam

worker/pool/{cpu_gate,memory_gate}.rs are the cheapest, earliest admission predicates. They run inside try_claim_spawn_slot before any permit acquisition or lease reservation, so a saturated host doesn’t burn through the global-permit semaphore on doomed spawns.

The gates run two distinct policies, named explicitly per PoolGateState (derived in worker/pool/lifecycle.rs):

  • ColdStart: first worker for a (profile, lang, engine) class. Both gates bypass unconditionally: back-pressure has nothing to push against on an empty pool, and refusing here leaves the pool dead-on-arrival on memory-tight hosts (the laptop-class failure mode that motivated the split).
  • Warm: N+1 worker for a class with existing workers. Both gates run their projection.
GateWarm predicateSource
CPU loadavggetloadavg(3).one < available_parallelism()cpu_gate.rs
Memory floor + projectionavailable_mb − new_worker_estimate_mb > host_min_free_mb_threshold_for_tier(tier)memory_gate.rs

The Warm-gate floor is tier-scaled: 1024 MB Small / 2048 MB Medium / 4096 MB Large / 4096 MB Fleet. The historical fixed MIN_FREE_MEMORY_MB = 2048 constant is now the Medium-tier value and the JobStore-level gate’s default; it is no longer the single admission floor.

The new_worker_estimate_mb is the average RSS of same-profile idle peers (Mode B, rss_observer.rs) when peers exist; otherwise falls back to the canonical per-tier MemoryTier::*_startup_mb (Mode A fallback), tier.gpu_startup_mb, tier.stanza_startup_mb, tier.io_startup_mb. By design (Principle 1), MemoryTier is the sole canonical source of these values; an architectural-invariant test in types/runtime.rs prevents reintroduction of parallel constants.

The Mode-A fallback is further engine-aware for the IO profile, via memory_gate::engine_aware_startup_reservation_mb. The IO baseline (tier.io_startup_mb: 2 GB Small/Medium, 4 GB Large/Fleet) is correct for engines that are thin API clients (Google Translate via googletrans) but under-reserves for engines that load large local models in the worker process, SeamlessM4T (~2.4 GB resident) and NLLB-200-distilled-1.3B (~5 GB resident). The helper takes the MAX of the profile baseline and the engine’s resident footprint as declared by TranslateEngineName::resident_memory_mb, so the admission gate refuses to spawn an NLLB worker on a Medium-tier host that doesn’t have ~5 GB of headroom. Workers running the lightweight Google engine continue to pay the IO baseline.

Admission is back-pressure, not safety (Principle 5). The correctness floor is worker/memory_guard.rs (per-spawn host-memory reservation + RSS observation + kill on overrun) plus the OS OOM killer. An over-permissive admission means a worker may die at spawn, bounded cost. An over-strict admission means the host can’t run at all, unbounded cost (jobs queue forever). The bias is toward over-permissive; ColdStart bypass implements the bias.

The eviction-side counterpart: worker/pool/idle_eviction.rs runs as a pre-pass in run_health_check and evicts idle workers largest-RSS first when available_mb falls at or below EVICTION_PRESSURE_THRESHOLD_MB = 4096 MB (= 2× the admission floor). There is no idle_timeout_s knob, eviction is purely pressure-driven.

The host’s available-memory reading is shared across all five sysinfo-touching paths (admission gate, eviction pre-pass, in-spawn guard, host-facts probes, info logs) via the TTL-cached host_memory::system_memory_snapshot: at most one /proc/meminfo (Linux) / host_statistics64 (macOS) read per second across the whole pool.

Layer 1: Host-wide coordinator (prevents cross-process overcommit)

A machine-local JSON ledger guarded by an exclusive file lock coordinates memory across local batchalign3 processes on the same host. This covers:

  • multiple server ports,
  • CLI auto-daemons,
  • pre-scaled vs foreground jobs,
  • independent Rust test binaries.

The coordinator tracks three lease types:

  • worker startup leases for the model-loading spike,
  • job execution leases for in-flight file parallelism,
  • machine-wide ML test locks so real-model test runs do not stampede the host.

The reserve/headroom policy comes from ServerConfig.memory_gate_mb, which now means “keep at least this much RAM free after reservations” rather than a standalone job gate. The default is MIN_FREE_MEMORY_MB = 2048 (the JobStore-gate constant). The previous tier-derived per-host default (Small=2 GB, Medium=4 GB, Large/Fleet=8 GB) was retired on 2026-05-08 for this knob; workload-sized headroom now comes from Layer 0’s per-process RSS observation. Note: this is independent of the Layer-0 Warm-gate floor host_min_free_mb_threshold_for_tier, which is tier-scaled (1024/2048/4096/4096) and protects the per-spawn admission decision rather than the per-job admission decision.

Layer 2: Spawn semaphore (prevents in-process TOCTOU race)

A process-global tokio::sync::Semaphore serializes all worker spawns. This still matters even with the host-wide coordinator because one server process can otherwise race with itself:

sequenceDiagram
    participant T1 as Test Binary 1
    participant T2 as Test Binary 2
    participant Sem as Spawn Semaphore
    participant Mem as Memory Check
    participant Py as Python Worker

    Note over T1,T2: WITHOUT semaphore (old behavior - CRASHES)
    T1->>Mem: Check: 20 GB free ✓
    T2->>Mem: Check: 20 GB free ✓
    T1->>Py: Spawn worker (loads 15 GB)
    T2->>Py: Spawn worker (loads 15 GB)
    Note over Py: 30 GB loaded on 20 GB free → OOM PANIC

    Note over T1,T2: WITH semaphore (new behavior - SAFE)
    T1->>Sem: Acquire permit
    Sem-->>T1: Granted
    T1->>Mem: Check: 20 GB free ✓
    T1->>Py: Spawn worker (loads 15 GB)
    T2->>Sem: Acquire permit
    Sem-->>T2: Wait (T1 holds permit)...
    T1->>Sem: Release permit
    Sem-->>T2: Granted
    T2->>Mem: Check: 5 GB free ✗
    Note over T2: MemoryGuardError returned, no spawn

Location: crates/batchalign/src/worker/memory_guard.rs

Layer 3: Explicit startup reservations (before every worker spawn)

Worker startup budgets are now tier-adaptive, scaled by a MemoryTier derived from total system RAM:

TierTotal RAMGPU StartupStanza StartupIO StartupHeadroom
Small< 24 GB6 GB3 GB2 GB2 GB
Medium24-48 GB3 GB (LazyProfile)6 GB3 GB4 GB
Large48-128 GB16 GB12 GB4 GB8 GB
Fleet≥ 128 GB16 GB12 GB4 GB8 GB

These values are defined exactly once in MemoryTier::from_total_mb() in crates/batchalign/src/types/runtime.rs , the sole canonical source per Principle 1. Operator overrides flow in via RuntimeOverridesConfig.{gpu,stanza,io}_startup_mb, which override the tier-derived values. The Medium tier uses LazyProfile for GPU: the worker starts with only process overhead and loads model weights on demand. A lazy worker key still includes the selected engine recipe; two engines cannot share one task-only process. The host permit gate and idle-worker eviction bound how many recipe-specific workers remain resident. This trades a small amount of process overhead for a hard correctness guarantee: an already-loaded model cannot silently satisfy a request for another engine. The startup reservations are intentionally more conservative than the per-command execution budgets. They protect the model-loading spike where Whisper, Stanza, or related engines can temporarily consume far more memory than steady-state request handling.

Note: On macOS, sysinfo::available_memory() undercounts because it only reports free + purgeable pages, not inactive pages. The kernel can reclaim inactive pages, so the real headroom is larger. We use the conservative number.

Layer 4: Job execution reservations (before a job starts running)

The runner no longer uses a separate memory_gate() plus independent memory-based auto-tune formula. Instead it:

  1. computes a requested worker count from file count, CPU, and category caps,
  2. asks the host coordinator for a job execution plan,
  3. receives a granted worker count plus a lease held for the job lifetime,
  4. re-queues the job if the host cannot safely fit that plan.

This makes worker startup and job execution share one memory story instead of two unrelated heuristics.

Layer 5: Machine-wide ML test lock

The live ML fixture now acquires a machine-wide test lock before preparing warm workers. This prevents concurrent cargo test or IDE runs from each building their own model pool on the same machine.

This lock complements, rather than replaces:

  • RUST_TEST_THREADS=1,
  • the ML golden suite,
  • the single-binary ml_golden layout.

Layer 6: SIGKILL Follow-Through in Drop

Both WorkerHandle::Drop and SharedGpuWorker::Drop now send SIGTERM, wait 200ms, then send SIGKILL if the worker is still alive. This prevents zombie Python processes when the worker is stuck in a C extension (PyTorch, NumPy) that ignores SIGTERM.

Layer 7: Periodic Orphan Reaping

The health check background task now calls reap_orphaned_workers() on every tick (default: 30s). This catches orphaned workers from server crashes without waiting for the next server restart. Previously, orphans only got cleaned up when a new server instance started.

Layer 8: Test-level skip (bail out before any setup)

Every test file that spawns workers has a require_python!() macro that checks available memory BEFORE attempting to spawn:

#![allow(unused)]
fn main() {
macro_rules! require_python {
    () => {{
        let available_mb = batchalign::worker::memory_guard::available_memory_mb();
        if available_mb < 4096 {
            eprintln!("SKIP: insufficient memory ({available_mb} MB)");
            return;
        }
        // ... resolve python path ...
    }};
}
}

Layer 9: Test isolation and explicit ML opt-in

The default Rust suite includes test-echo integration tests, which may spawn lightweight Python workers but do not load ML models. Related tests share a small number of executables and warmed fixtures. ML model tests are gated behind their own per-host opt-in feature and must only be run on a Fleet/Large-tier host with ≥ 256 GB RAM:

make test                                            # Default suite, no ML models
cargo test -p batchalign --test worker_integration_suite worker_integration:: -- --test-threads=1
# ML golden tests: Fleet/Large-tier hosts only

Environment Variables

VariableDefaultWhat it does
BATCHALIGN_SPAWN_MIN_MEMORY_MB4096Minimum free RAM (MB) to allow a worker spawn
BATCHALIGN_MAX_CONCURRENT_SPAWNS1Max concurrent worker spawns (semaphore size)
BATCHALIGN_HOST_MEMORY_LEDGERtemp-dir pathOverride the shared host-memory ledger path
RUST_TEST_THREADSRust harness defaultOptional cap on parallel test functions; pass --test-threads=1 for isolated worker probes

Key config knobs

SettingDefaultWhat it does
memory_gate_mb2048 MBHost reserve/headroom preserved after reservations. Same constant the worker-pool admission gate enforces; operator override accepted but rarely needed.
max_concurrent_worker_startups1Host-wide limit for simultaneous worker/model startups
gpu_thread_pool_size4In-process GPU request concurrency, now forwarded into Python

How to Run Tests Safely

On a developer machine (≤ 64 GB)

# Default suite: Rust plus test-echo workers, no ML models
make test

# Worker integration tests (test-echo mode, no ML models): safe with
# the memory guard; spawn real Python workers in test-echo mode
# without model loading
cargo test -p batchalign --test worker_integration_suite worker_integration:: -- --test-threads=1

# NEVER run ML golden tests on a 64 GB machine; they will OOM.

On a Fleet/Large-tier host (≥ 256 GB RAM, e.g. an M3 Ultra Mac Studio)

# Default suite + focused worker integration + ML golden tests
make test
cargo test -p batchalign --test worker_integration_suite worker_integration:: -- --test-threads=1
cargo test -p batchalign --test ml_golden -- --test-threads=1

Running a specific integration test

# One module within the shared worker suite, single thread, memory guard active
cargo test -p batchalign --test worker_integration_suite worker_integration:: -- --test-threads=1

# Run only ignored tests (if any)
cargo test -p batchalign --test worker_integration_suite worker_integration:: -- --ignored --test-threads=1

What requires explicit opt-in

# Loads real ML models and may use hosted services; large hosts only
cargo test -p batchalign --features ml-golden --test ml_golden -- --test-threads=1

Plain Cargo runs integration executables sequentially; each executable’s Rust test harness may run its test functions concurrently. The shared fixtures and memory gates make the default suite safe without depending on a third-party runner. Use a focused module filter and --test-threads=1 when diagnosing ordering or resource-admission behavior.

Implementation Files

FileWhat
crates/batchalign/src/worker/pool/cpu_gate.rsLayer 0 admission: getloadavg(3) vs available_parallelism()
crates/batchalign/src/worker/pool/memory_gate.rsLayer 0 admission: available − reservation > host_min_free_mb_threshold_for_tier(tier) (Warm); ColdStart bypasses. Tier-scaled floor: 1024/2048/4096/4096 by tier
crates/batchalign/src/worker/pool/rss_observer.rsPer-process RSS sampling for the Mode B admission estimate
crates/batchalign/src/worker/pool/idle_eviction.rsPressure-driven idle-worker eviction (largest-RSS first when available <= 4096 MB)
crates/batchalign/src/host_memory.rsHost-wide ledger, startup leases, job execution leases, ML test lock; TTL-cached system_memory_snapshot shared by every memory poll
crates/batchalign/src/worker/memory_guard.rsLocal spawn semaphore plus host-memory startup reservation
crates/batchalign/src/worker/handle/mod.rs and crates/batchalign/src/worker/handle/spawn.rsWorkerHandle::spawn() (mod.rs:73) and spawn_tcp_daemon() (spawn.rs:135) both call acquire_spawn_permit()
crates/batchalign/src/runner/mod.rsCoordinator-backed job execution planning and requeue
crates/batchalign/tests/common/mod.rsMachine-wide ML fixture lock
crates/batchalign/tests/worker_integration.rsrequire_python! macro with memory check
crates/batchalign/tests/gpu_concurrent_dispatch.rsSame
crates/batchalign/tests/worker_protocol_matrix.rsSame
.cargo/config.tomlRUST_TEST_THREADS = "1"
MakefileTiered test targets: test-rust, test-workers, test-ml

This page last changed: 2026-09-05 (commit 7f5e86f1). The whole book last changed: 2026-09-16 (commit 34d249d8).