SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training
Zhuang Wang
cs.DC, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/LMResiliency/lm-resiliency
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: SCOUT is a unified runtime failure-localization framework for LLM pre-training, built on the design principle of "identify outliers through strict-majority consensus among equivalent replicas." It
Terminology
Summary
SCOUT is a unified runtime failure-localization framework for LLM pre-training, built on the design principle of identify outliers through strict-majority consensus among equivalent replicas.
It addresses the problem that in LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin.
The paper focuses on three representative latent failure manifestations: hangs, stragglers, and silent data corruption (SDC), which fall outside this category because the launcher or device-health interface supplies an authoritative rank or device signal
for fail-stop failures.
The core observation is that failures are rare relative to healthy ranks: an individual failure usually leaves most ranks operating normally,
and practical LLM pre-training jobs typically use data parallelism to create equivalent replicas that execute the same model computation graph at the same training coordinate.
Therefore, a rank-local failure breaks the behavioral symmetry among these equivalent replicas.
SCOUT groups replicas that share a computation graph and aligns their training coordinate, input, and randomness, using majority behavior [as] a live, workload-relative reference for progress, elapsed time, or numerical output.
SCOUT's architecture has three layers: an integration layer, an evidence layer, and a decision layer. The integration layer normalizes framework structure and constructs peer groups
by mapping framework-specific parallelism into logical mesh addresses,
where peers retain the same model-parallel role while differing only along an eligible data-replica or state-shard dimension.
The evidence layer uses two mechanisms: bounded in-situ replay on the live accelerator
for SDC and stragglers, and an out-of-band (OOB) CPU observer
for hangs. The decision layer uses the Consensus Collective Communication (C3)
abstraction, which combines a standard AllGather with deterministic consensus to return the disagreeing ranks.
C3 maps diagnostic objects to a result containing a status, an outlier bitmap, and gathered evidence. It distinguishes three statuses: Agree means that the evidence satisfies the selected agreement rule and B = 0N,
Attributed means that the rule identifies one or more outliers,
and Inconclusive means that peers disagree but the evidence does not justify rank attribution.
For exact consensus on deterministic values, SCOUT first selects the most frequent value v∗ and its count c∗,
and "If c∗ > N/2, SCOUT marks every rank whose evidence differs from v∗. For timing evidence, C3 uses
robust statistical consensus, computing
their median m and a robust scale d and marking outliers where
τi − m > κd."
For in-situ replay, SCOUT replays selected forward, backward, and optimizer work to produce rank-distinguishing numerical and timing evidence.
For dense models, it builds recipes for embedding, hidden layers, output layers, and optimizer updates. The replay broadcasts the source input and RNG state, computes on copied state, and compares the materialized parameters and module output outside the training result.
SCOUT also separates computation and communication timing: An outlier in tM localizes the slowdown to module M and an outlier in tgather instead identifies FSDP parameter communication as the affected surface.
Cross-PG validation is used to localize faulty machines: the slow FSDP group r4, r5 and slow TP group r4, r12 intersect only at r4,
identifying the machine hosting both.
For MoE layers, SCOUT handles router, AllToAll, and expert computation. It compresses the shape search space for expert replay using coverage-based compression,
where Shape α covers shape β when their fingerprints match and α's path-count vector is no smaller in any component.
The paper reports that the exhaustive 128×128 campaign compressed 3,457 admitted shapes into 16 representatives,
with reductions of 97.66% to 99.54% across configurations.
For hang localization, SCOUT uses an independent observer
that reads the latest progress record from shared memory
and communicates through a Gloo group created separately from the training communicator.
The observer enters C3 only after visible progress remains silent beyond a configured threshold.
When a minority of ranks publishes a different collective fingerprint, SCOUT attributes the hang to rank-local software or control-flow divergence,
whereas when all ranks report matching progress and fingerprints, SCOUT instead classifies a group-scoped runtime or transport stall.
For checkpoint certification, SCOUT couples replay evidence to checkpoint eligibility and resumes training jobs only from certified checkpoints that preserve numerical trust.
It tracks three checkpoint states: latest, candidate, and verified. SCOUT promotes a candidate only after two consecutive accepted cycles, comprising 2K checks.
On failure, A confirmed SDC selects the verified checkpoint,
and for other failures, SCOUT runs the complete compressed recipe catalog and selects the latest checkpoint only if the sweep completes without SDC.
The implementation is implemented as a Python library over PyTorch
and supports PyTorch, Megatron-Core, TorchTitan, and DeepSpeed through public interfaces,
modifying neither training loops nor framework source.
It also integrates checkpoint certification with GEMINI.
Evaluation results show that across DDP, FSDP2, and HSDP, all nine 16-GPU cells for checkpoint recovery, SDC exclusion, and compute-straggler localization passed,
with each SDC-contaminated boundary was excluded globally, and a fresh runtime restored the preceding model, AdamW, caller-owned, CPU-RNG, and CUDA-RNG state bitwise.
For fault injection, 344/344 localized; 30/30 fault-free runs stayed clean
for dense numerical SDC, and 5,960/5,960 changed output or gradient; persistent faults localized ranks 7 and 15 on 8 and 16 GPUs
for MoE kernel-role SDC.
The paper acknowledges limitations: The current evaluation exercises software fault injection and dense and MoE mechanisms on at most 16 A100 GPUs across two hosts,
and It does not measure end-to-end throughput and resource overhead across replay cadences, or report recovery time and rollback distance.
Also, SCOUT targets permanent faults and intermittent faults that recur under a similar operating regime; a one-shot fault that does not recur during replay falls outside this coverage contract.
Improvements for AI systems
Improvements to AI Systems:
-
Self-Healing Pre-Training Infrastructure: Integrate SCOUT’s consensus-based failure localization directly into LLM training frameworks. The improved system can automatically detect, isolate, and attribute rank-local hangs, stragglers, and silent data corruption (SDC) during pre-training, then trigger checkpoint rollback to the last verified state without human intervention. This reduces job downtime and prevents corrupted model weights from propagating.
-
Fault-Tolerant Data-Parallel Execution: Use SCOUT’s strict-majority consensus among equivalent replicas to create a “voting” mechanism for numerical outputs. The improved system can mask or exclude outlier ranks in real time, ensuring that a single faulty GPU or communication link does not corrupt the global gradient or model state. This enables reliable training on heterogeneous or aging hardware clusters.
-
Proactive Straggler Mitigation: Leverage SCOUT’s timing-based consensus (median + robust scale) to detect compute or communication stragglers early. The improved system can dynamically rebalance workloads, migrate data shards, or throttle faster ranks to match the slowest healthy peer, minimizing idle time and improving overall throughput without waiting for job-level symptoms.
-
Certified Checkpointing for Crash Recovery: Adopt SCOUT’s three-state checkpoint lifecycle (latest, candidate, verified) with replay-based certification. The improved system guarantees that any resumed training starts from a numerically trustworthy state, eliminating silent corruption after crashes. It can also automatically select the most recent certified checkpoint, reducing rollback distance and wasted compute.
-
Cross-Dimension Fault Attribution: Use SCOUT’s cross-parallel-group intersection logic (e.g., FSDP group ∩ TP group) to pinpoint faulty physical machines or network links. The improved system can map logical rank failures to hardware locations, enabling automated node eviction, RMA, or network reconfiguration during live training.
-
MoE-Specific Fault Localization: Apply SCOUT’s coverage-based shape compression and expert replay to mixture-of-experts models. The improved system can detect which expert kernel, router, or AllToAll communication is faulty, even when the fault manifests only under specific input shapes, and can exclude only the affected expert path rather than the entire model.
-
Out-of-Band Hang Detection: Integrate SCOUT’s independent observer (via a separate Gloo communicator) to monitor collective fingerprints and progress records. The improved system can distinguish between rank-local software hangs and group-wide transport stalls, allowing targeted restarts of only the faulty rank or a full collective group, rather than killing the entire job.
-
Low-Overhead Replay-Based Validation: Use SCOUT’s bounded in-situ replay (broadcasting input and RNG state, computing on copied state) to validate forward/backward/optimizer steps without affecting the live training result. The improved system can run continuous, low-cost health checks on every checkpoint boundary, catching SDC before it contaminates the global model.
-
Adaptive Fault Coverage Contracts: Implement SCOUT’s limitation-aware design (permanent and recurring intermittent faults) into system monitoring. The improved system can explicitly track which fault types are covered and escalate one-shot anomalies to a higher-level diagnostic layer, preventing false confidence in unverified states.
-
Framework-Agnostic Observability: Build a unified diagnostic layer that maps PyTorch, Megatron-Core, TorchTitan, and DeepSpeed parallelism into logical mesh addresses. The improved system can provide consistent failure localization across different training stacks, enabling portability and easier adoption in diverse production environments.
What the Improved AI System Can Do:
-
Run uninterrupted pre-training on unreliable hardware, automatically excluding faulty ranks and resuming from certified checkpoints within seconds, with no silent corruption.
-
Scale to thousands of GPUs with a live, workload-relative reference for health, rather than relying on static thresholds or launcher-level signals.
-
Pinpoint the exact module, parameter-shard, or expert kernel responsible for a slowdown or numerical error, reducing debugging time from days to minutes.
-
Guarantee numerical trust for every checkpoint saved, ensuring that any resumed job produces scientifically valid results.
-
Operate across multiple frameworks without modifying training code, making fault tolerance a drop-in feature for existing LLM pipelines.
Sources
- Exploring Silent Data Corruption as a Reliability Challenge in LLM Training
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Understanding Silent Data Corruption in LLM Training
- OLMoE: Open Mixture-of-Experts Language Models
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
- GLM-5: from Vibe Coding to Agentic Engineering
- The Llama 3 Herd of Models
- Kimi K2.5: Visual Agentic Intelligence
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
- GLM-130B: An Open Bilingual Pre-trained Model
- OPT: Open Pre-trained Transformer Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing