Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
Xing Zhang, Yanwei Cui, Guanghui Wang, Peiyang He
AWS Generative AI Innovation Center
cs.MA, cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Project page: https://langchain-ai.github.io/langgraph
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 95/100
The gist: This paper presents a deployment-oriented two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing.
Terminology
Summary
This paper presents a deployment-oriented two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic “librarian” continuously ingests public, timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph into an always-current source of truth, not per-query RAG over raw chunks. A portable multi-agent “writer” runtime then composes a long, contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as of ≤ T (no look-ahead); red-team verdicts flow back into the librarian, closing the loop.
The system runs as two decoupled phases coupled by a timestamped store and closed by a write-back loop. Phase A: the librarian (deterministic) ingests public sources, recording each one’s true publication date and trust tier. The librarian maintains knowledge at three progressively-distilled layers: quote-grounded evidence cards (raw facts, each pinned to a source quote), an authoritative metric ledger (one reconciled value per company-metric pair), and a claim graph of contradicts, supersedes, and qualifies edges over those values. Each layer is derived deterministically from the one below, so a number traces down to its source quote and up to its conflicts. The core costs essentially nothing to re-run; an LLM is used only at one clearly-bounded seam: a cheap refinement pass (Claude Haiku 4.5) that corrects a numeric card’s value/unit against its own quote or demotes it to qualitative.
The bridge: point-in-time projection. Given a cutoff T, the bridge projects the store into the writer’s four artifacts (outline, evidence cards, metric ledger, claim graph), filtering to as of ≤ T. This is the no-look-ahead seam: a report “as of” a past date sees exactly the evidence that existed then.
Phase B: the portable writer runtime. The writer is a self-contained, headless runtime over a tiered LLM provider and a bounded-concurrency pool, with no dependency on an external agent framework. Its orchestration wraps a set of deterministic scripts (slice, tag-normalize, QC, render) around the LLM calls. The workflow is a fixed directed acyclic graph (DAG): slice each section → compose (one LLM call per section, tier-routed) → normalize → red-team (a Claude Opus 4.8 “prosecutor” per section that returns holds/weak/refuted verdicts) → apply verdicts → a bounded rewrite of affected sections → deterministic convergence backstop → QC gate → render.
Write-back loop. When the red-team refutes a card, the verdict maps to a librarian source override/claim refuted, so the next run at the next cutoff inherits the correction: the store is a living library, not a static dump. The write-back is guarded: an override never invents a value, it only demotes the refuted source so the ledger falls back to a pre-existing same-kind alternative, and every change appends a row to an append-only audit log, so promotion is human-gated and reversible. Regeneration is idempotent: it carries a human promotion forward rather than resetting it, but snapshots the exact evidence approved, so if that evidence later vanishes, even when back-filled cards keep the count unchanged (an “anchor swap”), the claim is flagged for re-validation, not silently kept.
The distributed multi-agent design: (1) Parallel per-section agents. The report outline is a partition: each section is an independent compose task, and the sections fan out across a bounded-concurrency pool rather than being written serially. Sections are the natural unit of parallelism because the outline contract makes them near-independent: the only shared state is the metric ledger, read-only at compose time. (2) Heterogeneous agents. A difficulty router sends conflict-touching sections to a stronger, costlier model (Opus) and routine sections to a cheaper one (Sonnet), so compute is spent where the reasoning is hard. (3) Separation of powers. Composition and criticism are different agents with opposing objectives: a composer writes to satisfy the section contract, an independent Opus “prosecutor” red-teams the draft to break it. Neither grades its own work, and the arbiter that decides delivery is the deterministic QC gate, not an LLM. (4) Coordination through a shared store, not messages. The agents never talk to each other directly; they coordinate stigmergically through the trust-tiered store: composers read the same authoritative ledger, and the red-team writes verdicts back to it. This adds parallelism without the usual multi-agent failure mode of concurrent writers diverging: because the single authoritative value lives in the shared ledger rather than in each agent’s context, two sections physically cannot cite different numbers for the same metric.
Why the fan-out is safe. The consistency guarantee is structural, not a matter of scheduling. The bridge emits an immutable, point-in-time snapshot at cutoff T, and every composer reads from that single snapshot, so within one run there are no concurrent writers: no lock, barrier, or two-phase commit is needed, and the classic shared-memory races (write–write, torn reads, deadlock) cannot arise. The only writer is the red-team’s write-back, deferred to the next cutoff, never mid-run. Read-only fan-out over an immutable snapshot makes each compose step idempotent and order-independent, so the worker count K trades latency against cost but cannot change the delivered numbers.
Trust & Consistency Mechanisms. Source tiers and permitted use. Every source is typed from text/path cues and assigned a trust tier together with a permitted use that governs whether its numbers may be cited: U.S. Securities and Exchange Commission (SEC) filings become official (usable as hard evidence), U.S. Bureau of Labor Statistics (BLS) macro-statistics releases become gov stat (supporting evidence: authoritative, but for macro context, not a company’s own figures), and Wikipedia becomes media (routing only). The tiers are strictly ordered (official > gov stat > sell side > media). Routing-only sources may inform entity/topic routing and context but can never become a citable hard-evidence number, and a supporting-tier macro value can never displace a company’s own official figure: the “official-first” rule.
Metric ledger. For each (company, metric) the ledger selects one authoritative value by a fixed policy: tier dominates, then corroboration (distinct-source count) breaks ties within a tier, then recency (as of). Competing values are retained as alternatives and flagged as a conflict when a comparable same-kind figure materially disagrees (a fixed > 15% threshold, so unit-equal restatements do not spuriously fire), but the report cites the single authoritative value, which is what eliminates cross-section drift.
QC gate (six deterministic checks). The delivery gate is language-neutral and LLM-free: (1) orphan citations, (2) unsourced numbers, (3) numeric drift across sections, (4) buried contradictions (a claim-graph conflict whose two endpoints are not reconciled together), (5) unregistered metrics, (6) cross-section contradiction. A report is deliverable only when the error set is empty.
Deployment & Dataset. The authors self-collected a public, redistributable, English-only corpus at production scale: 6,130 sources extracting to 555,926 evidence cards (457,561 numeric) and a metric ledger of 2,589 authoritative company-metric values, 2,132 of them carrying recorded conflicts. The three tiers are 5,397 SEC EDGAR filings (official) for 295 issuers across 11 sectors, 672 U.S. Bureau of Labor Statistics (BLS) macro releases (gov stat), and 61 Wikipedia articles (media). The design point is one library, many reports: a single maintained store serves multiple report theses rather than being purpose-built for one. From this store the authors generate four flagship point-in-time reports at a common cutoff (2025-12-31) (AI-compute, energy, healthcare/pharma, and banks), each projected by the sector-scoped bridge, and each passing the QC gate with zero errors.
Evaluation. Every headline metric is machine-computed (no LLM decides a reported number) and, where an ablation applies, compared on the identical corpus against an explicit baseline arm.
E1: a shared ledger removes cross-section drift. Without a shared ledger, a writer grounding each section independently surfaces every competing value for a metric; the ledger collapses each to one authoritative value. On the real store at the final cutoff, the no-ledger baseline would emit 6,845 contradictory figures across 2,105 metrics with competing values; ours emits 0. Replayed across seven cutoffs, the ledger’s authoritative value changed 4,732 times, and every change was justified by newer, higher-tier, or more-corroborated evidence (0 unexplained).
E2: grounding. Across all four flagship theses composed from the one library (AI-compute, energy, healthcare/pharma, banks; cutoff 2025-12-31), every numeric-bearing body line must carry an evidence citation or a metric annotation. Aggregate grounding is 202/203 (99.5%) with 0 orphan citations and 0 unregistered metrics; three reports are 100%, and the lone exception is a synthesis sentence whose figures are each cited earlier in the same section.
E3: trust tiering suppresses rumor and quarantines macro context. End-to-end on the real store, all three tiers classify correctly (5,397 filings as hard evidence, 672 BLS releases as supporting evidence (gov stat), 61 Wikipedia articles as routing only, 0 misclassified), and 0 of the 457,561 numeric cards trace to a routing-only source. The third tier is genuinely mined (2,352 numeric gov stat cards: CPI, PPI, payrolls, unemployment), yet because a macro statistic carries no company attribution, 0 gov stat values displace a company’s own official figure and the per-company ledger stays 100% official.
E4: tier-first selection beats popularity. The authors run two selection policies on one labeled gold set of metric clusters: our tier-first ledger, and a popularity-first baseline that takes the value with the most distinct backing sources (the “most-cited”/semantic-layer heuristic), ignoring tier. Tier-first is correct on 22/22 cases; popularity-first scores only 9/22. Thirteen cases are popularity traps: a widely-repeated lower-tier value (a rumor echoed by several media sources, or a corroborated macro statistic) competes with a single official filing; the tier rule survives all thirteen, popularity adopts the wrong value every time. The set spans the full configured lattice (official > gov stat > sell side > media), including the invariant that a newer, more-corroborated gov stat value still cannot displace an official figure (which gov stat may anchor only when none exists), and corroboration breaking ties only within a tier.
E5: the checker is trustworthy (recall and precision). A gate is only trustworthy if it both catches real defects and stays quiet on clean text. The authors clone a clean, QC-passing run and (i) inject five defect classes (orphan citation, unsourced number, broken cross-reference, unregistered metric, buried contradiction) and (ii) apply three negative controls: defect-free perturbations that must not fire (a paragraph reusing only already-valid citations and metric tokens, a duplicated grounded line, number-free prose). Recall is 1.0 (5/5) and precision is 1.0 with a 0 false-positive rate on the controls. Crucially, the authors separate delivery-blocking from advisory detection: the delivery gate is “error set empty,” and 3/3 error-level defects raise a blocking error, while the two warning-level defects are caught but advisory by design.
E6: in the distributed writer, parallelism and routing cut cost at comparable quality. Isolating the multi-agent compose fan-out on identical slices over three repeats, bounded-concurrency parallelism across the per-section agents runs 3.7× faster than serial, and difficulty-tiered routing (conflict-touching sections → Opus, the rest → Claude Sonnet 5) costs 4.1% less than sending every section to Opus. The cost gap is deliberately modest on this flagship: it is conflict-heavy, so 5 of 6 sections touch an unresolved edge and correctly route to Opus, so routing saves little precisely when the report is hard, and the same dial saves far more on a low-conflict thesis. All four variants pass QC with zero errors, but a binary gate cannot rank them, and all-Sonnet is the cheapest, so “equal quality” needs an independent signal. The authors add a deterministic, graded quality score (grounding coverage + conflict-pair coverage + output-contract adherence, computed on each variant’s actual prose, independent of the QC gate). The counterintuitive result: tiered routing scores above the all-Opus ceiling (+0.079) and far above all-Sonnet (+0.262). Spending the strong model only where reasoning is hard beats spending it everywhere, so the cheaper all-medium point is not the default: it saves on easy sections but degrades exactly the conflict-heavy synthesis that routes to Opus.
E7: the living library grows without look-ahead. Replaying the store across seven cutoffs yields 0 look-ahead violations and monotonic growth (235,373→555,312 cards; 1,659→6,054 card-bearing sources), capturing 4,395 post-initial evidence-arrival events (new filings, restatements, new conflicts) and lifting recorded metric conflicts from 1,770 to 2,132. This closes end to end at the report level: regenerating the AI-compute flagship at three advancing cutoffs yields a report that grows in lockstep (27,104→42,900 cards, 255→276 metrics, 562→655 reconciled conflict edges) while every cutoff stays deliverable (QC errors = 0), and all four flagship theses compose at the shared 2025-12-31 cutoff from this one library.
E8: the write-back loop self-corrects a later run. The authors trace one worked case end to end using the librarian’s real override machinery. In this illustrative scenario the report-side red-team challenges the authoritative interest figure (19mn) after flagging its backing filing as low-confidence. The verdict maps to a librarian source override (status retracted); on re-ingestion the 284 evidence cards from that filing inherit the non-active status, and because the metric ledger considers only active-source cards, the authoritative value self-corrects to 9mn, an already-recorded alternative, with 0 manual value edits. Tracing the same override path on a batch of auto-discovered conflicted metrics, 5 of 6 write-backs self-correct, each to a pre-existing same-kind alternative (0 manual edits), so the loop closes on many values, not one, deterministically and auditably.
Limitations. The corpus is English-only and three-tier as collected (official SEC filings, gov stat BLS macro statistics, media Wikipedia; sell side omitted). The gov stat tier is authoritative but macro-only: its releases carry no company attribution, so by construction it enriches context, not the per-company ledger. Entity linking is substring-based, so it occasionally over-attributes a metric when a company’s short name is a substring of unrelated filing text (e.g. “3M”); this is an auditable extraction artifact, not a ledger-policy error. The E4 cross-tier gold set is designed rather than sampled (the deployed corpus has no cross-tier clusters), though the within-tier rule is corroborated on 2,132 real conflicts; E6 uses a deterministic quality proxy, not human judgement; and E8 traces write-back on a small batch. E1’s no-ledger arm isolates the ledger’s effect, not a strong shared-state competitor (a graph- or semantic-layer-backed retriever), the natural next comparison our design is built to host.
Lessons. (i) Govern by trust, not popularity: a most-cited-value heuristic adopts widely-repeated rumor, while an official-first ledger is simpler and correct (E1, E3, E4). (ii) Keep the core deterministic, put the LLM at the edges: deterministic selection, consistency, and QC make the headline metrics reproducible and the gate meta-evaluable (E5). (iii) Own the runtime: standard multi-agent building blocks (file-per-agent artifacts, contract-first outlines) carry over cleanly to a headless service without binding to any external agent framework. (iv) Routing buys quality, not just cheapness: tiered composition matches or beats the all-Opus ceiling while a uniform-cheap baseline degrades, at modest dollar saving on a conflict-heavy report (E6).
Conclusion. Separating a maintained, trust-tiered, point-in-time library from report writing turns long-form generation’s chronic drift, provenance loss, and trust flattening into mechanically-checkable, largely eliminated properties.
Improvements for AI systems
Improvements to AI systems:
-
Deterministic knowledge library with trust-tiered ontology. Replace per-query RAG over raw chunks with a continuously-ingested, timestamped store organized into three layers: quote-grounded evidence cards, an authoritative metric ledger (one reconciled value per entity-metric pair), and a claim graph of contradicts/supersedes/qualifies edges. Each layer is derived deterministically from the one below, so every number traces to a source quote. LLM usage is confined to a single cheap refinement pass (correcting value/unit against its own quote).
-
Point-in-time projection for no-look-ahead generation. Given a cutoff date T, project the store into writer artifacts filtered to evidence with as of ≤ T. This guarantees reports
as of
a past date see exactly the evidence that existed then, eliminating look-ahead bias. Replay across advancing cutoffs yields monotonic growth with zero violations. -
Structural consistency via immutable snapshots, not locks. Emit an immutable point-in-time snapshot per run; all composer agents read from that single snapshot. This eliminates write-write races, torn reads, and deadlock without barriers or two-phase commit. Worker count trades latency vs. cost but cannot change delivered numbers.
-
Tier-first value selection over popularity. Select authoritative values by fixed policy: trust tier dominates, then corroboration (distinct-source count), then recency. This suppresses widely-repeated rumor (e.g., media-echoed figures) in favor of a single official filing. In evaluation, tier-first is correct on 22/22 cases vs. 9/22 for popularity-first.
-
Separation of powers: composer vs. prosecutor agents. Use independent agents with opposing objectives—a composer writes to satisfy a section contract; a stronger
prosecutor
red-teams the draft to break it. Neither grades its own work; a deterministic QC gate decides delivery. Red-team verdicts write back to the store, closing a self-correction loop. -
Stigmergic coordination through shared store, not messages. Agents never talk directly; they coordinate through the authoritative ledger. This prevents cross-section drift: two sections physically cannot cite different numbers for the same metric because the single value lives in shared state, not per-agent context.
-
Difficulty-tiered routing for cost-quality tradeoff. Route conflict-touching sections to a stronger, costlier model and routine sections to a cheaper one. This beats both all-strong and all-cheap baselines on a graded quality score (grounding coverage + conflict-pair coverage + contract adherence), not just cost.
-
Deterministic QC gate with six checks. Deliver only when error set is empty: (1) orphan citations, (2) unsourced numbers, (3) numeric drift across sections, (4) buried contradictions, (5) unregistered metrics, (6) cross-section contradiction. The gate is LLM-free, language-neutral, and meta-evaluable (recall 1.0, precision 1.0 on injected defects).
-
Write-back loop with guarded overrides. When red-team refutes a card, map verdict to a source override that demotes the refuted source (never invents a value), falls back to a pre-existing same-kind alternative, and appends to an append-only audit log. Regeneration is idempotent, human-gated, and reversible.
-
Anchor-swap detection for evidence stability. If evidence approved in a prior run later vanishes, even when back-filled cards keep counts unchanged, flag the claim for re-validation rather than silently keeping it.
What the improved AI system can do:
-
Generate long, contradiction-free, evidence-grounded reports at any historical cutoff with zero look-ahead, zero cross-section numeric drift, and 99.5% grounding (every numeric line cited).
-
Maintain a living knowledge library that grows monotonically (e.g., 235K→555K evidence cards across seven cutoffs) and self-corrects via write-back loops (e.g., authoritative value changes from 19M to 9M with 0 manual edits).
-
Produce multiple reports from one shared library (e.g., four flagship theses across AI-compute, energy, healthcare, banks) with zero QC errors.
-
Scale parallelism safely: per-section agents fan out over immutable snapshots, achieving 3.7× speedup over serial with no consistency risk.
-
Resist rumor and misinformation: official-first tiering blocks media-sourced numbers from ever becoming citable hard evidence, and macro statistics cannot displace company-specific official figures.
-
Provide auditable, reversible corrections with full provenance—every number traces to a source quote, every override is logged, and every change is human-gated.
-
Operate as a headless runtime with no external agent framework dependency, using deterministic scripts (slice, normalize, QC, render) wrapped around bounded LLM calls.
Abstract
Long-form research reports generated by large language models drift, contradict themselves, and lose provenance: the same metric appears with different values, and rumor is quoted as confidently as an audited filing. We present a two-tier agentic system that separates a maintained, point-in-time knowledge library from report writing. A deterministic "librarian" ingests timestamped sources into a trust-tiered ontology, layering evidence cards, an authoritative metric ledger, and a claim graph into an always-current source of truth, not per-query RAG over raw chunks. A portable multi-agent "writer" runtime then composes a contradiction-free, evidence-grounded report at any knowledge cutoff T, reading only evidence with as of <= T (no look-ahead); red-team verdicts flow back into the librarian. We evaluate on a self-collected, public corpus of 6,130 sources yielding 555,926 evidence cards (SEC EDGAR filings across 295 issuers and 11 sectors, U.S. Bureau of Labor Statistics releases, and Wikipedia). From the one library we compose four point-in-time reports on distinct theses and run eight reproducible experiments, whose headline metrics come from a deterministic quality-control gate, itself validated by defect-injection meta-evaluation at recall 1.0 and precision 1.0. A shared metric ledger removes 6,845 cross-section contradictions to zero. Tier-first selection is correct on 22/22 gold cases where a popularity-first baseline scores only 9/22; trust tiering leaks zero media-sourced numbers, and no government statistic displaces a company's own filing. A red-team refutation propagates back and self-corrects a later run with zero manual edits. Replay exhibits zero look-ahead violations across seven cutoffs while the library grows from 235,373 to 555,312 cards. Difficulty-tiered model routing exceeds the all-Opus quality ceiling while running 3.7x faster than serial.
Sources
- REGAL: A Registry-Driven Architecture for Deterministic Grounding of Agentic AI in Enterprise Telemetry
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Retrieval-Augmented Generation for Large Language Models: A Survey
- FinanceBench: A New Benchmark for Financial Question Answering
- MemGPT: Towards LLMs as Operating Systems
- OntoMetric: An Ontology-Driven LLM-Assisted Framework for Automated ESG Metric Knowledge Graph Generation
- Verified Multi-Agent Orchestration: A Plan-Execute-Verify-Replan Framework for Complex Query Resolution
- FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
- FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning