From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents

summary

Video file (mp4)

The gist

The paper studies the problem of trajectory-to-evidence conversion in industrial research-agent settings, asking "what a completed research process has actually established." The authors observe that

In short

The episode discusses a paper by Kuaishou Technology about converting experimental trajectories into auditable evidence for industrial research agents. The hosts detail a framework involving source adaptation, claim qualification, and downstream reuse to ensure that claims are supported by verifiable evidence, focusing on the need for rigor in production settings.

Key concepts

Trajectory vs. Evidence
A trajectory is just a recording of an AI agent's actions during experiments. Evidence is what those actions actually prove. The paper argues that a completed trajectory does not automatically constitute valid evidence; it must be converted into qualified claims.
Bounded Generate-Verify-Repair Loop
This is the first stage of the framework where an agent produces something consequential, and a separate verifier checks it against requirements. If there is a problem, the producer fixes it in a loop until it passes verification or hits a limit. This prevents agents from rationalizing their own mistakes.
Claim Qualification
This stage involves grouping experiments addressing the same intervention to determine what evidence actually supports the result. Claims are labeled as Repairs (actionable improvements), Guards (warnings), or Withheld, based on execution validity and attribution.

Terminology used across episodes

This episode discusses

The paper

From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents · Read on arXiv

Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang, Yingzhi He, Peng Zhang, Jiangxia Cao, Yusheng Huang, Guohong Mu, Jian Liang, Ruiming Tang, Shuang Yang, Zhaojie Liu, Wenwu Ou, Kun Gai

Kuaishou Technology

Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study trajectory-to-evidence conversion, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate--verify--repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents".

Jane: The paper was written by Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Ruochen Yang et al. from Kuaishou Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we're cracking open a paper that's got a mouthful of a title — "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Jane, what's your first read on that title?

Jane: Tom, I love it because it's asking a question most of us never think to ask. When an AI agent runs a bunch of experiments, we tend to assume the log of what it did is the same as what we learned. This paper says no — a trajectory is just a recording of actions, and evidence is what those actions actually prove.

Tom: Right, and that distinction is huge. The authors are from Kuaishou Technology — that's a massive Chinese short-video and live-streaming platform — and they're dealing with recommendation systems that serve millions of users. Their agents run experiments constantly, tweaking ranking models, and they need to know which changes are real improvements.

Jane: Exactly. And the word "auditable" in the title is doing a lot of work. They're not just saying "trust the agent's summary." They're building a system where every claim has a paper trail — which experiment, which code change, which measurement, under what conditions.

Tom: So it's like the difference between a scientist saying "this works" and a scientist showing you the lab notebook with every run recorded, every failure noted, every assumption flagged.

Jane: That's the whole thing. And the industrial part matters too. In a research lab, you can afford to re-run things. In production, every experiment costs GPU time and engineering hours. You want to know before you deploy whether the evidence actually supports the change.

Tom: And that's why they're calling it "trajectory-to-evidence conversion." The agent leaves behind a trail of proposals, code states, execution logs, measurements, failures, retries. The paper's job is to sift through all of that and figure out what's actually established.

Jane: Which is harder than it sounds, because a completed experiment doesn't automatically mean a valid result. The code might not match the proposal. The comparison might be confounded. A later modification might erase an earlier gain. The trajectory has all of that mixed together.

Tom: So the title is really a promise — we're going to turn messy process into clean, defensible conclusions. And the authors have built a whole framework to do it. We'll dig into how they verify artifacts, how they qualify claims, and what happens when they actually test this in production.

Jane: And I have to say, the fact that they're doing this at Kuaishou scale — with real users, real A/B tests, real deployment decisions — makes this more than an academic exercise. This is about whether we can trust AI to run our experiments for us.

Tom: Stay with us — next segment, we're breaking down the abstract and what they actually claim to have achieved.

Summary: Tom: We're back with "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Jane, let's get into the abstract — what are these folks actually claiming?

Jane: So the core claim is that a completed research trajectory is not automatically evidence. You have to convert it. And they've built a framework that does that conversion in three stages — source adaptation, claim qualification, and downstream reuse.

Tom: And the first stage is this "bounded generate-verify-repair" loop. Every time the agent produces something consequential — a proposal, a code change, an experiment spec — a separate verifier checks it against the evidence and the requirements of the next step. If there's a problem, the producer fixes it, and they loop until it passes or they hit a limit.

Jane: Right, and the clever part is that the verifier is context-isolated from the producer. The producer's private reasoning doesn't leak into the verification step. They communicate only through typed artifacts and explicit issue sets. That's a real design choice — it prevents the agent from just rationalizing its own mistakes.

Tom: Then stage two is claim qualification. After the experiments run, they group rounds that address the same intervention and ask — what does the evidence actually support? They check execution validity — did the code do what it was supposed to? Did the comparison follow protocol? And they check attribution — can we actually credit the intervention for the change?

Jane: And based on that, they label each candidate claim as a Repair — that's an actionable improvement with evidence; a Guard — that's a warning about a failure mode; or Withheld — not enough evidence to support anything. Only Repairs and Guards become persistent records.

Tom: And here's a number that stuck with me — out of fourteen candidate claims from their experiments, only nine became records. Eight Repairs and one Guard. Five were withheld. So nearly a third of what the agent tried just didn't hold up.

Jane: That's the whole point, Tom. If you just copied the trajectory, you'd treat all fourteen as knowledge. The framework says no — five of those are not supported. And the records that do survive carry explicit applicability boundaries. They say under what conditions the finding holds.

Tom: Then stage three is downstream reuse. When a new target task comes in, the controller looks at the frozen records and decides — Apply, Defer, or Reject. Apply means we run an experiment based on this record. Defer means we need more information. Reject means the record doesn't fit.

Jane: And this is where it gets interesting, because their controller is not great at saying yes. We'll get into the numbers later, but the affirmative precision is only twenty-five percent. That's a bottleneck they're honest about.

Tom: So the framework works — it filters out bad claims — but the reuse side still has room to grow. And that's a really honest result for an industrial paper.

Jane: It is. They're not claiming perfection. They're showing what works and where the remaining problems are. That's the kind of research that actually moves the field forward.

Tom: Next segment, we're going to look at the improvements they're proposing over existing systems — what's actually new here.

Improvements: Tom: Welcome back to the show, still on "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Jane, let's talk about what's actually new here — what improvements does this paper bring over what already exists?

Jane: So the key improvement is separating the evidence from the trajectory. Existing systems like Reflexion or ExpeL store past experience as reflections or reusable snippets. The AI Scientist and Agent Laboratory use trajectories to guide future work. But none of them explicitly ask — what does this trajectory actually prove?

Tom: Right, and that's the gap. The closest thing they mention is NOVA, which does semantic verification and keeps round-level records. But even NOVA treats the records more like a history of what happened, not a qualified set of claims with applicability boundaries.

Jane: Exactly. And the improvement here is the two-part verification harness. First, the evidence-and-constraint check — does this artifact conflict with what we know? Second, the downstream requirement check — does this artifact omit information the next step needs? That second check is really clever because it catches things the producer never even thought to include.

Tom: And the context isolation is a real improvement too. The producer, the evidence verifier, and the requirement reviewer all run in separate contexts. They can't share rationales. They only communicate through typed artifacts and explicit issue sets. That prevents the agent from talking itself into believing its own output.

Jane: And then the claim qualification step — that's another improvement. They don't just accept a result because the experiment ran. They check execution validity through four gates: implementation fidelity, execution completion, evaluation protocol, and mechanism preservation. If any gate fails, the round is excluded from persistent claims.

Tom: And the mechanism preservation gate is interesting. These are paper-to-target adaptations. The agent takes a method from a research paper and adapts it to a production recommendation system. The defining mechanism is the core idea that must survive the adaptation. If the agent's changes remove that mechanism, it's not a valid adaptation.

Jane: Right, and that's a really practical concern. An agent could optimize a metric by stripping out the thing that made the paper interesting. This gate prevents that.

Tom: And then the downstream reuse — they freeze the records before any target adaptation. So the memory can't be contaminated by the very experiments it's supposed to inform. That's a clean experimental design.

Jane: It is. And it lets them evaluate the controller separately from the framework. They can ask — is the controller making good decisions? Without worrying about whether the records themselves are sound.

Tom: So the improvements are really about rigor — separating verification from generation, separating qualification from execution, separating reuse from construction. Each separation makes the system more auditable.

Jane: And that's the word — auditable. You can trace every claim back to its evidence. You can see why a claim was withheld. You can see the applicability boundary. That's what makes this trustworthy enough for production.

Tom: Next segment, we're going to look at the first page of the paper itself — the intro and the problem setting — and see how they frame this whole challenge.

First Page: Tom: We're still on "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Jane, let's actually look at the first page of the paper — how do they frame the problem?

Jane: So the first page opens with a really sharp observation. Research agents are increasingly running multi-round machine-learning experiments in industrial settings. They retain the trajectories to guide later decisions. But a completed trajectory is not automatically evidence.

Tom: And they spell out why. Generated artifacts may be unsupported or incomplete. Executed rounds may be invalid or confounded. Later modifications may obscure earlier findings. So the question becomes — what has this research process actually established?

Jane: And that's the framing for the whole paper. They're not asking "did the agent finish the task?" They're asking "what can we defensibly claim after all of this?"

Tom: They also introduce the problem setting formally. A paper-to-target adaptation task takes a paper, a target setting, and a defining mechanism. The target setting includes the initial repository, data, evaluation utility, and constraints. The defining mechanism is the part of the paper that must survive adaptation.

Jane: And there's a resource limit — up to B machine-executed experiments. That's a really industrial constraint. You can't just run experiments forever. You have a budget. And agent-side verification runs before dispatch to reduce the risk of wasting that budget on invalid artifacts.

Tom: And then they make a key distinction. They have source adaptation tasks and a target adaptation task. Source and target are roles in the evaluation protocol, not disjoint paper identities. The same paper could be a source in one episode and a target in another, as long as record construction and downstream use happen in separately initialized episodes.

Jane: That's a careful experimental design choice. It prevents contamination — you can't use records from the very experiment you're evaluating.

Tom: And they also define the record structure. Each record has five semantic blocks: source context, intervention, measured outcome, supporting evidence, and applicability boundary. That's the auditable unit.

Jane: And the source context identifies the mechanism, the target condition, and the triggering observation. The intervention includes both the action and its locus — where in the model it acts. The boundary limits reuse.

Tom: So the first page is really setting up the problem with precision. They're not hand-waving about "AI agents doing science." They're defining exactly what a claim is, what evidence supports it, and how you know when it applies elsewhere.

Jane: And that precision is what makes the rest of the paper work. When they report that five out of fourteen claims were withheld, you know exactly what that means. When they report a twenty-five percent affirmative precision on the controller, you know exactly what they measured.

Tom: Next up — the conclusion. We'll wrap up what this paper means for the field and where it leaves us.

Conclusion: Tom: Alright, we're wrapping up our discussion of "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Jane, give us the final summary.

Jane: So the paper's core message is that a trajectory is a source of candidate evidence, not knowledge itself. You have to verify artifacts before they affect downstream steps. You have to qualify claims after execution — checking execution validity, attribution, and applicability boundaries. And you have to preserve only the claims that survive as auditable records.

Tom: And the results back that up. Their verification loop improved proposal gate acceptance from seventy-three percent to one hundred percent. Code completeness went from ninety-two percent to ninety-seven percent. Observability from sixty-seven percent to one hundred percent. But the audit also showed that coverage of source papers' Method sections was incomplete — eighty-seven percent of blocks, not one hundred percent.

Jane: And the trajectory analysis was really striking. Out of thirty paper-to-target adaptations, twenty-six had a later round outperform the first. But twenty-two had the final round underperform an earlier best. So the final state is often not the best state. That's a direct argument for building memory from validated intermediate results, not the endpoint.

Tom: And the claim qualification — fourteen candidates, nine records. Eight Repairs, one Guard, five Withheld. The framework is genuinely filtering.

Jane: And the downstream controller — that's where the honest limitations show. Affirmative precision of twenty-five percent. The controller said Apply four times, and only one matched the reference. They're treating Apply as a hypothesis for execution, not a confirmation of applicability.

Tom: But even with those limitations, two candidates from the complete workflow produced positive online lifts — zero point seven five percent on page duration for live streaming and six point three four percent on net growth utility for user growth. So the pipeline can produce deployable improvements.

Jane: And the limitations section is refreshingly honest. The production case studies establish candidate-level performance, not framework-level attribution. The record-use diagnostic has only eight pairs. The controller evaluation is class-imbalanced. They're not overclaiming.

Tom: So what's the takeaway for the field? I think it's that we need to treat AI-generated experimental trajectories with the same skepticism we'd apply to a human scientist's lab notebook. Not everything in it is a finding.

Jane: And the framework they built — bounded verification, claim qualification, auditable records — gives us a concrete way to do that. It's not perfect, but it's a real step toward trusting AI to run experiments at industrial scale.

Tom: Well said, Jane. That's our show on "From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents." Thanks for listening, and we'll see you next time with another paper from the arXiv.

Jane: Take care, everyone.

More episodes

← Home