Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

arXiv:2604.10800 · cs.SE, cs.AI, cs.CR, cs.LG, cs.PL · Submitted 2026-04-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Verify Before You Fix".

Jane: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we've been talking about this paper now called "Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis." Essentially, this paper tackles the reliability issue where AI predictions in agentic pipelines aren't actually verified conclusions, and acting on those without evidence just leads to more problems down the line.

Jane: That sounds really important, Tom; it gets right to the core problem of making sure these powerful tools give us trustworthy results instead of just guessing.

Lu: I'm really interested in how they tackle that uncertainty because building robust systems across different programming languages usually introduces a ton of complexity and potential for error.

Meng: From an engineering standpoint, when a model suggests something, we need to know if it actually works before we waste time trying to fix the wrong thing in the actual code base.

Lalam: I think this paper's focus on grounding predictions in observable evidence is what will really help improve how our culture handles these complex tasks by making decisions based on facts rather than just high-confidence guesses.

Tom: Exactly, and they propose a unified framework with three main reasoning stages—hybrid structural-semantic detection, execution-grounded agentic validation, and then validation-aware iterative repair—all held together by this rule that no repair action is taken without confirmation from the execution stage.

Jane: That sounds like a really solid structure for tackling the problem of probabilistic inferences becoming actual actions.

Lu: What caught my eye immediately was their use of a Universal Abstract Syntax Tree, or uAST, which normalizes Java, Python, and C++ into one shared schema so they can do zero-shot transfer across languages with high accuracy.

Meng: A universal representation sounds neat conceptually for cross-language work; does that mean the model learns the underlying structure once and can apply it everywhere?

Lalam: It’s like having a common dictionary for code, which should make understanding vulnerabilities across different languages much more consistent for the AI.

Tom: Right, and they show this uAST approach gives them a seventy-four point four three to eighty point one two percent F1 score in zero-shot cross-language tasks without needing to retrain separate models for each language, which is a big deal.

Jane: That suggests they've found a way to achieve significant cross-language generalization just by standardizing the structural representation upfront.

Paper summary: Lu: And on top of that, they fuse structural and semantic embeddings using a learned two-way gating mechanism, and those per-sample weights give them intrinsic explainability without adding extra cost to the process.

Tom: That intrinsic explainability is crucial; it means we can actually see *why* the model is making a certain prediction at each step, which helps build trust in the system's reasoning.

Meng: I’m curious about that gating mechanism; how does that fusion of GraphSAGE and Qwen2 point 5-Coder-1 point 5B embeddings translate into tangible performance gains in detection accuracy <ref:2604.10800#pg0>?

Lalam: It suggests a very nuanced understanding of both the code's structure and its actual meaning, which should lead to much more precise vulnerability identification overall.

Tom: And then they move into the execution-grounded agentic validation stage, where an agent generates an exploit hypothesis and synthesizes instrumented test programs using tools like AddressSanitizer for C++ or AspectJ for Java.

Jane: That step sounds like the crucial bridge; it’s where they take that probabilistic output and turn it into a binary, execution-backed verdict instead of just a prediction.

Lu: I think that plan-execute-verify loop is what really makes this system agentic and trustworthy; it moves beyond simple detection to actual verification through running tests in isolated containers.

Meng: From my perspective, that level of isolation and instrumentation sounds like the necessary practical step to ensure the agent isn't just hallucinating a fix or a vulnerability report.

Lalam: For me, seeing those probabilistic outputs convert into concrete execution evidence is what elevates this from an interesting research paper to something that could genuinely improve the safety culture in our development process.

Tom: And they follow that up with validation-aware iterative repair, where repairs only happen for samples confirmed by execution, and a fine-tuned model generates minimal patches using tools like redbaron for Python.

Jane: That conditional repair strategy is very smart; it prevents the system from wasting effort trying to fix things based on unverified guesses.

Lu: The fact that they show that removing the uAST normalization degrades cross-language F1 by twenty-three point four two percent really hammers home how foundational that structural abstraction is for their entire framework, and I think that's a significant finding <ref:2604.10800#pg1>.

Meng: That degradation number shows you exactly how much reliance the system has on having that shared structural understanding across different languages to maintain accuracy.

Paper summary: Lalam: It really shows that without normalizing the structure first, the benefits of cross-language capability just don't materialize as strongly for this kind of analysis.

Tom: So, we’ve seen how they unify detection, validation, and repair under that strict confirmation rule in "Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis."

Jane: And when we talk about the implications of this paper, it seems to be about moving AI from making strong suggestions to providing actionable evidence for decisions.

Lu: The big picture here is that this execution grounding mechanism could become a general principle for reducing uncertainty in any multi-stage LLM pipeline, not just code analysis.

Meng: If we can reliably ground agentic reasoning in execution results, it opens up possibilities where AI can participate in complex decision-making workflows with much higher confidence than current systems allow.

Lalam: I think the most profound implication for our field is shifting the focus from simply getting a high accuracy score to ensuring that every output has a verifiable trace leading back to an actual run or a confirmed test result.

Tom: That's right, it means we can start building agentic systems where human oversight is focused on reviewing structured diagnostic traces instead of just blindly trusting the model's summary.

Jane: It moves the conversation toward designing for safety by making the verification step an integral, non-optional part of the entire reasoning lifecycle.

Lu: The authors also pointed out that while they achieve high accuracy, there are still limitations, specifically mentioning a bounded reasoning horizon in multi-hop inference chains and incomplete environmental reconstruction in sandboxed verification.

Meng: Those limitations are important for us because they tell us precisely where the current capabilities of this approach stop working right now.

Lalam: It means we can’t expect this framework to solve every single edge case, like complex concurrency issues or cryptographic misuse, and that's an honest assessment of its current state.

Tom: So, while the paper shows a very effective mechanism for reducing overconfidence in agentic pipelines, it’s clear there’s still work ahead on handling those deeper reasoning complexities.

Jane: It really sets a high bar by showing how to implement execution grounding as a principled way to reduce uncertainty in complex, multi-step AI processes.

Conclusion: Tom: So, we've seen how this paper tackles the reliability issue where AI predictions in agentic pipelines aren't actually verified conclusions, and acting on those without evidence just leads to more problems down the line.

Jane: That sounds like a really important focus; it gets right to the core problem of making sure these powerful tools give us trustworthy results instead of just guessing.

Lu: I'm really interested in how they tackle that uncertainty because building robust systems across different programming languages usually introduces a ton of complexity and potential for error.

Meng: From an engineering standpoint, when a model suggests something, we need to know if it actually works before we waste time trying to fix the wrong thing in the actual code base.

Lalam: I think this paper's focus on grounding predictions in observable evidence is what will really help improve how our culture handles these complex tasks by making decisions based on facts rather than just high-confidence guesses.

Tom: Exactly, and they propose a unified framework with three main reasoning stages—hybrid structural-semantic detection, execution-grounded agentic validation, and then validation-aware iterative repair—all held together by this rule that no repair action is taken without confirmation from the execution stage.

Jane: That sounds like a really solid structure for tackling the problem of probabilistic inferences becoming actual actions.

Lu: What caught my eye immediately was their use of a Universal Abstract Syntax Tree, or uAST, which normalizes Java, Python, and C++ into one shared schema so they can do zero-shot transfer across languages with high accuracy.

Meng: A universal representation sounds neat conceptually for cross-language work; does that mean the model learns the underlying structure once and can apply it everywhere?

Lalam: It’s like having a common dictionary for code, which should make understanding vulnerabilities across different languages much more consistent for the AI.

Tom: Right, and they show this uAST approach gives them a seventy-four point four three to eighty point one two percent F1 score in zero-shot cross-language tasks without needing to retrain separate models for each language, which is a big deal.

Jane: That suggests they've found a way to achieve significant cross-language generalization just by standardizing the structural representation upfront.

Wrap-up: Lu: And on top of that, they fuse structural and semantic embeddings using a learned two-way gating mechanism, and those per-sample weights give them intrinsic explainability without adding extra cost to the process.

Tom: That intrinsic explainability is crucial; it means we can actually see *why* the model is making a certain prediction at each step, which helps build trust in the system's reasoning.

Meng: I’m curious about that gating mechanism; how does that fusion of GraphSAGE and Qwen2 point five-Coder-one point 5B embeddings translate into tangible performance gains in detection accuracy?

Lalam: It suggests a very nuanced understanding of both the code's structure and its actual meaning, which should lead to much more precise vulnerability identification overall.

Tom: And then they move into the execution-grounded agentic validation stage, where an agent generates an exploit hypothesis and synthesizes instrumented test programs using tools like AddressSanitizer for C++ or AspectJ for Java.

Jane: That step sounds like the crucial bridge; it’s where they take that probabilistic output and turn it into a binary, execution-backed verdict instead of just a prediction.

Lu: I think that plan-execute-verify loop is what really makes this system agentic and trustworthy; it moves beyond simple detection to actual verification through running tests in isolated containers.

Meng: From my perspective, that level of isolation and instrumentation sounds like the necessary practical step to ensure the agent isn't just hallucinating a fix or a vulnerability report.

Lalam: For me, seeing those probabilistic outputs convert into concrete execution evidence is what elevates this from an interesting research paper to something that could genuinely improve our culture.

Tom: And they follow that up with validation-aware iterative repair, where repairs only happen for samples confirmed by execution, and a fine-tuned model generates minimal patches using tools like redbaron for Python.

Jane: That conditional repair strategy is very smart; it prevents the system from wasting effort trying to fix things based on unverified guesses.

Lu: The fact that they show that removing the uAST normalization degrades cross-language F1 by twenty-three point four two percent really hammers home how foundational that structural abstraction is for their entire framework, and I think that's a significant finding.

Meng: That degradation number shows you exactly how much reliance the system has on having that shared structural understanding across different languages to maintain accuracy.

Wrap-up: Lalam: It really shows that without normalizing the structure first, the benefits of cross-language capability just don't materialize as strongly for this kind of analysis.

Tom: So, we’ve seen how they unify detection, validation, and repair under that strict confirmation rule in "Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis."

Jane: And when we talk about the implications of this paper, it seems to be about moving AI from making strong suggestions to providing actionable evidence for decisions.

Lu: The big picture here is that this execution grounding mechanism could become a general principle for reducing uncertainty in any multi-stage LLM pipeline, not just code analysis.

Meng: If we can reliably ground agentic reasoning in execution results, it opens up possibilities where AI can participate in complex decision-making workflows with much higher confidence than current systems allow.

Lalam: I think the most profound implication for our field is shifting the focus from simply getting a high accuracy score to ensuring that every output has a verifiable trace leading back to an actual run or a confirmed test result.

Tom: That's right, it means we can start building agentic systems where human oversight is focused on reviewing structured diagnostic traces instead of just blindly trusting the model's summary.

Jane: It moves the conversation toward designing for safety by making the verification step an integral, non-optional part of the entire reasoning lifecycle.

Lu: The authors also pointed out that while they achieve high accuracy, there are still limitations, specifically mentioning a bounded reasoning horizon in multi-hop inference chains and incomplete environmental reconstruction in sandboxed verification.

Meng: Those limitations are important for us because they tell us precisely where the current capabilities of this approach stop working right now.

Lalam: It means we can’t expect this framework to solve every single edge case, like complex concurrency issues or cryptographic misuse, and that's an honest assessment of its current state.

Tom: So, while the paper shows a very effective mechanism for reducing overconfidence in agentic pipelines, it’s clear there’s still work ahead on handling those deeper reasoning complexities.

Jane: It really sets a high bar by showing how to implement execution grounding as a principled way to reduce uncertainty in complex, multi-step AI processes.

Department of Computer Science, The George Washington University

cs.SE, cs.AI, cs.CR, cs.LG, cs.PL

Submitted: 2026-04-12

Updated: 2026-10-02

Code: https://github.com/PyCQA/redbaron

Importance score: 86/100

The gist: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in

Key concepts

Universal Abstract Syntax Tree (uAST)
This is a standardized structure that translates code from Java, Python, and C++ into one shared format. It allows the AI to reason about code across different languages without needing separate training for each language, improving cross-language understanding significantly.
Execution-Grounding Agentic Validation
Instead of just guessing if a bug exists, an agent creates a plan, writes specific test programs using tools like AddressSanitizer or monkeypatching, and runs them in secure containers. This turns uncertain predictions into concrete, binary proof based on actual code execution.
Validation-Aware Iterative Repair
Code fixes are only attempted if the execution validation confirms a fault. A specialized model then generates minimal patches using AST manipulation tools. The process repeats until the test passes, ensuring that repairs are always based on verified, observable evidence.
Learned Two-Way Gating
This technique fuses structural information (like code structure) and semantic meaning (the actual code content) into a single representation. By using learned weights to decide how much to trust each source for any given piece of data, the system gains intrinsic explainability.

Terminology

Summary

Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified conclusions, and acting on them without grounding in observable evidence leads to compounding failures across downstream stages.

The gist

Execution-grounded closed-loop reasoning is a principled and practically deployable mechanism for trustworthy LLM-driven agentic AI.

How it works

The framework is built around three LLM-driven reasoning stages: hybrid structural-semantic detection, execution-grounded agentic validation, and validation-aware iterative repair. The governing principle is that no repair action is taken without execution-based confirmation.

  1. Universal structural representation for cross-language LLM reasoning: A Universal Abstract Syntax Tree (uAST) normalizes Java, Python, and C++ into a shared schema to enable zero-shot transfer across languages at 74.43–80.12% F1, showing a 23.42% improvement over language-specific representations without per-language retraining.

  2. Interpretable hybrid reasoning with modality-aware gating: The model fuses structural (GraphSAGE) and semantic (Qwen2.5-Coder-1.5B) embeddings via a learned two-way gating, whose per-sample weights provide intrinsic explainability at no additional cost. This fusion achieves 89.84–92.02% intra-language detection accuracy.

  3. Execution-grounded agentic validation as a trustworthy AI mechanism: An LLM-driven agent implements a plan-execute-verify reasoning loop by generating an exploit hypothesis, synthesizing instrumented test programs (using AddressSanitizer for C++, AspectJ for Java, and monkeypatching for Python), and executing them within Docker containers under strict isolation. This stage converts probabilistic detector outputs into binary, execution-backed verdicts.

  4. Validation-aware iterative repair: Repair activates exclusively for execution-confirmed samples. A fine-tuned LLM generates minimal patches using AST manipulation libraries (like redbaron for Python), and a re-detection pass terminates the loop with success if the flag is zero, or triggers the next iteration if it is one.

Key Findings and Contributions

The framework achieves 89.84–92.02% intra-language detection accuracy and 74.43–80.12% zero-shot cross-language F1, resolving 69.74% of vulnerabilities end-to-end at a 12.27% total failure rate (Figure 3). Ablations confirm the necessity of both core components: removing uAST normalization degrades cross-language F1 by 23.42%, and disabling execution-grounded validation increases unnecessary repairs by 131.7%.

Architectural Principles

The paper establishes three core architectural principles for trustworthy agentic AI:

execution grounding as a mechanism for uncertainty reduction in multi-stage LLM pipelines.

structural abstraction for cross-domain transfer in learned systems,

and the use of intrinsic explainability through fusion design via per-sample gating weights.

Practical Implementation Details

The system is implemented in Python 3.11 and utilizes a pipeline that processes approximately 600–800 files per minute on an A100 GPU. The validation agent employs language-specific instrumentation: Java uses AspectJ for bytecode weaving, C++ uses compiler-level sanitizers (AddressSanitizer), and Python uses monkeypatching for runtime monitoring. The final repair model is a LoRA fine-tuned Qwen2.5-Coder-1.5B with rank 16 parameters.

Conclusion and Generalizability

The results demonstrate that execution-grounded feedback loops are a principled and empirically effective mechanism for reducing overconfidence in LLM-driven agentic pipelines. This approach is shown to be transferable to any domain where a binary oracle exists, such as theorem proving or scientific hypothesis testing. The framework positions human oversight as a first-class design principle by surfacing structured diagnostic traces for non-convergent cases.

Limitations and Future Work

Current limitations include the bounded reasoning horizon in multi-hop inference chains, incomplete environmental reconstruction in sandboxed verification, and coverage gaps in areas like concurrency and cryptographic misuse. Future work includes training the validation agent via reinforcement learning with exploit confirmation as a reward signal, integrating formal verification for provably safe patch synthesis, and expanding language coverage using Tree-sitter grammars.

References

[1] Rie Ando and Tong Zhang. Learning on graph with laplacian regularization. Advances in neural information processing systems, 19, 2006.

[2] Daniel Arp, Erwin Quiring, and Feargus et al. Pendlebury.

Improvements for AI systems

Based on the provided research paper, here are specific, actionable improvements for existing AI systems:


  1. Improvements to existing LLM-driven Agentic Pipelines (General Principle)

The core improvement is shifting from probabilistic prediction to execution-grounded decision-making. This applies to any multi-stage pipeline where an LLM makes a prediction and then acts on it without verification.

  1. Specific System Enhancement: Implementing the Three-Stage Closed Loop Framework

An existing agentic pipeline should be redesigned into a mandatory three-stage process:

  1. A Detection Stage (Structural/Semantic).

  2. A Validation Stage (Execution Grounding).

  3. A Repair Stage (Validation-Aware Iterative).

  4. Specific System Enhancement: Integrating Execution Grounding via Sandboxed Verification

Instead of allowing an LLM to act immediately after detection, the system must generate a concrete exploit hypothesis and execute it within a strictly isolated environment (Docker/Sandbox) using language-specific harnesses (ASan for C++, AspectJ for Java, monkeypatching for Python).

  1. Specific System Enhancement: Enforcing the Invariant: No Repair Without Execution Confirmation

The system must be architecturally constrained so that the repair module is physically gated by a binary execution result from the validation stage. This prevents acting on unverified findings.

  1. Improvement to Cross-Language Reasoning: Universal AST Normalization (uAST)

To enable reliable transfer across languages (Java, Python, C++), AI systems must be equipped with a shared structural schema that normalizes syntax into a universal Abstract Syntax Tree (uAST). This ensures that vulnerability patterns are recognized regardless of the surface language constructs.

  1. Improvement to Detection Accuracy: Hybrid Structural-Semantic Fusion with Interpretable Gating

Detection models should fuse structural reasoning (GraphSAGE on uAST) and semantic reasoning (Qwen2.5-Coder embeddings) using a learned two-way gating mechanism.

  1. Specific System Enhancement: Intrinsic Explainability via Modality Weights

The fusion mechanism must produce per-sample weighting scores that explicitly indicate whether the decision was driven by structural patterns (e.g., buffer overflows) or semantic context (e.g., logic flaws). This provides intrinsic explainability at no additional inference cost, replacing unstable post-hoc attribution methods.

  1. Improvement to Repair Efficiency: Validation-Aware Iterative Repair Loop

The repair mechanism must operate in a closed loop:

  1. Generate minimal patches based on confirmed exploit evidence (payload, behavior).

  2. Apply patches using language-specific AST manipulation libraries (RedBaron, LibClang).

  3. Re-run detection to verify the fix (re-detection).

  4. If the re-detection confirms success, terminate; otherwise, iterate up to five times with accumulated context and structured diagnostic traces for human review.

  5. Improvement to Human Collaboration: Structured Diagnostic Traces for Non-Convergence

When an autonomous repair loop fails to converge after its iteration budget (e.g., 5 iterations), the system must pause and present a structured diagnostic trace detailing all attempted patches, rejection reasons, and persistent indicators to a human reviewer, transforming the system into a collaborative tool rather than a black box that fails silently.

Sources

Related papers