LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Zhixin Zhang, Xinke Jiang, Zhibang Yang, Weixuan Xu, Guohong Qiu, Xu Chu, Junfeng Zhao, Yasha Wang
Peking University · National Engineering Research Center for Software Engineering, Peking University · Key Laboratory of High Confidence Software Technologies, Ministry of Education · Center on Frontiers of Computing Studies, Peking University · Peking University Information Technology Institute (Tianjin Binhai)
cs.LG, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 15 pages, 8 figures
Code: https://github.com/jiangxinke/Agentic-RAG-R1
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation Abstract Large language model agents increasingly rely on long-horizon reasoning to solve complex
Terminology
Summary
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Abstract
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. Learning effective reflection, however, is challenging because reflection is performed locally within the current branch, whereas its utility can only be determined by its contribution to the final trajectory outcome. This local–global mismatch makes outcome-based reinforcement learning provide only local, sparse and delayed supervision for reflective decisions. To solve these, we propose LoongReflect, a training framework that formulates reflection as a memory-control policy. The agent operates over a reversible trajectory tree using explicit and actions. Reflection consolidates verified facts, missing evidence, and branch-specific risks into working memory, while backtracking removes an unreliable branch from the active context and preserves a concise corrective lesson. To learn this policy, LoongReflect combines two complementary signals through a look-ahead, extragradient-style coordination mechanism. A fast channel distills globally informed reflective behavior from a privileged teacher, with supervision restricted to reflection and backtracking tokens. A slow channel optimizes complete trajectories using outcome-based GRPO, aligning local control decisions with final task success. Experiments on multi-hop retrieval-augmented generation and mathematical reasoning benchmarks demonstrate consistent improvements over outcome-only reinforcement learning and self-distillation baselines.
Introduction
Reflection is not merely additional reasoning text; it is a control process over the trajectory itself. An effective reflection step should consolidate verified facts, identify missing evidence, diagnose branch-specific risks, and decide whether the current branch should be continued, revised, or abandoned. This capability is particularly important over long horizons, where localized errors—such as irrelevant retrievals, spurious entity associations, or stale memory updates—can enter the active context and influence many subsequent decisions. Without an explicit mechanism for isolating and correcting such errors, state contamination compounds as the trajectory grows, becoming a major obstacle to reliable agent behavior.
Despite its importance, learning effective reflection faces two fundamental challenges:
-
(C1) Learning-signal dilemma: The value of a reflective decision is mediated by many subsequent actions and is revealed only through the final task outcome. Outcome-based reinforcement learning therefore provides sparse, delayed, and weakly attributable supervision for reflection. Yet assigning reflection an explicit intermediate reward is equally problematic and prone to reward hacking, as it may encourage excessive or superficial reflection without improving task success.
-
(C2) Local–global perspective gap: Reflection is performed from the context of the current branch, whereas its true value depends on how that branch contributes to the complete trajectory. A local reflector cannot directly observe whether continuing, revising, or abandoning the branch will ultimately improve the outcome.
Existing approaches address only part of this problem. One line of work improves intermediate judgment through step-level verification or action evaluation. Another provides denser learning signals through privileged feedback or error-localized supervision. However, these approaches either treat reflection primarily as local judgment or rely on misaligned supervision signals that are not fully consistent with long-horizon outcomes.
To address C1&C2, we propose LoongReflect, a training framework that formulates reflection as a memory-control policy, designed for long-horizon setting. Rather than treating reflection as unconstrained verbal feedback, LoongReflect turns it into explicit, structured decisions over what the agent should retain, revise, or discard from its reasoning memory. Specifically, the agent operates over a reversible trajectory tree with an explicit active path and two control actions: and. The action consolidates verified facts, missing evidence, and branch-specific risks into working memory. When the current branch is deemed unreliable, the action removes its contaminated suffix from the active context, restores a validated state, and preserves a concise corrective lesson for subsequent decisions.
We train this memory-control policy through two complementary learning channels designed to resolve the two challenges jointly:
-
A fast channel distills reflective behavior from a privileged teacher, with supervision restricted to reflection and backtracking tokens. This token-level supervision provides the dense learning signal missing from outcome-based reinforcement learning, thereby addressing C1. Moreover, because the teacher observes the trajectory tree and terminal outcome, it can evaluate a local branch from a global perspective, thereby addressing C2. To prevent answer imitation, the teacher produces answer-masked feedback that focuses exclusively on state diagnosis, missing evidence, branch risks, and control decisions.
-
A slow channel applies outcome-based GRPO to complete trajectories, ensuring that the distilled reflective behavior remains aligned with final task success.
We coordinate the two channels through a look-ahead, extragradient-style update, in which the slow objective evaluates and calibrates the update direction proposed by the fast channel before the combined update is committed.
Contributions:
-
We formulate reflection as a memory-control policy for long-horizon agents and instantiate it through explicit and actions over a reversible trajectory tree.
-
We introduce a two-channel learning framework that combines answer-masked teacher distillation for dense, globally informed reflective supervision with outcome-based GRPO for trajectory-level alignment, coordinated through a look-ahead, extragradient-style update.
-
LoongReflect consistently outperforms outcome-only RL and self-distillation baselines on multi-hop RAG and mathematical benchmarks, demonstrating the contributions of structured reflection, reversible backtracking, and two-channel optimization.
Method
Problem Formulation
Given a task x, LoongReflect models a long-horizon agent as both an execution policy and a memory-control policy. In addition to reasoning, retrieval, and answer generation, the agent must determine which information accumulated during execution should be retained, revised, or discarded. We therefore equip the agent with a reversible trajectory state that can be explicitly modified through reflection and backtracking. At step t, the agent state is defined as zt = (x, Tt, Pt, mt), where Tt denotes the accumulated trajectory tree, Pt the active execution path, and mt the working memory compressed from that path.
Only the active path Pt and its compressed memory mt are exposed to subsequent generation. Inactive branches are excluded from the active context to prevent unreliable states from influencing future decisions, but remain in Tt for error diagnosis and reflective supervision. This separation enables the agent to recover from a corrupted branch without discarding the information needed to learn from it.
Reflection as Memory Control
To support reflection over intermediate states, LoongReflect represents agent memory as a reversible trajectory tree. At step t, the execution state consists of a trajectory tree Tt = (Vt, Et), an active path Pt, compressed memory mt, and archived branches Bt. The LLM context is serialized as ct = Serialize(x, Pt, mt). Thus, the model accesses only the task, active path, and compressed memory, while Tt and Bt remain internal states for future diagnosis and recovery.
Beyond task-solving actions, the agent may trigger reflection controls: actrl t ∈ Actrl =,. These actions form the memory-control mechanism:
-
diagnoses the active state by summarizing evidence, identifying missing information or faulty assumptions, and proposing the next control decision.
-
executes recovery when reflection determines that the current state is unreliable, rolling the active trajectory back to a trustworthy prefix and removing the unreliable suffix.
Concretely, produces a structured summary: rt = (ever t, jret t, qt risk, dctrl t), where ever t, jret t, qt risk, and dctrl t denote evidence, risk, return point, and control intent, respectively. The control intent determines whether to continue along the active path or invoke for recovery, making reflection a state-control diagnosis rather than a generic critique.
When backtracking, the transition is Pt+1 = Pj ⊕ uj:t, Bt+1 = Bt ∪ Pj+1:t, where Pj is the recovered active prefix, Pj+1:t is the invalidated suffix removed from the active context, and uj:t is a compact corrective update distilled from the current reflection, such as the decisive contradiction, the falsified assumption, or the constraint that the next attempt must satisfy.
Two-Channel Optimization for Reflection Learning
Fast Local Supervision: To provide dense supervision for reflection learning, LoongReflect introduces a fast channel over reflection control tokens. At step t, the student generates an active-state prefix with context ct = (x, Pt, mt). A teacher then constructs an answer-masked structured control label from the execution history: ht = Hϕ(x, T≤t, Pt, mt), where T≤t includes both the current active path and archived historical branches, and Hϕ denotes a privileged hint constructor implemented by either an auxiliary LLM or a rule-based feedback module. Its output shares the same schema as the reflect summary but is constructed from privileged access to the global execution record while masking the final answer, so the fast channel supervises local diagnosis and recovery rather than answer generation.
The student and teacher then evaluate the same continuation. Let yt,k denote the k-th generated token at step t. We define lt,k = log πθ(yt,k ct, yt,<k) and l̄t,k = log qθ̄(yt,k ct, ht, yt,<k), where qθ̄ is an EMA teacher of the policy. The teacher and student share the same on-policy prefix and continuation; the only additional information available to the teacher is the structured hint ht. Since the fast channel supervises only reflection-related spans, we define the token mask mref t,k = I(yt,k ∈ Span(,)), with normalization factor Z = Σt,k mt,k. Based on the teacher–student log-probability gap δt,k = lt,k − l̄t,k, we optimize a masked reverse-KL objective (a k3-style unbiased estimator) restricted to the reflection span: Lfast = (1/Z) Σt,k mref t,k min(exp(−δt,k) − 1 + δt,k, c), where the clipping constant c suppresses extreme token-ratio estimates. This objective is applied only to and tokens, providing dense local supervision for state diagnosis and recovery decisions.
Slow Global Optimization: To calibrate the global utility of reflection, LoongReflect introduces a slow channel that optimizes reflection decisions based on terminal trajectory outcomes. Unlike the fast channel, which supervises local state diagnosis, the slow channel evaluates whether these decisions improve final task success. Specifically, for each task x, we sample complete trajectories τ1, τ2,..., τG ∼ πθ(· x) with terminal rewards R1, R2,..., RG, where Ri is a task-level reward returned by the environment or an answer verifier. Given the trajectory group, we define the group-relative advantage as Ai = (Ri − mean g(Rg)) / std g(Rg). This normalizes rewards within the sampled group and assigns credit based on relative trajectory quality. To ensure clear credit assignment, the slow objective is applied only to policy-generated tokens in the final active execution. Let ξi denote the policy token sequence of trajectory τi, including execution and reflection control tokens (and), while excluding tool outputs, external observations, and controller updates. The slow channel is optimized by Lslow(θ) = −(1/G) Σi (1/ξi) Σk∈ξi min(ρi,k(θ)Âi, ρ̄i,k(θ)Âi) + β DKL(πθ ∥ πref), where ρi,k(θ) = πθ(yi,k ci,k) / πold(yi,k ci,k) and ρ̄i,k(θ) = clip(ρi,k(θ), 1 − ε, 1 + ε). Since Âi is derived from terminal outcomes, the slow channel rewards reflection only when it improves overall performance.
Look-Ahead Coordination: The fast and slow channels provide local and global signals, but their update directions may conflict. LoongReflect introduces look-ahead coordination, where the slow channel calibrates the fast update before optimization. Given θ, we obtain a provisional fast policy θe = U K fast(θ) by applying K inner fast-channel updates with estimated direction gf = (θe − θ)/α, where α is accumulated inner step size. We then evaluate trajectories under θe to derive slow-channel calibration direction gs = ∇θe Lslow(θe). Here, gs identifies fast-policy updates that improve final outcomes, while gf represents local reflection supervision. If gf and gs conflict, we remove the opposing component of gf along gs, yielding the calibrated fast direction gf LA = gf − (⟨gf, gs⟩ / ∥gs∥22) gs if ⟨gf, gs⟩ < 0, otherwise gf LA = gf. This operation preserves fast updates aligned with the global objective while removing only conflicting components. After calibration, we return to the original parameters θ and apply the fused update θ+ = θ − ηs gs − ηf gf LA, where ηs and ηf denote the outer-step sizes for the slow and calibrated fast directions.
Experiments
Experimental Setup
Training and Evaluation Benchmarks: Following AgenticRAG-R1, we train on HotpotQA and 2WikiMultiHopQA after filtering questions that can be answered without retrieval or with a single trivial lookup. We evaluate on seven retrieval-augmented QA benchmarks. 2WikiMultiHopQA and HotpotQA are treated as in-domain, while Bamboogle, FRAMES, MuSiQue, Natural Questions (NQ), and TriviaQA measure transfer to compositional, open-domain, and distribution-shifted questions. We additionally use MATH and GSM8K to examine transfer to non-retrieval multi-step reasoning.
SFT Data Construction: Before RL, we distill SFT trajectories from a locally deployed Qwen3-32B model. To elicit reflection during generation, whenever the teacher produces an incorrect answer before exhausting the maximum step budget, we replace answer action with and let the rollout continue until it reaches the correct answer within the budget. We then apply two-stage rejection sampling to retain long, informative rollouts. First, we sample LoongReflect trajectories and retain successful trajectories with at least one reflection and at least five interaction turns. Second, we resample the same questions with Search-R1 and keep if it succeeds while Search-R1 fails.
Models and Baselines: We use Qwen2.5-3B and 7B instruction models; the ablations use Qwen2.5-3B. We compare methods from four paradigms: no RAG (Base and CoT), naive RAG (FS-RAG and FL-RAG), agentic RAG (ReAct, IRCoT, TCRAG, and ReSearch), and RL-based agentic RAG (Search-R1, AEPO, ARPO, Mem1, and AgenticRAG-R1). For math reasoning, we further compare with RLSD.
Retrieval Setup and Metric: All retrieval methods use the same English Wikipedia snapshot dated November 1, 2023, together with the same retriever, top-k setting, context budget, tool-call budget, decoding configuration, and answer normalizer. We report answer-level F1 (%).
Main Result Analysis
To answer RQ1, Table 1 compares LoongReflect with representative baselines across the seven QA benchmarks. LoongReflect achieves the highest F1 on every benchmark with both model sizes. Its average F1 reaches 46.15 with Qwen2.5-3B and 49.21 with Qwen2.5-7B, exceeding the strongest baseline, AgenticRAG-R1, by 12.60 and 12.61 points, respectively. The improvement extends beyond the training distribution. On Qwen2.5-3B, LoongReflect raises the in-domain average from 38.46 to 52.09 and the out-of-domain average from 31.59 to 43.77 relative to AgenticRAG-R1. On Qwen2.5-7B, the corresponding averages increase from 41.75 to 53.86 and from 34.54 to 47.35. The consistent gains across model scales and evaluation regimes indicate that explicit reflection generalizes beyond the training benchmarks rather than merely fitting the in-domain tasks.
Transfer to Mathematical Reasoning
To answer RQ2, Table 3 evaluates whether the learned reflection policy transfers beyond retrieval. With Qwen2.5-3B, LoongReflect obtains 56.0 F1 on MATH and 82.4 F1 on GSM8K. It improves over AgenticRAG-R1 by 1.2 and 1.8 points, and over RLSD by 2.4 and 1.7 points, respectively. Although LoongReflect is motivated by long-horizon search, the gains on both datasets suggest that its learned diagnosis and recovery behavior also benefits multi-step reasoning without external retrieval.
Contributions of the Training Stages
To answer RQ3, Table 4 separates the effects of curated supervised fine-tuning and two-channel reinforcement learning. SFT improves average F1 over the raw instruction model by 4.43 points for Qwen2.5-3B and 3.54 points for Qwen2.5-7B, establishing an initial policy for reflection and backtracking. Applying two-channel RL yields a further 11.39-point gain for 3B and a 7.94-point gain for 7B. Overall, the complete pipeline improves over the raw backbones by 15.82 and 11.48 points. These results show that SFT provides a useful reflective prior, while globally calibrated two-channel optimization contributes the larger performance gain.
Ablation and Analysis
Component Ablation: To answer RQ4, Table 2 first examines the two memory-control actions. Removing reflection reduces average F1 from 46.15 to 30.84 (−15.31), while removing backtracking lowers it to 33.09 (−13.06). The former result shows the importance of diagnosing the active state; the latter confirms that diagnosis alone is insufficient without an explicit mechanism for discarding an unreliable suffix and resuming from a validated state. The optimization ablations are also consistently worse than the complete method. Removing slow outcome optimization, fast reflection distillation, or look-ahead coordination decreases average F1 by 7.04, 5.64, and 4.94 points, respectively. Thus, dense local teacher guidance and trajectory-level outcome optimization provide complementary supervision, while look-ahead coordination is necessary to suppress locally preferred updates that conflict with final task success.
Optimization Dynamics: Figure 3 tracks the two learning signals over 100 training steps. Task reward rises over the course of training, while the distillation reward moves upward from a strongly negative initial value toward zero. Their concurrent improvement indicates that the policy increasingly follows the teacher's local reflection guidance without sacrificing complete-trajectory outcomes, consistent with the intended division of labor between the fast and slow channels.
Hyperparameter Sensitivity: Figure 4 studies the number of inner fast updates K and the relative fast-direction weight w = ηf/ηs in the fused outer update. With w=1, F1 rises from 41.26 at K=1 to 46.15 at K=3, before declining to 44.71 at K=4. With K=3, w=1 outperforms both w=0.5 and w=2.0. Too few inner updates provide insufficient local adaptation, whereas an excessive relative fast weight can dominate the outcome-aligned slow direction. We therefore use K=3 and w=1 in the main experiments.
Conclusion
Long-horizon reflection faces the two challenges identified in the introduction: reflective decisions receive sparse and delayed outcome supervision, yet must be made from a branch-local context whose value depends on the complete trajectory. We introduced LoongReflect to address this learning-signal dilemma and local–global perspective gap by treating reflection as an explicit memory-control policy. A reversible trajectory tree, together with and actions, enables the agent to diagnose its active state, remove unreliable branches, and retain compact corrective information. For learning, the fast channel supplies dense, answer-masked supervision from a globally informed teacher, while the slow channel uses outcome-based GRPO to align local control with final task success. Look-ahead coordination reconciles these signals before the update is committed. Across seven retrieval-augmented QA benchmarks, LoongReflect consistently improves both Qwen2.5-3B and Qwen2.5-7B, including on out-of-domain tasks; with Qwen2.5-3B, it also transfers to two mathematical reasoning benchmarks. Training-stage and component ablations further verify the complementary contributions of structured reflection, reversible backtracking, local distillation, global outcome optimization, and their coordination.
Improvements for AI systems
Improvements to AI Systems:
-
Explicit Memory-Control Actions: Add
andas first-class actions in agent architectures. This enables the AI to (a) consolidate verified facts, missing evidence, and branch-specific risks into working memory, and (b) roll back to a validated state while preserving a corrective lesson—preventing error contamination from propagating through long trajectories. -
Reversible Trajectory Tree State: Implement a trajectory tree where only the active path and compressed memory are exposed to generation, while inactive branches remain archived for diagnosis. This allows the AI to recover from corrupted branches without losing the information needed to learn from mistakes, improving robustness in multi-step tasks.
-
Two-Channel Learning with Look-Ahead Coordination: Train agents using (a) a fast channel that distills globally informed reflective behavior from a privileged teacher (with answer-masked feedback to avoid answer imitation), and (b) a slow channel using outcome-based GRPO for trajectory-level alignment. Coordinate these via extragradient-style updates that remove conflicting components of the fast update against the slow objective, ensuring local reflection decisions improve final task success.
-
Answer-Masked Teacher Distillation: When providing dense supervision, mask the final answer and only supervise reflection/backtracking tokens. This prevents the student from copying answers while still receiving dense, globally informed guidance on state diagnosis, missing evidence, and branch risks—solving the sparse-reward problem without reward hacking.
-
Token-Level Reverse-KL Objective for Reflection Spans: Apply a clipped reverse-KL loss only to
andtokens, using the teacher–student log-probability gap. This provides stable, dense local supervision for control decisions without over-regularizing execution tokens. -
Group-Relative Advantage for Credit Assignment: Use group-relative advantages (normalized within sampled trajectory groups) in the slow channel, so the AI assigns credit to reflection decisions based on relative trajectory quality—rewarding reflection only when it improves outcomes, not merely for appearing.
-
Curated SFT Prior via Rejection Sampling: Before RL, distill SFT trajectories by (a) replacing incorrect answers with `` and continuing rollouts until success, and (b) retaining successful trajectories that beat a baseline (e.g., Search-R1) on the same questions. This provides a strong reflective prior, reducing RL exploration burden.
What the Improved AI System Can Do:
-
Self-Correct Long-Horizon Reasoning: Detect and discard unreliable intermediate states (e.g., irrelevant retrievals, spurious entity associations) before they contaminate subsequent decisions, improving accuracy on multi-hop QA and complex math problems.
-
Learn from Global Outcomes with Local Dense Feedback: Balance immediate reflective guidance (from a privileged teacher) with long-term task success (via outcome-based RL), avoiding both sparse-reward blindness and reward hacking from naive intermediate rewards.
-
Transfer Reflection Skills Across Domains: Apply learned diagnosis and recovery behavior beyond training distribution—e.g., from retrieval-augmented QA to mathematical reasoning without external tools—demonstrating generalizable reflection.
-
Efficiently Recover from Errors: Instead of restarting or continuing with corrupted context, the AI can roll back to a validated prefix, apply a corrective lesson, and resume—reducing wasted computation and improving success rates on tasks with long horizons.
-
Avoid Answer Imitation During Training: Learn reflective control without copying teacher answers, ensuring the policy generalizes to novel tasks where the teacher’s answer is unavailable.
-
Achieve Consistent Performance Gains: Outperform outcome-only RL and self-distillation baselines by 12+ F1 points on average across seven QA benchmarks, with further gains on out-of-domain tasks and math reasoning, all while using smaller models (3B/7B) than typical teacher models.
Abstract
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. Learning effective reflection, however, is challenging because reflection is performed locally within the current branch, whereas its utility can only be determined by its contribution to the final trajectory outcome. This local-global mismatch makes outcome-based reinforcement learning provide only local, sparse and delayed supervision for reflective decisions. To solve these, we propose LoongReflect, a training framework that formulates reflection as a memory-control policy. The agent operates over a reversible trajectory tree using explicit reflect and backtrack actions. Reflection consolidates verified facts, missing evidence, and branch-specific risks into working memory, while backtracking removes an unreliable branch from the active context and preserves a concise corrective lesson. To learn this policy, LoongReflect combines two complementary signals through a look-ahead, extragradient-style coordination mechanism. A fast channel distills globally informed reflective behavior from a privileged teacher, with supervision restricted to reflection and backtracking tokens. A slow channel optimizes complete trajectories using outcome-based GRPO, aligning local control decisions with final task success. Experiments on multi-hop retrieval-augmented generation and mathematical reasoning benchmarks demonstrate consistent improvements over outcome-only reinforcement learning and self-distillation baselines.
Sources
- Reinforcement Learning for Long-Horizon Interactive LLM Agents
- Training Verifiers to Solve Math Word Problems
- Agentic Entropy-Balanced Policy Optimization
- Agentic Reinforced Policy Optimization
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- Measuring Mathematical Problem Solving With the MATH Dataset
- Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Generalization through Memorization: Nearest Neighbor Language Models
- Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning
- Agentic Critical Training
- Qwen2.5 Technical Report
- Self-Distilled RLVR
- ReAct: Synergizing Reasoning and Acting in Language Models
- StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Where LLM Agents Fail and How They can Learn From Failures
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks