LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- GPT-4 Technical Report
- Unlocking the Working Memory of Large Language Models for Latent Reasoning
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- Titans: Learning to Memorize at Test Time
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
- Soft Tokens, Hard Truths
- Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Training Verifiers to Solve Math Word Problems
- Universal Transformers
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
- MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- Group-in-Group Policy Optimization for LLM Agent Training
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Think before you speak: Training Language Models With Pause Tokens
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Training Large Language Models to Reason in a Continuous Latent Space
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering