MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
cs.CL, cs.AI, cs.MA, cs.SE
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- Controlling Tool Use with Heading-Specific Activation Steering
- Diagnosing LLM Arbitration Behavior over Pre-evidence Epistemic States in RAG-based Fact-Checking
- Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
- Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts
- From Context-Aware to Conflict-Aware: Generalizing Contrastive Decoding for Knowledge Conflict in LLMs
- Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
- Trust or Abstain? A Self-Aware RAG Approach
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering