Tracekit: Tamper-Evident Intent-Reasoning-Action Auditing for Autonomous Coding Agents
cs.CE, cs.CR
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/LoopGlitch26/Tracekit
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Ctrl-Z: Controlling AI Agents via Resampling
- Reasoning Models Don't Always Say What They Think
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Defeating Prompt Injections by Design
- AI Control: Improving Safety Despite Intentional Subversion
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Measuring Faithfulness in Chain-of-Thought Reasoning
- NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Constrained Sensing and Reliable State Estimation with Shallow Recurrent Decoders on a TRIGA Mark II Reactor
- Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval
- Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad
- Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness
- RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
- Wildfire Suppression: Complexity, Models, and Instances