Rufus-Air: An Open LLM Post-Training Recipe
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-24
Updated: 2026-09-25
Code: https://github.com/aws-samples/sample-e2b-on-aws
Terminology
Sources
- OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Llama-Nemotron: Efficient Reasoning Models
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
- Evaluating Large Language Models Trained on Code
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
- Learning to Reason at the Frontier of Learnability
- Mach-Mind-4-Flash Technical Report
- Endless Terminals: Scaling RL Environments for Terminal Agents
- From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
- Scaling Laws for Reward Model Overoptimization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering