Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-09-08
Comments: 12 pages
Code: https://github.com/THUDM/slime
License: http://creativecommons.org/licenses/by/4.0/
The gist: Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks.
Terminology
Abstract
Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synthetic tool-use trajectories, optimized with a Muon + Adam hybrid. The recipe is deliberately conservative and deliberately controlled: 626 trajectories, a single epoch, a low learning rate, and a KL anchor to the frozen base. The model shows substantial gains over the previous default model for Writer Agent, and compares favorably with several recent models on public benchmarks, scoring the highest on BFCL Core at 0.785 and posts the highest six-benchmark mean of the cohort. Furthermore, the model has shown itself to be competitive or leading relative to comparators in our bias and safety evaluations.
Sources
- Anchored Supervised Fine-Tuning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- DeepSeek-V3 Technical Report
- Self-Distillation Enables Continual Learning
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
- Muon is Scalable for LLM Training
- GLM-5: from Vibe Coding to Agentic Engineering
- FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering