LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
cs.SE, cs.AI, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- Evaluating Large Language Models Trained on Code
- The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
- Quantifying Variance in Evaluation Benchmarks
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Dynamic Stability of LLM-Generated Code
- Effectively Controlling Reasoning Models through Thinking Intervention
- Logics-STEM: Empowering LLM Reasoning via Failure-Driven Post-Training and Document Knowledge Enhancement
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties