Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
cs.CL, cs.LG
Submitted: 2025-12-16
Updated: 2026-08-25
Code: https://github.com/Mega4alik/ollm
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Establishing Task Scaling Laws via Compute-Efficient Model Ladders
- CST: Calibration Side-Tuning for Parameter and Memory Efficient Transfer Learning
- Training Deep Nets with Sublinear Memory Cost
- Language models scale reliably with over-training and on downstream tasks
- Training Verifiers to Solve Math Word Problems
- OpenThoughts: Data Recipes for Reasoning Models
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- Scaling Laws for Neural Language Models
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
- Skip a Layer or Loop it? Test-Time Depth Adaptation of Pretrained LLMs
- Reducing Activation Recomputation in Large Transformer Models
- s1: Simple test-time scaling
- EvoLM: In Search of Lost Language Model Training Dynamics
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering