LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu
University of Oklahoma · Imperial College London · University of Michigan · Tencent · University of Edinburgh · University of Southern California · Beijing Normal University · Beijing Normal-Hong Kong Baptist University
cs.AI
Submitted: 2026-08-07
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 61/100
The gist: LoCA (Local Credit Assignment) is introduced as "a two-stage method for small-shift adaptation" designed to address the limitations of parameter-efficient post-training, which "still requires
Terminology
Summary
LoCA (Local Credit Assignment) is introduced as a two-stage method for small-shift adaptation
designed to address the limitations of parameter-efficient post-training, which still requires repeated end-to-end backpropagation through the frozen backbone.
This repeated backpropagation necessitates backward-capable hardware
and requires the system to store or recompute activations.
LoCA aims to replace this repeated backward chain... by a one-time calibration.
The LoCA Framework
The method operates in two distinct stages:
-
Calibration:
One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction.
This process fits alow-rank feedback operator
F for each block, whichmaps the top-layer error e to an estimate of g
(the hidden-state gradient). The paper notes that while random feedback provides little useful direction, LoCAfits F to the frozen model
once. The resulting F is arank- k approximation
(with k=8 in experiments) of the regularized fit. -
Adaptation:
LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required.
During this stage, the target for each block is defined as tau = h - eta F e, where e is thetop-layer error from the frozen head.
Mathematical Implementation
For each block, LoCA solves a strictly convex
ridge regression problem to find the adapter B. The objective is to minimize the error between the adapted hidden state and the local target: sum n B p - rho squared + lambda B 2 F. This problem has the unique closed-form minimizer B = C (G + lambda I r)-1,
where G and C are sufficient statistics that are additive over batches.
Consequently, the block parameter uses no gradient optimizer, learning-rate schedule, or optimizer state.
To ensure the method is robust across different model scales, LoCA employs scale normalization: we reduce this sensitivity by normalizing the target correction to the residual-stream scale,
using the formula = eta times RMS(h-1) over RMS(F e). This allows a shared scale-normalized candidate set
to be reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B.
Empirical Results
LoCA was evaluated on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B.
The findings include:
-
Performance:
In 16 of 25 reported task–scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run.
-
Resource Efficiency:
Its measured full-run GPU peak, including calibration, is 26–29% lower than LoRA’s. After calibration, its CPU steady-state memory is 36–52% lower and its per-pass time is 43–48% lower.
-
Comparison to MeZO: While
Adapter-MeZO uses the least memory,
it requires10 cubed – 10 4 perturbation steps,
whereasLoCA uses at most 40 outer iterations.
-
Generalization: The scale-normalized approach demonstrated that
the candidate range transfers across the tested models,
showing successful recovery on the SmolLM2-1.7B model family.
In summary, LoCA amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical.
Improvements for AI systems
1. Inference-Only Edge Adaptation Engines
These systems enable real-time, on-device fine-tuning on hardware that lacks backward-pass capabilities (such as mobile NPUs or specialized inference-only chips). By replacing gradient-based updates with forward-only, closed-form ridge regression, these engines allow personalized AI to learn from user interactions directly on local hardware without requiring the full computational overhead of a training-capable GPU.
2. Low-Latency Streaming Learning Systems
These systems can adapt to non-stationary data distributions—such as live conversational shifts or continuous sensor telemetry—with 43–48% lower per-pass latency than LoRA. This allows an AI to undergo on-the-fly
adaptation, shifting its behavior or knowledge base instantaneously without the computational pauses or training freezes
typically required for backpropagation.
3. Memory-Optimized Large-Context Training Pipelines
By reducing GPU peak memory by 28% and CPU memory by up to 52%, these systems can utilize the reclaimed resources to significantly increase batch sizes or extend context window lengths during the adaptation phase. This enables the fine-tuning of massive models on consumer-grade hardware that would otherwise be unable to store the activations required for standard LoRA.
4. Cross-Scale Knowledge Transfer Frameworks
Using the scale-normalized calibration method, these systems can transfer learned feedback operators from large-scale teacher
models (e.g., Qwen2.5-14B) to smaller student
models (e.g., SmolLM2). This allows small, efficient models to achieve high-performance task adaptation by inheriting the structural credit-assignment logic of much larger models through a single calibration step.
5. Optimizer-Free Autonomous Agents
These systems eliminate the need for optimizer states (such as Adam's momentum and variance tensors) and complex learning-rate schedules. This results in highly stable, memory-efficient continuous learning for autonomous agents, preventing the training instability and catastrophic forgetting
often caused by poorly tuned hyperparameters in long-running, real-world deployments.
Abstract
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task--scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run. Its measured full-run GPU peak, including calibration, is 26--29% lower than LoRA's. After calibration, its CPU steady-state memory is 36--52% lower and its per-pass time is 43--48% lower. A shared scale-normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical. The code associated with this paper is available here.
Sources
- The Forward-Forward Algorithm: Some Preliminary Investigations
- Training Deep Nets with Sublinear Memory Cost
- Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
- Qwen2.5 Technical Report
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection