LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

arXiv:2608.03020 · cs.AI · Submitted 2026-08-07 · Read on arXiv

Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu

University of Oklahoma · Imperial College London · University of Michigan · Tencent · University of Edinburgh · University of Southern California · Beijing Normal University · Beijing Normal-Hong Kong Baptist University

cs.AI

Submitted: 2026-08-07

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 61/100

The gist: LoCA (Local Credit Assignment) is introduced as "a two-stage method for small-shift adaptation" designed to address the limitations of parameter-efficient post-training, which "still requires

Terminology

Summary

LoCA (Local Credit Assignment) is introduced as a two-stage method for small-shift adaptation designed to address the limitations of parameter-efficient post-training, which still requires repeated end-to-end backpropagation through the frozen backbone. This repeated backpropagation necessitates backward-capable hardware and requires the system to store or recompute activations. LoCA aims to replace this repeated backward chain... by a one-time calibration.

The LoCA Framework

The method operates in two distinct stages:

  1. Calibration: One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. This process fits a low-rank feedback operator F for each block, which maps the top-layer error e to an estimate of g (the hidden-state gradient). The paper notes that while random feedback provides little useful direction, LoCA fits F to the frozen model once. The resulting F is a rank- k approximation (with k=8 in experiments) of the regularized fit.

  2. Adaptation: LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. During this stage, the target for each block is defined as tau = h - eta F e, where e is the top-layer error from the frozen head.

Mathematical Implementation

For each block, LoCA solves a strictly convex ridge regression problem to find the adapter B. The objective is to minimize the error between the adapted hidden state and the local target: sum n B p - rho squared + lambda B 2 F. This problem has the unique closed-form minimizer B = C (G + lambda I r)-1, where G and C are sufficient statistics that are additive over batches. Consequently, the block parameter uses no gradient optimizer, learning-rate schedule, or optimizer state.

To ensure the method is robust across different model scales, LoCA employs scale normalization: we reduce this sensitivity by normalizing the target correction to the residual-stream scale, using the formula = eta times RMS(h-1) over RMS(F e). This allows a shared scale-normalized candidate set to be reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B.

Empirical Results

LoCA was evaluated on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. The findings include:

  • Performance: In 16 of 25 reported task–scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run.

  • Resource Efficiency: Its measured full-run GPU peak, including calibration, is 26–29% lower than LoRA’s. After calibration, its CPU steady-state memory is 36–52% lower and its per-pass time is 43–48% lower.

  • Comparison to MeZO: While Adapter-MeZO uses the least memory, it requires 10 cubed – 10 4 perturbation steps, whereas LoCA uses at most 40 outer iterations.

  • Generalization: The scale-normalized approach demonstrated that the candidate range transfers across the tested models, showing successful recovery on the SmolLM2-1.7B model family.

In summary, LoCA amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical.

Improvements for AI systems

1. Inference-Only Edge Adaptation Engines

These systems enable real-time, on-device fine-tuning on hardware that lacks backward-pass capabilities (such as mobile NPUs or specialized inference-only chips). By replacing gradient-based updates with forward-only, closed-form ridge regression, these engines allow personalized AI to learn from user interactions directly on local hardware without requiring the full computational overhead of a training-capable GPU.

2. Low-Latency Streaming Learning Systems

These systems can adapt to non-stationary data distributions—such as live conversational shifts or continuous sensor telemetry—with 43–48% lower per-pass latency than LoRA. This allows an AI to undergo on-the-fly adaptation, shifting its behavior or knowledge base instantaneously without the computational pauses or training freezes typically required for backpropagation.

3. Memory-Optimized Large-Context Training Pipelines

By reducing GPU peak memory by 28% and CPU memory by up to 52%, these systems can utilize the reclaimed resources to significantly increase batch sizes or extend context window lengths during the adaptation phase. This enables the fine-tuning of massive models on consumer-grade hardware that would otherwise be unable to store the activations required for standard LoRA.

4. Cross-Scale Knowledge Transfer Frameworks

Using the scale-normalized calibration method, these systems can transfer learned feedback operators from large-scale teacher models (e.g., Qwen2.5-14B) to smaller student models (e.g., SmolLM2). This allows small, efficient models to achieve high-performance task adaptation by inheriting the structural credit-assignment logic of much larger models through a single calibration step.

5. Optimizer-Free Autonomous Agents

These systems eliminate the need for optimizer states (such as Adam's momentum and variance tensors) and complex learning-rate schedules. This results in highly stable, memory-efficient continuous learning for autonomous agents, preventing the training instability and catastrophic forgetting often caused by poorly tuned hyperparameters in long-running, real-world deployments.

Abstract

Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen backbone. Every adaptation step therefore needs backward-capable hardware and must store or recompute activations. We ask whether this repeated backward chain can be replaced by a one-time calibration. We introduce Local Credit Assignment (LoCA), a two-stage method for small-shift adaptation. One probe backward pass fits a low-rank map at each transformer block from the final prediction error to a local hidden-state correction. LoCA then reuses these maps to form blockwise regression targets from forward activations and fits low-rank adapters with closed-form ridge solves. No further backbone backward pass is required. We evaluate LoCA on five discriminative benchmarks with Qwen2.5 models from 0.5B to 14B. In 16 of 25 reported task--scale comparisons, LoCA yields lower evaluation cross-entropy than the corresponding LoRA run. Its measured full-run GPU peak, including calibration, is 26--29% lower than LoRA's. After calibration, its CPU steady-state memory is 36--52% lower and its per-pass time is 43--48% lower. A shared scale-normalized candidate set is reused across all tested Qwen2.5 sizes and on SmolLM2-1.7B. LoCA thus amortizes global credit assignment into one calibration and enables later forward-only tuning when repeated backpropagation is impractical. The code associated with this paper is available here.

Sources

Related papers