SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
cs.CL, cs.AI, cs.LG, math.SP
Submitted: 2025-11-11
Updated: 2025-11-11
Journal ref: 2025 IEEE International Conference on Big Data (BigData)
DOI: 10.1109/BigData66926.2025.11402347
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) adapted to financial domains often suffer from catastrophic forgetting of general reasoning capabilities essential for customer interactions and complex financial
Terminology
Abstract
Large language models (LLMs) adapted to financial domains often suffer from catastrophic forgetting of general reasoning capabilities essential for customer interactions and complex financial analysis. We introduce Selective Parameter Evaluation and Restoration via Model Merging (SPEAR-MM), a practical framework that preserves critical capabilities while enabling domain adaptation. Our method approximates layer-wise impact on external benchmarks through post-hoc analysis, then selectively freezes or restores transformer layers via spherical interpolation merging. Applied to LLaMA-3.1-8B for financial tasks, SPEAR-MM achieves 91.2% retention of general capabilities versus 69.7% for standard continual pretraining, while maintaining 94% of domain adaptation gains. The approach provides interpretable trade-off control and reduces computational costs by 90% crucial for resource-constrained financial institutions.
Sources
- Training Verifiers to Solve Math Word Problems
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Instruction-Following Evaluation for Large Language Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Automated Driving Without Ethics: Meaning, Design and Real-World Implementation
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
- Spectrum: Targeted Training on Signal to Noise Ratio
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering