ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding
cs.CL, cs.AI
Submitted: 2026-04-20
Updated: 2026-08-29
Terminology
Sources
- Towards better understanding of gradient-based attribution methods for Deep Neural Networks
- We Should Chart an Atlas of All the World's Models
- Qwen Technical Report
- Transferring Backdoors between Large Language Models by Knowledge Distillation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Language Models are General-Purpose Interfaces
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
- Physical Symmetries Embedded in Neural Networks
- Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
- RouteLLM: Learning to Route LLMs with Preference Data
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Mapping 1,000+ Language Models via the Log-Likelihood Vector
- Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- SmoothGrad: removing noise by adding noise
- Qwen2 Technical Report
- Protect Your Prompts: Protocols for IP Protection in LLM Applications
- Gradient based Feature Attribution in Explainable AI: A Technical Review
- Qwen3 Technical Report
- PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering