Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Listening to the Wise Few".
Tom: This research introduces novel scoring mechanisms, specifically the Query-Key Score (QK-score) and Attention Score, derived from select-and-copy attention heads within Large Language Models (LLMs),
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, this paper is titled "Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models," and the authors are Tulchinskii, Kushnareva, Kuznetsov, Voznyuk, Andriiainen, Piontkovskaya, Burnaev. It sounds a bit academic at first because it’s deep into the mechanics of how these models operate.
Jane: It definitely has a very technical title; "Listening to the Wise Few" suggests they are focusing on a small subset of the model's internal workings rather than trying to analyze every single component equally.
Lu: That aligns with what they’re doing, Jane; they are looking for those specific attention heads that actually do the selection process effectively, instead of just looking at all the layers blindly. They are essentially tuning into the most relevant signals.
Meng: Tuning into a subset sounds useful if we need to optimize performance on specific tasks without having to overhaul the entire model architecture, which is good for iterative improvements. But what makes this "wise few" different from just picking randomly?
Lalam: The difference is that these heads are performing a very specific action—selecting and copying—which suggests they hold concentrated knowledge about how to correctly map a question to an answer option. This kind of focused understanding could really improve the reliability of our AI systems.
The paper's summary: Tom: Okay, so the core summary is that standard evaluation methods have limitations because even if a model knows the right answer, it might struggle with the rigid format of multiple-choice questions, and this paper introduces new scores to reveal that underlying knowledge.
Jane: They introduce two main scoring metrics: the Query-Key Score and an Attention Score, which are derived from how specific attention heads interact with the query and key representations during the process.
Lu: The Query-Key Score is calculated by taking a dot product between the query vector and the last key vector corresponding to an option-representative token, but they explicitly state that this calculation doesn't apply positional transformation to account for shifts in position.
Meng: So, they’re focusing on the immediate interaction at that specific moment rather than trying to predict what happens much later in the model's processing pipeline? That’s a very direct approach.
Lalam: And the Attention Score focuses purely on the attention weights themselves, specifically looking at how those weights connect between the answer option tokens and both the query and prompt tokens. It's a way of estimating probabilities for each option based on these internal connections.
The paper's improvements: Tom: The paper suggests that identifying these select-and-copy heads allows us to significantly boost performance because these heads are present across various model sizes, from seven billion to seventy billion parameters.
Jane: They found that using the best-performing heads selected based on validation set accuracy leads to substantial gains; they report up to a sixteen percent gain for LLaMA2-7B and up to a ten percent gain for larger models on popular multiple-choice question answering benchmarks.
Lu: What's really interesting is that they noted these select-and-copy heads process MCQA tasks more effectively in the middle layers, but then those same representations tend to get revised in the later layers, which actually hurts performance if we don't account for that.
Meng: That layering observation is important; it suggests we might need a strategy that leverages these early, effective representations while carefully managing the information flow as it moves through deeper parts of the model.
Lalam: Furthermore, they proposed an algorithm to find stable heads without needing a labeled validation set by scoring them based on things like the sum of average attention weights to all symbols after options and the frequency with which a head chooses any option other than the most popular one.
Conclusion: Tom: So, wrapping up, the main implication is that by isolating these select-and-copy heads using QK-score and Attention Score, we can unlock latent correct answers in multiple-choice tasks that standard formats often hide.
Jane: It really shows us a 'white-box' technique where we use the internal representations of LLMs directly to extract solutions for given tasks, rather than relying solely on external output logits.
Lu: The research suggests that these specialized heads aren't just task-specific quirks; they are relatively stable across different datasets and model scales, which opens up new avenues for general understanding of how LLMs function.
Meng: From a practical standpoint, this means we can potentially fine-tune or guide models by activating the best performing heads we’ve identified, which could give us a way to improve performance without needing massive retraining cycles.
Lalam: This work on "Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models" gives us a clear path toward understanding LLMs better and improving their accuracy on these structured tasks. We are so excited about what this means for the future of AI applications.
Skolkovo Institute of Science and Technology · AI Foundation and Algorithm Lab · Moscow Institute of Physics and Technology · CNRS, Universite Paris Citée de l'Université de Paris
cs.CL, cs.LG
Submitted: 2024-10-03
Updated: 2026-09-30
Project page: https://wilburone.github.io/cosmos/14
Importance score: 78/100
The gist: This research introduces novel scoring mechanisms, specifically the Query-Key Score (QK-score) and Attention Score, derived from select-and-copy attention heads within Large Language Models (LLMs),
Key concepts
- Select-and-Copy Heads
- These are specific attention heads within LLMs that perform the core task of answering multiple-choice questions: they compute semantic representations, select the best option using a query-key mechanism, and then copy that chosen option. The paper identifies these as crucial for performance.
- QK-score
- This metric measures alignment by calculating the dot product between a query vector and the last key vector corresponding to an answer token. It is designed to estimate the probability of selecting an option, specifically avoiding positional transformations to maintain accuracy.
- Attention Score
- This score analyzes attention weights focusing on connections between answer option tokens and the question or prompt tokens. By maximizing this score, the method estimates which option is most likely correct based on how the model attends to different parts of the input.
Terminology
Summary
This research introduces novel scoring mechanisms, specifically the Query-Key Score (QK-score) and Attention Score, derived from select-and-copy attention heads within Large Language Models (LLMs), to unlock latent correct answers in Multiple-Choice Question Answering (MCQA) tasks. These methods aim to mitigate the limitations of standard evaluation formats by focusing on internal model representations, demonstrating significant performance gains across various benchmarks and synthetic datasets.
The Core Concept: Select-and-Copy Heads
The paper identifies a principal elementary algorithmic operation performed by pretrained Transformer models when answering multiple-choice questions: first, computing a representation of semantic information from the question and options within specific heads, followed by selecting the most appropriate option using a query-key alignment mechanism, and finally copying and outputting the option. The authors propose identifying select-and-copy heads
that perform this selection operation on aggregated embeddings of possible answers. They suggest these heads are present in all models examined (7B to 70B parameters) and are crucial for performance, noting that they process MCQA tasks more effectively in the middle layers but tend to revise this information in later layers.
Novel Scoring Mechanisms: QK-score and Attention Score
The method introduces two key scoring metrics derived from these heads. The Query-Key Score (QK-score) is calculated as a dot product of the query vector and the last key vector corresponding to an option-representative token, explicitly stated as not applying positional transformation to mitigate relative position shift effects. The Attention Score is based on attention weights, specifically focusing on the weights between the answer option tokens and the query/prompt tokens. These scores are used to estimate probabilities for each option: we take the option, for which the score gives maximum.
Experimental Methodology and Performance Gains
The study evaluates these scores across four challenging real-world MCQA datasets: MMLU, CosmosQA, HellaSwag, and HaluDialogue. The authors demonstrate that using the best-performing heads selected based on validation set accuracy yields substantial improvements. Specifically, they report up to 16% gain for LLaMA2-7B and up to 10% for larger models on popular MCQA benchmarks.
Furthermore, on a simple synthetic dataset (SSD) where the model explicitly knows the right answer, accuracy increases by almost 60%, achieving nearly perfect accuracy,
proving efficiency in mitigating format limitations.
Identification of Universal and Stable Heads
The research moves beyond task-specific heads to find more robust mechanisms. They identify universal heads
that perform well across multiple datasets and model scales. By analyzing the performance of these best heads, they determine that the most robust heads are (14, 24) and (14, 20)
in LLaMA2-7B for zero-shot setups. They also propose an algorithm to find stable heads without a labeled validation set by scoring them based on the sum of average attention weights to all n symbols after options on this head
and a frequency of 'choosing' any option aside of the most popular one.
Insights into Attention Patterns and Selection Bias
Analysis of attention patterns confirms that select-and-copy heads concentrate their attention on option-representative tokens, namely newline symbols after options, with the highest weight on the correct option,
which is expected. Furthermore, they observe complementary biases between different best heads; for instance, one head might be biased toward options 'A' and 'D', while another is biased toward 'B' and 'C'. This investigation provides a step towards better understanding how the LLMs work in general.
Conclusion
The paper concludes that the introduction of QK-score and Attention Score significantly improves MCQA performance, revealing a subset of attention heads—the select-and-copy heads—that are relatively stable across datasets and model scales. These specialized heads have the potential to deepen our understanding of LLMs’ capabilities not only for MCQA but for other reasoning tasks as well.
Improvements for AI systems
Based on this scientific paper, here are specific improvements that can be made to AI systems:
-
Develop a novel,
white-box
internal mechanism for multiple-choice question answering (MCQA) by explicitly identifying and utilizingselect-and-copy heads
within LLMs. -
Implement the proposed QK-score and Attention Score as novel scoring metrics derived from specific attention head interactions to better capture knowledge extraction compared to standard output logits or simple attention weights.
-
Enhance LLM performance on MCQA benchmarks by selectively activating or weighting the outputs of identified high-performing select-and-copy heads, leading to up to a 16% gain for LLaMA2-7B and 10% for larger models, especially in zero-shot scenarios.
-
Achieve near-perfect accuracy (up to 60% improvement) on synthetic MCQA datasets designed specifically to test the model's ability to adhere to rigid multiple-choice formats, effectively mitigating the limitations of LLMs in following strict output constraints.
-
Improve robustness against adversarial question formatting, including option permutations, renaming options, and the addition of uncertainty options (
I don't know
,None of the above
), by using QK-score as a more stable metric than simple baseline methods (like PriDe). -
Create a method for automatically discovering universal select-and-copy heads across different LLM scales (7B to 70B) and various MCQA datasets, allowing for efficient adaptation without requiring labeled validation data.
-
Improve model interpretability by analyzing the specific attention patterns of these selected heads to understand which semantic information is being
copied
from the query/key representations into the final answer selection process. -
Develop a method to diagnose
selection bias
in LLMs—where they tend to choose specific options rather than objectively correct ones—by identifying and compensating for biased attention patterns, especially in zero-shot settings.
Sources
- The Llama 3 Herd of Models
- Language Models (Mostly) Know What They Know
- Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
- GPT-4 Technical Report
- A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Benchmarking LLMs via Uncertainty Quantification
- Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
- Attention Heads of Large Language Models: A Survey
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering