Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER".
Jane: The paper was written by Ahmed Ewais, Ahmed Hashish, Amr Ali and WitnessAI from WitnessAI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Core Mechanism: Tom: We just established that "Just Pass Twice" is a game-changer for how we see LLMs performing Named Entity Recognition. So, let’s look deeper into the core mechanism of this paper.
Jane: The paper explains that the biggest obstacle to accurate token classification in LLMs is their causal attention, which only allows tokens to look at what came before them.
Lu: That lack of access to future context is exactly what makes distinguishing between a person named Paris and a city called Paris so difficult for the model.
Meng: So, when the model sees "Paris released a new album," if it's looking at "Paris," it has no idea that "album" will clarify that this person isn't a location yet.
Lalam: And then, JPT solves this by duplicating the input sequence, which is quite clever because it’ allows the the second pass of each token to see the whole picture.
Tom: It's like giving the model a full view of a movie while it’s only watching half an hour in real time.
Jane: That's a perfect analogy, Tom; in this research, that second pass lets all future context flow back into the first pass.
Lu: The authors are essentially bypassing the causal mask without requiring any massive structural changes to the backbone of duplicating it, which is incredibly efficient.
Meng: That efficiency is critical for me; we aren't adding entirely new layers, just utilizing existing parallel computation in a way that speeds up the process dramatically.
Lalam: This allows AI to perform this complex disambiguation task with a clear vision of the entire input, making it feel more human.
Method Details: Tom: So we know how it works by duplicating the input, but now let’s talk about the specific improvements in "Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER."
Jane: The authors didn't just rely on the duplication; they also introduced definition-guided entity typing to make it work without retraining.
Lu: That means we are leveraging natural language descriptions of what an entity should be, not just memorizing fixed labels like "PERSON."
Meng: For me, that makes zero-shot generalization much more robust because we aren't limited to a vocabulary the model has seen before.
Lalam: It’s about teaching the AI to understand *semantics* of definitions rather than just asking it to recall a specific label.
Tom: The paper is using lightweight LoRA adapters, which is another key technical detail that helps keep the system efficient while performing this classification.
Jane: And these adapters are trained on a huge Wikipedia-derived dataset, ensuring they capture diverse examples for the model's learning process.
Lu: It’s a combination of keeping the massive knowledge of the LLM and then adding targeted, trainable expertise through LoRA adapters.
Meng: From an engineering view, that approach is smart because we are only updating a tiny fraction of parameters while achieving significant performance gains.
Lalam: This allows for extremely flexible deployment, where we can define new entity types on the fly without massive retraining efforts.
Improvements and Impact: Tom: We've covered the mechanism and now we need to look at the actual improvements made in "Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER."
Jane: The results are truly remarkable, achieving a significant F1 score increase over state-of-the-art methods.
Lu: The authors claim an average of +seven point nine F1 across various benchmarks, which is a massive leap in accuracy for this type of task.
Meng: And what it’s faster than generative methods is even more impressive; they are saying twenty times faster, which is a huge efficiency win for the industry.
Lalam: This speed difference means we can integrate this powerful AI into applications that need real-time performance, like live data stream processing.
Tom: The authors highlight that this isn't just one specific model, but consistent improvements across nineteen out of twenty extended benchmarks.
Jane: That tells us the system is incredibly robust and doesn't just work on one type of text; it’ handles diverse domains like medical or social media.
Lu: The combination of this performance stability with the dual-channel definition injection really speaks to the power, achieving a level of generalization that was previously unattainable.
Meng: It feels like we are finally bridging the gap between highly accurate but slow generative AI and fast but less capable discriminative models.
Lalam: We' can deploy an AI that is both incredibly smart and incredibly quick, which is a massive cultural shift for information processing.
Conclusion: Tom: So, we have covered the mechanics, the specific technical improvements, and the impressive results of "Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER."
Jane: It's clear that this is a very strong contender for it solves those fundamental problems that generative models have.
Lu: The ability to leverage massive world knowledge while achieving this speed is genuinely exciting, and I can't wait to see how far this technology takes us.
Meng: From an engineering standpoint, the fact that we are using only a fraction of the backbone parameters means deployment will be much more resource-efficient.
Lalam: The ability to define and recognize entities based on human language definitions allows for a level of understanding in AI that is truly transformative for cultural applications.
Tom: It’s amazing how this simple idea—just passing it twice—can unlock so many possibilities in this field.
Jane: We'll be keeping a very close eye on the code and model weights, too, since the authors have made them available.
Lu: This paper marks a significant moment in AI, achieving SOTA performance with incredible efficiency.
Meng: It's a practical solution to problems that have plagued generative NER for years.
Lalam: We’re confident that "Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER" is going to be used in many different ways soon, improving how we interact with data.
WitnessAI
cs.CL
Submitted: 2026-04-06
Updated: 2026-08-26
Project page: https://witness-ai-jpt-ner.hf.space
Importance score: 88/100
The gist: The paper details an investigation into efficient token classification for Zero-Shot Named Entity Recognition (NER), particularly addressing failure modes observed when entity mentions appear early
Key concepts
- Causal Attention Limitation
- LLMs suffer because their causal attention only allows tokens to look at what came before them. This makes it difficult for them to distinguish between entities that have the same name, such as a person named Paris or a city called Paris, because they lack future context.
- Just Pass Twice Mechanism
- The core solution involves duplicating the input sequence. This allows the second pass of each token to see the entire picture, effectively bypassing the causal mask. This gives the model a full view of all future context during processing.
- Definition-Guided Entity Typing
- Instead of relying on fixed labels like 'PERSON,' this method uses natural language descriptions (semantics) of what an entity should be. This makes zero-shot generalization robust by allowing the AI to understand the definition rather than just memorizing a label.
- LoRA Adapters
- The system uses lightweight LoRA adapters, which are small, trainable components added to the LLM. This allows for targeted expertise and significant performance gains without requiring massive retraining or updating large portions of the model parameters.
Terminology
Summary
The paper details an investigation into efficient token classification for Zero-Shot Named Entity Recognition (NER), particularly addressing failure modes observed when entity mentions appear early in a sentence, where the absence of future context can lead to type confusion or missed predictions.
The research analyzes various failure modes, concluding that errors primarily stem from surface-form ambiguity and fine-grained semantic overlap rather than lack of contextual understanding.
Specific error patterns identified include:
-
Type confusions often occur between closely related categories with overlapping definitions.
-
Missed entities frequently appear in descriptive or adjectival forms.
-
Over-predictions arise when domain terms resemble entity mentions.
To mitigate these issues, the authors suggest that Refining type definitions and adding targeted training examples for ambiguous and rare constructions may reduce these errors.
The core benefit demonstrated by the model discussed is its ability to leverage type definitions, as shown by the statement: These examples highlight the benefit of definition-guided modeling for resolving polysemy in context.
The paper provides extensive comparative evidence across multiple datasets (CrossNER-Science, CrossNER-Politics, etc.) and demonstrates that while models like GLiNER and UniNER exhibit various errors—such as GLiNER confusing adjacent categories (e.g., POLITICAL PARTY/ORGANIZATION, POLITICIAN/PERSON)—the model under study successfully resolves these distinctions.
Specific representative examples illustrate the superior performance of the proposed approach:
-
Dutch Universities (Example 1): The ground truth entities were four universities requiring the ORGANIZATION type. Model errors noted that both GLiNER and UniNER
Missed all four universities entirely,
whereas the model correctly predicted all entities. -
Canadian Political Parties (Example 2): For two political parties, the ground truth was POLITICAL PARTY. The model showed that GLiNER
Missed both political parties entirely,
and UniNEROver-predicted types for both parties.
-
US Political Leaders (Example 3): When identifying five US political leaders, the ground truth type was POLITICIAN. Model errors were noted that GLiNER
Labeled all five as PERSON instead of POLITICIAN.
-
Historical Sovereign State (Example 4): For the Empire of Japan, the ground truth type was COUNTRY. The model showed that GLiNER
Labeled as ORGANIZATION instead of COUNTRY,
while the model correctly identified it as COUNTRY.
In summary, the analysis confirms that errors are often due to fine-grained boundary cases and overlapping semantic definitions rather than arbitrary label flips, and the proposed methodology effectively uses type definitions to resolve these complex ambiguities.
Improvements for AI systems
Based on a rigorous analysis of the paper Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER,
I have identified several critical improvements and capabilities that can be integrated into next-generation AI information extraction systems.
The primary deficiency in current generative LLM-based NER is the trade-off between reasoning power (which requires large models) and efficiency/speed (which requires discriminative, single-pass methods). JPT resolves this conflict.
Here are the specific improvements and the resulting capabilities of the enhanced AI system:
The Improvement: Instead of forcing a complex architectural overhaul (like full bidirectional transformers), we utilize Input Duplication. By concatenating the input sequence (x) with itself, x' = (x, [SEP], x), the system leverages the inherent causality of LLMs to achieve bidirectionality. In the second pass, every token in the sequence gains access to all tokens from both its preceding and succeeding contexts (the original first pass).
What this enables in the improved AI system:
-
Accurate Disambiguation: The system can resolve highly ambiguous entities that require future context. For example, it can definitively classify
Paris
as a Location if the full sentence context reveals it is a city, rather than a Person, even though the initial tokens appear in sequence. -
Elimination of Causal Blindness: It overcomes the fundamental limitation where standard LLMs cannot see past their own
future
to determine correct entity boundaries or types.
The Improvement: The system uses Definition-Guided Entity Typing. Instead of relying on a fixed, memorized label (e.g., PERSON
), the the system encodes rich, natural language definitions for each entity type and inject these definitions into the classification process via two channels: (1) as specialized embedding vectors (p j) and (2) directly into the LLM's input prompt.
The Improvement: The system employs a Discriminative Token Classifier built upon the frozen LLM backbone (Qwen3) using Light-weight LoRA adapters and a Bilinear Classifier. This approach replaces slow, sequential generative decoding with a single, highly parallel forward pass.
The resulting improved AI system is a high-speed, zero-shot, bidirectional NER engine capable of:
-
Processing massive volumes of data: Due to its parallel, single-pass nature.
-
Extracting complex entities with high fidelity: By leveraging full context (bidirectionality) and precise definitions (zero-shot typing).
-
Adapting instantly to new domains: By using natural language descriptions instead of fixed labels, ensuring maximum versatility and minimizing retraining costs.
Sources
- PromptNER: Prompting For Named Entity Recognition
- Efficient Training of Language Models to Fill in the Middle
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Splitwise: Efficient generative LLM inference using phase splitting
- DeBERTa: Decoding-enhanced BERT with Disentangled Attention
- LoRA: Low-Rank Adaptation of Large Language Models
- Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study
- Evaluating ChatGPT's Information Extraction Capabilities: An Assessment of Performance, Explainability, Calibration, and Faithfulness
- UL2: Unifying Language Learning Paradigms
- A Systematic Characterization of LLM Inference on GPUs
- InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction
- ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT
- A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering