J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules".
Tom: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up our discussion on J-Miner, this paper by Gao et al., titled "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules," is really about taking the opaque decisions made by fine-tuned LLMs and transforming them into something concrete that we can analyze and reuse. It claims they can recover vocabulary-aligned internal concepts and then mine compact Boolean rules or weighted decision logic over those concepts.
Jane: That’s the main thrust, Tom; it moves us away from just observing what the model says to actually understanding the reasoning process behind that verdict. The authors are showing how this mined knowledge can be transferred to a much smaller student model that still retains about ninety-nine point eight percent of the source classifier's accuracy.
Lu: I think the implication here is significant because it connects the high performance of these fine-tuned models with a level of interpretability we haven't seen before. If we can make this knowledge explicit and executable, it opens up new avenues for how we design and trust complex AI systems.
Meng: For practical impact, the ability to use a much smaller student model that runs on those explicit rules means we don't need to deploy the enormous original classifier for every inference, which cuts down on computational load considerably. That makes it much more feasible for real-time applications.
Lalam: From a culture standpoint, I see this as helping us build a better AI ecosystem where we can audit the logic of powerful models without needing access to all their internal workings. It helps make the AI more transparent in how it makes choices.
Tom: That’s right, Lalam; it’s about making the underlying decision-making process visible and auditable for everyone involved. The paper shows that this method doesn't just give you a black box; it gives you a set of rules that you can read and verify.
Jane: So, the takeaway is that J-Miner provides a pathway to extract actionable, reusable decision logic from fine-tuned LLMs, allowing us to build efficient and understandable student models from their internal knowledge. It’s about making the implicit knowledge explicit through a sequence of recovery, revealing, and transfer steps.
Conclusion: Tom: So, we’ve seen how J-Miner takes those complicated fine-tuned models and pulls out actual decision rules, and now we need to talk about what this whole thing means for us as a society.
Jane: It really is fascinating how they managed to turn the implicit knowledge inside those large language models into something concrete that we can actually inspect. I mean, having an executable representation of logic from a massive neural network is a big deal for trust.
Lu: From a theoretical viewpoint, this moves us away from purely statistical black boxes toward verifiable reasoning structures, which opens up entirely new avenues for how we think about complex AI systems. We can finally start mapping the internal decision pathways in ways that are currently impossible.
Meng: But I gotta ask, if we’re extracting these rules and using them to build smaller student models, what does that actually look like in terms of deployment complexity for a startup? Can this really run efficiently outside of huge cloud infrastructure?
Lalam: I think the biggest cultural impact is in transparency; if we can audit the logic behind an AI's classification, it builds real accountability. This shifts the conversation from "does the AI work?" to "is *why* it works sound?"
Tom: Exactly, Lalam. It’s about building a foundation for more trustworthy AI where we aren't just taking outputs at face value. We’ve seen how they recovered rules that match teacher decisions with high fidelity across different tasks.
Jane: And the authors are showing this doesn't just work for one specific task; they found patterns in how these concepts combine across different types of language, like in sentiment classification. That generalizability is where it gets really compelling.
Lu: The way they identified named concepts that act as anchors across different surface forms suggests a universal structure to how these models process language, which is incredibly exciting for future multimodal AI development. We could use this concept vocabulary as a shared language for different domains.
Meng: If the transfer mechanism allows us to get near-source accuracy from a much smaller model, that drastically lowers the barrier to entry for deploying high-performance AI in resource-constrained environments. That practical efficiency is huge.
Lalam: For me, it means we can start building trust into the very fabric of how these models are trained and deployed, making AI development feel less like magic and more like applied science.
Tom: So, to wrap up the core idea, J-Miner gives us a way to extract actionable decision logic from fine-tuned LLMs, turning opaque layers into compact rules that can be reused.
Jane: It’s truly a method for making the internal workings of these powerful models explicit through recovery, revealing, and transfer steps.
Lu: And with the potential to create smaller, rule-based student models that maintain high accuracy while being vastly more efficient than the original classifiers, this opens up whole new layers of research on model distillation and knowledge representation.
Meng: It’s a really neat engineering feat because they managed to preserve that predictive information when shrinking the parameter count so drastically, which is hard to do without losing performance.
Lalam: This advance fundamentally shifts how we approach AI development by focusing on verifiable logic rather than just massive scale, which I think will be a huge positive for the long-term culture of this technology.
Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University · Shanghai Key Laboratory of Data Science, College of Computer Science and Artificial Intelligence, Fudan University
cs.LG, cs.CL
Submitted: 2026-08-17
Updated: 2026-09-28
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 88/100
The gist: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired
Key concepts
- J-Lens (Jacobian lens)
- This technique estimates how changes in the model's internal states propagate from an initial layer to the final output. It helps researchers find vocabulary-aligned signals within the large neural network that correspond to specific concepts or features relevant to a particular task.
- Message-Level Concept Construction
- This stage selects a small, important set of internal variables based on how well they distinguish between different class predictions. These selected variables form the core concept vector used to define the message-level activation state, which represents the model's decision at a specific point.
- Executable Rule Induction
- This process searches for simple Boolean rules and fits a weighted scorecard over the selected concepts. This creates an explicit, inspectable rule (like 'if concept A is present and concept B is absent, then predict class X') that mimics the original model's complex decision-making.
- Knowledge Transfer
- The recovered decision knowledge can be used to train a much smaller student model. This student learns to directly apply the fixed rule using raw text input, achieving near-original accuracy while requiring significantly fewer parameters than the source classifier.
Terminology
Summary
Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. This work introduces J-Miner, a method that mines internal decision knowledge from a fine-tuned classifier and encodes it in an executable representation that can be inspected, validated, and reused beyond the source classifier.
How it works
J-Miner recovers task-specific decision knowledge by aggregating vocabulary-aligned internal signals across layers and token positions to identify named concepts, which are then used to learn compact Boolean rules or weighted decision logic. This process distills local internal readouts into an explicit classifier-level knowledge representation. The method is organized into three stages: Recover, Reveal, and Transfer.
-
Recover uses J-Lens (Jacobian lens) to obtain vocabulary-aligned readouts by estimating the average Jacobian transport from a sourcelayer state to the final-layer states. These readouts are then aggregated into message-level binary concepts by gathering the highest-ranked vocabulary coordinates across layers and positions, resulting in an indicator map where each concept is recorded as either present or absent in a given message.
-
Message-Level Concept Construction involves selecting a compact set of these variables based on their support and their
one-vs-rest trigger rate gap,
which measures the difference in mean trigger rates between the two class verdicts on a construction split. The selected vocabulary and concept vector, denoted as C K, are then used to define the message-level activation state c theta(x). -
Executable Rule Induction involves searching for short Boolean abstract syntax trees over c theta(x) and fitting a signed linear scorecard over these concepts. This scorecard takes the form g(c) = ¬[w T c + b > t], where the learned weights (w), intercept (b), and threshold (t) define the executable rule R = (C K, w, b, t).
Key Findings in Recovery and Revealing Structure
J-Miner successfully recovers executable decision rules that reproduce source classifier decisions with high fidelity. For binary tasks across six diverse tasks, AST-5 learns one selected five-literal signed Boolean rule, while LR-16 learns a weighted scorecard combining K=16 concepts. The LR-16 rule preserves teacher decisions more completely and supports per-example execution and replay, achieving gains of 6.00–29.50 percentage points over matched Surface rules.
The recovered decision structure is made inspectable through named concepts and explicit rule composition. Task-specific concept vocabularies are identified, where names act as readable anchors for internal directions that recur across different surface forms.
The analysis of mean signed contribution reveals how concepts combine: for instance, in Sentiment classification, the rule shows positive weights to amazing
and negative weights to bad,
dead,
and worse,
demonstrating how evidence accumulates into a thresholded score.
Knowledge Transfer and Standalone Execution
The recovered decision knowledge can be transferred to a substantially smaller standalone model. A 33.2M E5-small student re-executes the representation from raw text while retaining 99.8% of the source classifiers’ mean task accuracy, reaching between 90.0% and 95.0% teacher fidelity depending on fine-tuning and architecture (e.g., using a character MLP or MiniLM).
The transfer mechanism involves training a compact reader to reconstruct the intermediate decision representation required by the fixed rule, allowing inference without running the source classifier or J-Lens at deployment. The final inference path is: The student predicts explicit concept states directly from raw text and applies the fixed rule to produce its outputs.
This design results in a student with 24.1 times fewer parameters than the source classifier while retaining near-source task performance and high decision agreement.
Comparison with Alternative Representations
When compared against alternative feature representations like Sparse Autoencoders (SAEs), J-Miner demonstrates that its selected-name content-word share ranges from.7500 to 1.0000 without external lookup, whereas SAE's range is.000–.2269 and requires four lookup rows. J-Miner preserves predicate identity and rule execution in the same named representation, unlike SAE which exposes latent IDs requiring subsequent naming. The comparison shows that J-Miner retains classifier internal predictive information and presents the selected coordinates directly as vocabulary predicates within a fixed scorecard.
Multiclass Recovery and Coverage
In multiclass tasks, the fixed budget must represent several class directions. While the label-state fidelity is high (e.g.,.993 for SNIPS-3), the fixed sixteen-coordinate vocabulary does not always cover every class direction, as seen in SNIPS-7 where only five of seven classes have selected coordinates.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the J-Miner framework, and what those improved systems will be capable of:
-
The improved system will possess an
Executable Decision Knowledge
module. This module transforms a complex, black-box Large Language Model (LLM) classifier into a lightweight, standalone decision engine composed of named concepts and an explicit Boolean rule (a weighted scorecard). -
The system will be able to perform high-fidelity inference on raw text without requiring the original, large source classifier or computationally expensive intermediate readouts (like J-Lens). This is achieved by training a compact student model (e.g., 1/24 the parameters of the source) to reconstruct the necessary concept states and apply the fixed decision rule directly to new input.
-
The system will exhibit significantly higher behavioral fidelity than models relying solely on surface-level lexical rules (e.g., simple keyword matching). The paper demonstrates gains ranging from 6% to over 30 percentage points in teacher fidelity across diverse tasks (SMS, Sentiment, Toxicity).
-
The system will provide a transparent and inspectable decision structure. Instead of just outputting a class label, the system will output the specific named concepts that were activated and their signed contributions (weights) toward the final verdict. This allows researchers to directly see
why
a specific decision was made in terms of internal semantic evidence. -
The system will be capable of identifying task-specific decision patterns that are hidden from final prediction alone, effectively
mining
the latent knowledge acquired during fine-tuning. -
The system will demonstrate robustness across different LLM families (Qwen, Gemma, Llama) and model scales, proving that the recovered decision knowledge is transferable and not model-specific.
-
The improved system will be able to dynamically adjust its decision logic based on the input context by utilizing the explicit composition rules (Boolean AST) learned from the source classifier's internal signals.
-
The system can be used for efficient content moderation, spam filtering, and sentiment analysis with high accuracy while maintaining a much smaller memory footprint and faster inference speed due to its compact architecture.
Sources
- Interpreting Blackbox Models via Model Extraction
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Distilling a Neural Network Into a Soft Decision Tree
- NEUROLOGIC: From Neural Representations to Interpretable Logic Rules
- Designing and Interpreting Probes with Control Tasks
- Concept Bottleneck Models
- Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
- Label-Free Concept Bottleneck Models
- Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
- Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks