J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
summary
The gist
Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired
In short
J-Miner extracts hidden decision logic from fine-tuned large language models by mining internal signals to create compact, executable Boolean rules. It recovers these rules through three stages—Recover, Reveal, and Transfer—allowing knowledge to be inspected and transferred to a much smaller model without needing the original complex classifier.
Key concepts
- J-Lens (Jacobian lens)
- This technique estimates how changes in the model's internal states propagate from an initial layer to the final output. It helps researchers find vocabulary-aligned signals within the large neural network that correspond to specific concepts or features relevant to a particular task.
- Message-Level Concept Construction
- This stage selects a small, important set of internal variables based on how well they distinguish between different class predictions. These selected variables form the core concept vector used to define the message-level activation state, which represents the model's decision at a specific point.
- Executable Rule Induction
- This process searches for simple Boolean rules and fits a weighted scorecard over the selected concepts. This creates an explicit, inspectable rule (like 'if concept A is present and concept B is absent, then predict class X') that mimics the original model's complex decision-making.
- Knowledge Transfer
- The recovered decision knowledge can be used to train a much smaller student model. This student learns to directly apply the fixed rule using raw text input, achieving near-original accuracy while requiring significantly fewer parameters than the source classifier.
Terminology used across episodes
This episode discusses
- J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules · Paper Radio
- Interpreting Blackbox Models via Model Extraction
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Distilling a Neural Network Into a Soft Decision Tree
- NEUROLOGIC: From Neural Representations to Interpretable Logic Rules
- Designing and Interpreting Probes with Control Tasks
- Concept Bottleneck Models
- Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
- Label-Free Concept Bottleneck Models
- Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
- Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
The paper
J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules · Read on arXiv
Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University · Shanghai Key Laboratory of Data Science, College of Computer Science and Artificial Intelligence, Fudan University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules".
Tom: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, wrapping up our discussion on J-Miner, this paper by Gao et al., titled "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules," is really about taking the opaque decisions made by fine-tuned LLMs and transforming them into something concrete that we can analyze and reuse. It claims they can recover vocabulary-aligned internal concepts and then mine compact Boolean rules or weighted decision logic over those concepts.
Jane: That’s the main thrust, Tom; it moves us away from just observing what the model says to actually understanding the reasoning process behind that verdict. The authors are showing how this mined knowledge can be transferred to a much smaller student model that still retains about ninety-nine point eight percent of the source classifier's accuracy.
Lu: I think the implication here is significant because it connects the high performance of these fine-tuned models with a level of interpretability we haven't seen before. If we can make this knowledge explicit and executable, it opens up new avenues for how we design and trust complex AI systems.
Meng: For practical impact, the ability to use a much smaller student model that runs on those explicit rules means we don't need to deploy the enormous original classifier for every inference, which cuts down on computational load considerably. That makes it much more feasible for real-time applications.
Lalam: From a culture standpoint, I see this as helping us build a better AI ecosystem where we can audit the logic of powerful models without needing access to all their internal workings. It helps make the AI more transparent in how it makes choices.
Tom: That’s right, Lalam; it’s about making the underlying decision-making process visible and auditable for everyone involved. The paper shows that this method doesn't just give you a black box; it gives you a set of rules that you can read and verify.
Jane: So, the takeaway is that J-Miner provides a pathway to extract actionable, reusable decision logic from fine-tuned LLMs, allowing us to build efficient and understandable student models from their internal knowledge. It’s about making the implicit knowledge explicit through a sequence of recovery, revealing, and transfer steps.
Conclusion: Tom: So, we’ve seen how J-Miner takes those complicated fine-tuned models and pulls out actual decision rules, and now we need to talk about what this whole thing means for us as a society.
Jane: It really is fascinating how they managed to turn the implicit knowledge inside those large language models into something concrete that we can actually inspect. I mean, having an executable representation of logic from a massive neural network is a big deal for trust.
Lu: From a theoretical viewpoint, this moves us away from purely statistical black boxes toward verifiable reasoning structures, which opens up entirely new avenues for how we think about complex AI systems. We can finally start mapping the internal decision pathways in ways that are currently impossible.
Meng: But I gotta ask, if we’re extracting these rules and using them to build smaller student models, what does that actually look like in terms of deployment complexity for a startup? Can this really run efficiently outside of huge cloud infrastructure?
Lalam: I think the biggest cultural impact is in transparency; if we can audit the logic behind an AI's classification, it builds real accountability. This shifts the conversation from "does the AI work?" to "is *why* it works sound?"
Tom: Exactly, Lalam. It’s about building a foundation for more trustworthy AI where we aren't just taking outputs at face value. We’ve seen how they recovered rules that match teacher decisions with high fidelity across different tasks.
Jane: And the authors are showing this doesn't just work for one specific task; they found patterns in how these concepts combine across different types of language, like in sentiment classification. That generalizability is where it gets really compelling.
Lu: The way they identified named concepts that act as anchors across different surface forms suggests a universal structure to how these models process language, which is incredibly exciting for future multimodal AI development. We could use this concept vocabulary as a shared language for different domains.
Meng: If the transfer mechanism allows us to get near-source accuracy from a much smaller model, that drastically lowers the barrier to entry for deploying high-performance AI in resource-constrained environments. That practical efficiency is huge.
Lalam: For me, it means we can start building trust into the very fabric of how these models are trained and deployed, making AI development feel less like magic and more like applied science.
Tom: So, to wrap up the core idea, J-Miner gives us a way to extract actionable decision logic from fine-tuned LLMs, turning opaque layers into compact rules that can be reused.
Jane: It’s truly a method for making the internal workings of these powerful models explicit through recovery, revealing, and transfer steps.
Lu: And with the potential to create smaller, rule-based student models that maintain high accuracy while being vastly more efficient than the original classifiers, this opens up whole new layers of research on model distillation and knowledge representation.
Meng: It’s a really neat engineering feat because they managed to preserve that predictive information when shrinking the parameter count so drastically, which is hard to do without losing performance.
Lalam: This advance fundamentally shifts how we approach AI development by focusing on verifiable logic rather than just massive scale, which I think will be a huge positive for the long-term culture of this technology.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck