J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules

arXiv:2608.17063 · cs.LG, cs.CL · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules".

Tom: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on J-Miner, this paper by Gao et al., titled "J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules," is really about taking the opaque decisions made by fine-tuned LLMs and transforming them into something concrete that we can analyze and reuse. It claims they can recover vocabulary-aligned internal concepts and then mine compact Boolean rules or weighted decision logic over those concepts.

Jane: That’s the main thrust, Tom; it moves us away from just observing what the model says to actually understanding the reasoning process behind that verdict. The authors are showing how this mined knowledge can be transferred to a much smaller student model that still retains about ninety-nine point eight percent of the source classifier's accuracy.

Lu: I think the implication here is significant because it connects the high performance of these fine-tuned models with a level of interpretability we haven't seen before. If we can make this knowledge explicit and executable, it opens up new avenues for how we design and trust complex AI systems.

Meng: For practical impact, the ability to use a much smaller student model that runs on those explicit rules means we don't need to deploy the enormous original classifier for every inference, which cuts down on computational load considerably. That makes it much more feasible for real-time applications.

Lalam: From a culture standpoint, I see this as helping us build a better AI ecosystem where we can audit the logic of powerful models without needing access to all their internal workings. It helps make the AI more transparent in how it makes choices.

Tom: That’s right, Lalam; it’s about making the underlying decision-making process visible and auditable for everyone involved. The paper shows that this method doesn't just give you a black box; it gives you a set of rules that you can read and verify.

Jane: So, the takeaway is that J-Miner provides a pathway to extract actionable, reusable decision logic from fine-tuned LLMs, allowing us to build efficient and understandable student models from their internal knowledge. It’s about making the implicit knowledge explicit through a sequence of recovery, revealing, and transfer steps.

Conclusion: Tom: So, we’ve seen how J-Miner takes those complicated fine-tuned models and pulls out actual decision rules, and now we need to talk about what this whole thing means for us as a society.

Jane: It really is fascinating how they managed to turn the implicit knowledge inside those large language models into something concrete that we can actually inspect. I mean, having an executable representation of logic from a massive neural network is a big deal for trust.

Lu: From a theoretical viewpoint, this moves us away from purely statistical black boxes toward verifiable reasoning structures, which opens up entirely new avenues for how we think about complex AI systems. We can finally start mapping the internal decision pathways in ways that are currently impossible.

Meng: But I gotta ask, if we’re extracting these rules and using them to build smaller student models, what does that actually look like in terms of deployment complexity for a startup? Can this really run efficiently outside of huge cloud infrastructure?

Lalam: I think the biggest cultural impact is in transparency; if we can audit the logic behind an AI's classification, it builds real accountability. This shifts the conversation from "does the AI work?" to "is *why* it works sound?"

Tom: Exactly, Lalam. It’s about building a foundation for more trustworthy AI where we aren't just taking outputs at face value. We’ve seen how they recovered rules that match teacher decisions with high fidelity across different tasks.

Jane: And the authors are showing this doesn't just work for one specific task; they found patterns in how these concepts combine across different types of language, like in sentiment classification. That generalizability is where it gets really compelling.

Lu: The way they identified named concepts that act as anchors across different surface forms suggests a universal structure to how these models process language, which is incredibly exciting for future multimodal AI development. We could use this concept vocabulary as a shared language for different domains.

Meng: If the transfer mechanism allows us to get near-source accuracy from a much smaller model, that drastically lowers the barrier to entry for deploying high-performance AI in resource-constrained environments. That practical efficiency is huge.

Lalam: For me, it means we can start building trust into the very fabric of how these models are trained and deployed, making AI development feel less like magic and more like applied science.

Tom: So, to wrap up the core idea, J-Miner gives us a way to extract actionable decision logic from fine-tuned LLMs, turning opaque layers into compact rules that can be reused.

Jane: It’s truly a method for making the internal workings of these powerful models explicit through recovery, revealing, and transfer steps.

Lu: And with the potential to create smaller, rule-based student models that maintain high accuracy while being vastly more efficient than the original classifiers, this opens up whole new layers of research on model distillation and knowledge representation.

Meng: It’s a really neat engineering feat because they managed to preserve that predictive information when shrinking the parameter count so drastically, which is hard to do without losing performance.

Lalam: This advance fundamentally shifts how we approach AI development by focusing on verifiable logic rather than just massive scale, which I think will be a huge positive for the long-term culture of this technology.

Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University · Shanghai Key Laboratory of Data Science, College of Computer Science and Artificial Intelligence, Fudan University

cs.LG, cs.CL

Submitted: 2026-08-17

Updated: 2026-09-28

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 88/100

The gist: Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired

Key concepts

J-Lens (Jacobian lens)
This technique estimates how changes in the model's internal states propagate from an initial layer to the final output. It helps researchers find vocabulary-aligned signals within the large neural network that correspond to specific concepts or features relevant to a particular task.
Message-Level Concept Construction
This stage selects a small, important set of internal variables based on how well they distinguish between different class predictions. These selected variables form the core concept vector used to define the message-level activation state, which represents the model's decision at a specific point.
Executable Rule Induction
This process searches for simple Boolean rules and fits a weighted scorecard over the selected concepts. This creates an explicit, inspectable rule (like 'if concept A is present and concept B is absent, then predict class X') that mimics the original model's complex decision-making.
Knowledge Transfer
The recovered decision knowledge can be used to train a much smaller student model. This student learns to directly apply the fixed rule using raw text input, achieving near-original accuracy while requiring significantly fewer parameters than the source classifier.

Terminology

Summary

Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. This work introduces J-Miner, a method that mines internal decision knowledge from a fine-tuned classifier and encodes it in an executable representation that can be inspected, validated, and reused beyond the source classifier.

How it works

J-Miner recovers task-specific decision knowledge by aggregating vocabulary-aligned internal signals across layers and token positions to identify named concepts, which are then used to learn compact Boolean rules or weighted decision logic. This process distills local internal readouts into an explicit classifier-level knowledge representation. The method is organized into three stages: Recover, Reveal, and Transfer.

  1. Recover uses J-Lens (Jacobian lens) to obtain vocabulary-aligned readouts by estimating the average Jacobian transport from a sourcelayer state to the final-layer states. These readouts are then aggregated into message-level binary concepts by gathering the highest-ranked vocabulary coordinates across layers and positions, resulting in an indicator map where each concept is recorded as either present or absent in a given message.

  2. Message-Level Concept Construction involves selecting a compact set of these variables based on their support and their one-vs-rest trigger rate gap, which measures the difference in mean trigger rates between the two class verdicts on a construction split. The selected vocabulary and concept vector, denoted as C K, are then used to define the message-level activation state c theta(x).

  3. Executable Rule Induction involves searching for short Boolean abstract syntax trees over c theta(x) and fitting a signed linear scorecard over these concepts. This scorecard takes the form g(c) = ¬[w T c + b > t], where the learned weights (w), intercept (b), and threshold (t) define the executable rule R = (C K, w, b, t).

Key Findings in Recovery and Revealing Structure

J-Miner successfully recovers executable decision rules that reproduce source classifier decisions with high fidelity. For binary tasks across six diverse tasks, AST-5 learns one selected five-literal signed Boolean rule, while LR-16 learns a weighted scorecard combining K=16 concepts. The LR-16 rule preserves teacher decisions more completely and supports per-example execution and replay, achieving gains of 6.00–29.50 percentage points over matched Surface rules.

The recovered decision structure is made inspectable through named concepts and explicit rule composition. Task-specific concept vocabularies are identified, where names act as readable anchors for internal directions that recur across different surface forms. The analysis of mean signed contribution reveals how concepts combine: for instance, in Sentiment classification, the rule shows positive weights to amazing and negative weights to bad, dead, and worse, demonstrating how evidence accumulates into a thresholded score.

Knowledge Transfer and Standalone Execution

The recovered decision knowledge can be transferred to a substantially smaller standalone model. A 33.2M E5-small student re-executes the representation from raw text while retaining 99.8% of the source classifiers’ mean task accuracy, reaching between 90.0% and 95.0% teacher fidelity depending on fine-tuning and architecture (e.g., using a character MLP or MiniLM).

The transfer mechanism involves training a compact reader to reconstruct the intermediate decision representation required by the fixed rule, allowing inference without running the source classifier or J-Lens at deployment. The final inference path is: The student predicts explicit concept states directly from raw text and applies the fixed rule to produce its outputs. This design results in a student with 24.1 times fewer parameters than the source classifier while retaining near-source task performance and high decision agreement.

Comparison with Alternative Representations

When compared against alternative feature representations like Sparse Autoencoders (SAEs), J-Miner demonstrates that its selected-name content-word share ranges from.7500 to 1.0000 without external lookup, whereas SAE's range is.000–.2269 and requires four lookup rows. J-Miner preserves predicate identity and rule execution in the same named representation, unlike SAE which exposes latent IDs requiring subsequent naming. The comparison shows that J-Miner retains classifier internal predictive information and presents the selected coordinates directly as vocabulary predicates within a fixed scorecard.

Multiclass Recovery and Coverage

In multiclass tasks, the fixed budget must represent several class directions. While the label-state fidelity is high (e.g.,.993 for SNIPS-3), the fixed sixteen-coordinate vocabulary does not always cover every class direction, as seen in SNIPS-7 where only five of seven classes have selected coordinates.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the J-Miner framework, and what those improved systems will be capable of:


  1. The improved system will possess an Executable Decision Knowledge module. This module transforms a complex, black-box Large Language Model (LLM) classifier into a lightweight, standalone decision engine composed of named concepts and an explicit Boolean rule (a weighted scorecard).

  2. The system will be able to perform high-fidelity inference on raw text without requiring the original, large source classifier or computationally expensive intermediate readouts (like J-Lens). This is achieved by training a compact student model (e.g., 1/24 the parameters of the source) to reconstruct the necessary concept states and apply the fixed decision rule directly to new input.

  3. The system will exhibit significantly higher behavioral fidelity than models relying solely on surface-level lexical rules (e.g., simple keyword matching). The paper demonstrates gains ranging from 6% to over 30 percentage points in teacher fidelity across diverse tasks (SMS, Sentiment, Toxicity).

  4. The system will provide a transparent and inspectable decision structure. Instead of just outputting a class label, the system will output the specific named concepts that were activated and their signed contributions (weights) toward the final verdict. This allows researchers to directly see why a specific decision was made in terms of internal semantic evidence.

  5. The system will be capable of identifying task-specific decision patterns that are hidden from final prediction alone, effectively mining the latent knowledge acquired during fine-tuning.

  6. The system will demonstrate robustness across different LLM families (Qwen, Gemma, Llama) and model scales, proving that the recovered decision knowledge is transferable and not model-specific.

  7. The improved system will be able to dynamically adjust its decision logic based on the input context by utilizing the explicit composition rules (Boolean AST) learned from the source classifier's internal signals.

  8. The system can be used for efficient content moderation, spam filtering, and sentiment analysis with high accuracy while maintaining a much smaller memory footprint and faster inference speed due to its compact architecture.

Sources

Related papers