Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio.
Tom: Next we'll be talking about the paper "Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies".
Jane: The paper was written by Isak Hwang, Yoon Pyo Lee and Syed Bahauddin Alam from Hanyang University and University of Illinois Urbana-Champaign.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: So, let's talk about the authors and their approach, Isak Hwang and Yoon Pyo Lee. They took a specific path to test this model, which is very methodical.
Jane: The researchers chose an open-weight multimodal model called Gemma four 31B-IT to handle the complexity of the NRC exam. The multimodal part is key because many questions involve diagrams or piping schematics that need visual understanding.
Tom: And they are testing it against the Generic Fundamentals Examination, or GFE, which is the standard assessment used by humans. They are holding this model to the exact same eighty percent passing threshold as human candidates.
Lu: This is brilliant because it lets us see if we can even compare an algorithmic performance against a human performance on identical terms. It’s a fair test of domain expertise, not just pattern matching.
Meng: The scale of the examination is quite large; they are testing fourteen different exams across two reactor types, which is a solid dataset to ensure robustness.
Lalam: I think this setup tells us that we aren't just looking at isolated successes but at sustained competence over multiple administrations, which is necessary for building trust in a long-term operational tool.
Summary: Tom: So, what did they actually find out when they ran the experiment? The key findings are pretty clear and tell a very specific story about performance.
Jane: They tested eight different configurations, combining various fine-tuning methods like supervised fine-tuning (SFT) and retrieval-augmented fine-tuning (RAFT) with two different chunking strategies.
Tom: And the standout winner was the SFT configuration using fixed-size chunking for eight of the fourteen examinations. That’s a significant number of successes.
Lu: That suggests that simply teaching the model specific rationales, or CoT distillation, is incredibly powerful for this domain adaptation. It's transferring complex reasoning abilities efficiently to a smaller student model.
Meng: But it also tells us that no configuration without some form of fine-tuning passed any exam, which means the base model is severely lacking in its initial knowledge base.
Lalam: That lack of foundational knowledge is why I think the cultural shift will be slow; we can't deploy systems that fundamentally fail to grasp core concepts.
Improvements: Tom: Now, let's talk about the improvements or the key design insights this paper offers. It’s not just about who won, but *why* they won.
Jane: The researchers found something called a "chunking-strategy reversal," which is a really interesting concept to explain. They discovered that structure-aware chunking works better for the base model, but fixed-size chunking works better for the fine-tuned models.
Tom: That's a fascinating dichotomy; it shows that effective retrieval design depends entirely on the model’s own training state.
Lu: I see this as a fundamental insight into how memory and context are organized within an LL's internal architecture, suggesting that the way we segment external knowledge must match the way we have trained the internal reasoning.
Meng: And it also found that RAFT underperformed plain SFT when comparing matching search environments. This is a major practical finding for building RAG systems.
Lalam: So, if I understand this right, RAFT often tries to force the model to use external evidence even when it knows the answer internally, which is exactly where its failure comes from being sub-optimal for cultural adoption.
Conclusion: Tom: We have covered a lot of ground in this paper, "Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies." It gives us some very clear conclusions about what's possible.
Jane: The best configuration reached eighty point two three percent on the PWR items, which is above the eighty percent passing mark, but overall, with pooled accuracy sitting around seventy-nine point six six percent, it falls just under the threshold when compared to human candidates.
Tom: And that's a critical distinction; the confidence interval for both scores spans that threshold, so we can't claim it's reliable yet as a regulatory standard.
Lu: I think this positions the AI not as a replacement but as an incredibly powerful assistive tool, which is where its most realistic and positive impact lies.
Meng: My takeaway is that you can close the gap between an off-the-shelf model and operational competence using techniques that run entirely on a single commodity workstation, which has huge implications for deployment in specialized industries.
Lalam: It's a powerful reminder that the goal isn' at all is to augment human judgment, not to replace it.
Tom: That's definitely the way to look at it! We are going to wrap up our discussion of this groundbreaking work and get ready for some news on other exciting developments in AI.
Isak Hwang, Yoon Pyo Lee, Syed Bahauddin Alam
Hanyang University · University of Illinois Urbana-Champaign
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
Code: https://github.com/ml-explore/mlx
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper presents a systematic benchmark of a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on the U.S.
Key concepts
- Multimodal Language Models
- These advanced AI models are capable of processing information beyond just text. They were necessary for the NRC exam because many questions involve visual elements, such as diagrams or piping schematics, requiring the model to possess visual understanding.
- NRC Reactor Operator Licensing Examination (GFE)
- This is the standard assessment used by human candidates in the field. The models were tested against this benchmark, which requires passing an eighty percent threshold. The examination covered fourteen different exams across two reactor types.
- Supervised Fine-Tuning (SFT)
- SFT is a method used to adapt a base model's knowledge for a specific domain. Researchers found that simply teaching the model specific rationales was incredibly powerful, efficiently transferring complex reasoning abilities to the student model.
Terminology
Summary
This paper presents a systematic benchmark of a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination. The study evaluates eight model-retrieval configurations against 14 Generic Fundamentals Examinations (GFE) from the March sittings of 2015 through 2021, comprising seven pressurized water reactor (PWR) and seven boiling water reactor (BWR) exams, totaling 700 items (698 scored after exclusions). The passing criterion is fixed by regulation: a candidate must score at least 80% on a single administered examination, and rounding up to reach that mark is not permitted.
The methodology involves several components. The evaluation dataset consists of all GFE examinations administered at the March sitting from 2015 through 2021 for both reactor types, with the study period terminating in 2021 because the NRC discontinued the GFE program on 16 March 2022. The training corpus is built from the NRC question banks (2,140 PWR questions and 2,149 BWR questions), deduplicated against evaluation items by normalized question text matching, leaving 3,577 training items. These items are augmented with chain-of-thought (CoT) rationales distilled from Gemini 3 Flash, a proprietary teacher model. The external retrieval corpus consists of seven volumes of the DOE Fundamentals Handbook series covering Thermodynamics, Heat Transfer, and Fluid Flow (volumes 1–3), Instrumentation and Control (volumes 1–2), and Nuclear Physics and Reactor Theory (volumes 1–2).
Two chunking strategies are compared for the retrieval corpus: fixed-size sliding-window chunking with 1,000-character chunks and 200-character overlap (20% overlap), and structure-aware chunking that partitions documents along typographic and table-of-contents boundaries, producing variable-length chunks with a mean of approximately 1,435 characters. Retrieval at inference uses BM25 sparse lexical retrieval with standard parameters (k1 = 1.5, b = 0.75), returning the top four chunks jointly capped at 7,000 characters. Fine-tuning uses LoRA with rank 16, scaling 32, dropout 0.05, targeting language-model linear layers while freezing the vision tower and multimodal projector. Training runs for 3 epochs with effective batch size 8, learning rate 2 × 10−4, cosine scheduling, and bfloat16 precision. Two fine-tuning paradigms are compared: supervised fine-tuning (SFT) and retrieval-augmented fine-tuning (RAFT), with RAFT training examples including retrieved context blocks matched to inference-time retrieval.
The primary results show that the strongest configuration, SFT combined with fixed-size chunking RAG, met the 80% criterion on 8 of 14 examinations (3 of 7 PWR and 5 of 7 BWR), outperforming all alternatives. No configuration without fine-tuning passed any examination. Pooled accuracy for this best configuration reached 79.66% with a Wilson score 95% confidence interval of [76.5, 82.5], which spans the threshold. On PWR items specifically, the accuracy reached 80.23% (280 of 349 items), clearing the criterion by a single item. The base model without any adaptation scored 51.86% pooled accuracy.
Three notable findings emerge. First, a chunking-strategy reversal phenomenon is observed: structure-aware chunking is preferred for the base model (64.47% vs. 62.18% for fixed-size), while fixed-size chunking is preferred for both fine-tuned paradigms (SFT: 79.66% vs. 77.51%; RAFT: 77.36% vs. 75.36%). Second, RAFT underperforms plain SFT under matched chunking and retrieval conditions by a margin of 2.3 percentage points under fixed-size chunking and 2.2 points under structure-aware chunking, a deficit that is consistent across reactor types. The authors attribute this provisionally to a mismatch in the RAFT training pipeline: the teacher's rationales were generated against dense-retrieved passages while the student's training context was sparse-retrieved, so the supervision signal may assert facts absent from the conditioning context. Third, retrieval benefits BWR items more than PWR items in all four available comparisons, though the authors do not claim a reactor-type effect due to sampling error.
The paper also reports substantial heterogeneity across examinations. The March 2019 sitting accounts for nine of the twenty-six passing cells in the results matrix, while the March 2017 sitting is the only one on which no configuration meets the criterion on either paper. The March 2015 BWR paper exhibits a training-coverage gap signature: the base model scores only 3.6 points below its BWR average, but SFT falls 10.5 points and SFT with fixed-size retrieval falls 11.1 points below their averages, indicating content underrepresented in the training corpus rather than intrinsic item difficulty.
The authors emphasize that the evaluation unit is the individual administered examination, matching the regulatory standard, rather than pooled accuracy, which has no direct regulatory interpretation. They note that pooled accuracy can invert qualitative conclusions: pooled figures rank PWR above BWR (80.23% vs. 79.08%), but the per-examination count shows the model met the criterion on 5 of 7 BWR papers against 3 of 7 PWR papers.
The computational framework runs entirely on Apple Silicon hardware with an M5 Max chip and 128 GB unified memory, using PyTorch with Hugging Face transformers and peft for fine-tuning, and MLX format with mlx-vlm for inference. All generation is deterministic at temperature zero with greedy decoding. The deployed system requires no proprietary model, no external inference endpoint, and no network access at run time, though one external dependency exists at construction time: the CoT rationales were distilled from a proprietary teacher model.
The authors conclude that the pipeline reaches the vicinity of the operator-level criterion rather than clearing it with margin, positioning the fine-tuned and retrieval-grounded model as an assistive tool that augments rather than replaces licensed operator judgment. Roughly one item in five is answered incorrectly even in the best configuration. Limitations include the instrument covering generic fundamentals only (not site-specific regulations, technical specifications, or emergency operating procedures), residual similarity between derived exam items and their bank ancestors, the March sitting only being evaluated, and the closed four-choice format not capturing open-ended reasoning or confidence calibration.
Improvements for AI systems
Based on this paper, I can implement the following specific improvements to AI systems, particularly for domain-specific question answering and RAG pipelines:
Improvement: Implement an adaptive chunking selector that detects the model's training state (base vs. fine-tuned) and automatically switches between structure-aware and fixed-size sliding-window chunking.
- What the improved system can do: Automatically select structure-aware chunking for off-the-shelf models (64.47% vs 62.18% accuracy) and fixed-size chunking for fine-tuned models (79.66% vs 77.51%), yielding a consistent 2-point accuracy gain without manual tuning.
Improvement: Build a training pipeline that uses a stronger teacher model (e.g., Gemini 3 Flash) to generate step-by-step rationales, then fine-tunes a smaller student model (e.g., 31B) using LoRA with masked loss over only the rationale tokens.
- What the improved system can do: Raise accuracy from 51.86% to 75.50% on unseen domain questions (a +23.6-point gain) using only 3,577 training examples, all executable on a single workstation with 128GB unified memory.
Improvement: Implement a diagnostic that compares a paper's accuracy deficit under the base model against its deficit after fine-tuning to distinguish intrinsic item difficulty from training-corpus coverage gaps.
- What the improved system can do: Automatically flag under-represented topics in the training corpus. For example, it detected an 11-point accuracy drop on the March 2015 BWR paper that was invisible to pooled accuracy, enabling targeted corpus augmentation before deployment.
Improvement: Replace pooled accuracy as the primary metric with per-examination pass/fail counts against the regulatory threshold (80%), since pooling can invert conclusions (PWR 80.23% pooled vs. 3/7 passed; BWR 79.08% pooled vs. 5/7 passed).
- What the improved system can do: Provide regulators and operators with decision-relevant pass rates per administered exam rather than misleading aggregate statistics, preventing incorrect conclusions about domain competence.
Improvement: Implement a three-tier answer extractor (explicit phrase match → last parenthesized option → terminal standalone option) with unextractable responses scored as incorrect.
- What the improved system can do: Achieve deterministic, reproducible scoring across configurations, ensuring that a 0.4-point accuracy difference (79.66% vs 80.1%) is attributable to model/retrieval changes, not parsing noise.
Improvement: Use BM25 with 20% overlap (1000-char chunks, 200-char overlap) for fine-tuned models, capping top-4 chunks at 7,000 characters.
- What the improved system can do: Increase the probability that the decisive fact appears in retrieved context, yielding 79.66% accuracy (vs. 77.51% for structure-aware) and passing 8/14 exams—the best configuration in the study.
Improvement: Add a pre-training check that measures passage overlap between the teacher's retrieval context and the student's training context to prevent RAFT degradation.
- What the improved system can do: Avoid the 2.3-point accuracy penalty observed when RAFT training uses mismatched contexts (teacher dense-retrieved vs. student sparse-retrieved), ensuring retrieval-conditioned training actually improves rather than degrades performance.
Improvement: Fine-tune only language-side LoRA modules (q/k/v/o/gate/up/down projections) while freezing the vision tower and multimodal projector.
- What the improved system can do: Preserve pretrained image understanding (critical for P&IDs, schematics, graphs) while achieving language-side specialization, enabling correct answers on image-bearing items without vision-catastrophic forgetting.
Improvement: Package the full pipeline (fine-tuning in PyTorch/MPS, adapter merging, MLX conversion, inference via mlx-vlm) to run entirely on Apple Silicon with no network access at inference time.
- What the improved system can do: Deploy an auditable, self-contained assistant in air-gapped environments (e.g., nuclear plants) with no proprietary model dependencies at runtime, supporting data-governance and on-premise requirements.
Improvement: Track retrieval benefit separately by subdomain (e.g., PWR vs. BWR) to detect unexplained asymmetries (base model: +7.2 PWR vs. +13.5 BWR with fixed chunking).
- What the improved system can do: Flag potential corpus biases or topic-specific retrieval failures early, prompting topic-level decomposition rather than relying on aggregate metrics that mask such patterns.
Sources
- Capabilities of GPT-4 on Medical Challenge Problems
- RAFT: Adapting Language Model to Domain Specific RAG
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering