Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
cs.IR, cs.AI
Submitted: 2025-05-17
Updated: 2026-09-06
Comments: The submitted version contains a significant error affecting the reported experimental results. The experiments must be repeated and the results re-evaluated. We therefore request withdrawal of the current submission, as we do not want researchers to rely on or cite potentially inaccurate results
Code: https://github.com/meta-llama/llama-models
License: http://creativecommons.org/licenses/by/4.0/
The gist: Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context.
Terminology
Abstract
Textbook question answering (TQA) is a complex task, requiring the interpretation of complex multimodal context. Although recent advances have improved overall performance, they often encounter difficulties in educational settings where accurate semantic alignment and task-specific document retrieval are essential. In this paper, we propose a novel approach to multimodal textbook question answering by introducing a mechanism for enhancing semantic representations through multi-objective joint training. Our model, Joint Embedding Training With Ranking Supervision for Textbook Question Answering (JETRTQA), is a multimodal learning framework built on a retriever--generator architecture that uses a retrieval-augmented generation setup, in which a multimodal large language model generates answers. JETRTQA is designed to improve the relevance of retrieved documents in complex educational contexts. Unlike traditional direct scoring approaches, JETRTQA learns to refine the semantic representations of questions and documents through a supervised signal that combines pairwise ranking and implicit supervision derived from answers. We evaluate our method on the CK12-QA dataset and demonstrate that it significantly improves the discrimination between informative and irrelevant documents, even when they are long, complex, and multimodal. JETRTQA outperforms the previous state of the art, achieving a 2.4% gain in accuracy on the validation set and 11.1% on the test set.
Sources
- Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
- Improving Retrieval-Augmented Large Language Models via Data Importance Learning
- Latent Retrieval for Weakly Supervised Open Domain Question Answering
- End-to-End Training of Neural Retrievers for Open-Domain Question Answering
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG