Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks

arXiv:2501.06286 · cs.CL, cs.AI · Submitted 2026-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks".

Jane: The paper was written by Iman Barati, Arash Ghafouri and Behrouz Minaei-Bidgoli from Iran University of Science & Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone! Today we're cracking open a fresh arXiv paper, and it's called "Bactrainus: Optimizing Large Language Models for Multihop Complex Question Answering Tasks." Jane, I gotta say, that name Bactrainus — what's the story there?

Jane: Tom, I looked into it, and it seems like a play on Bactrian camels — you know, the ones with two humps? And that actually fits perfectly, because this paper is all about a two-stage architecture. Two humps, two stages. I love it when a name actually means something.

Tom: That's a great catch! So we've got this two-hump camel model, and it's tackling something called multi-hop question answering. For our listeners who might not be deep in the weeds, what does multi-hop actually mean here?

Jane: So imagine you ask a question, and to answer it, you can't just look at one paragraph. You have to read one document, find a clue, then jump to another document, find another clue, and then combine them. That's the "multi-hop" — multiple jumps across different pieces of text. The paper uses the HotpotQA dataset, which is the standard benchmark for this kind of thing.

Tom: And that's where the two humps come in. The first stage, the selector, figures out which documents actually matter. The second stage, the reader, takes those chosen documents and actually reasons out the answer. It's like having a librarian find the right books, and then a detective piece together the clues from those books.

Jane: Exactly. And the authors — Iman Barati, Arash Ghafouri, and Behrouz Minaei-Bidgoli from Iran University of Science and Technology — they're asking a really fundamental question. Can we take these general-purpose large language models and make them actually good at this specific, complex task? Not just okay at it, but better than the dedicated systems that were built for this exact problem.

Tom: And that's the exciting part, because most of the time we see LLMs being tested on general knowledge, like trivia or summarization. But this is about structured reasoning. It's about the model having to prove it can connect dots across multiple sources. That's a much harder bar to clear.

Jane: Right, and they're not just testing them out of the box. They're fine-tuning them, they're using techniques like chain of thought, and they're even decomposing questions into smaller sub-questions. It's a full toolkit approach to see how much we can squeeze out of these models.

Tom: So the real question for the rest of the show is, does the two-hump camel actually outrun the competition? Stick around, because we're about to dig into the results, and trust me, they're pretty impressive.

Paper Summary: Tom: Alright, we're back with "Bactrainus: Optimizing Large Language Models for Multihop Complex Question Answering Tasks." Jane, we talked about the setup, but let's get into what they actually found. Give me the big picture.

Jane: Well Tom, the headline is that this modular approach — splitting the task into a selector and a reader — works. It beats the alternative, which is just taking one giant model and asking it to do everything at once. They call that the "all-in-one" approach, and it consistently scored lower.

Tom: How much lower? Give me some numbers.

Jane: So when they fine-tuned a single model to both find supporting facts and answer the question, it got an F1 score of about seventy-five point nine six on the joint task. But when they split it into the two-stage selector-reader, the best configuration jumped to seventy-nine point seven zero. That's nearly a four-point improvement, which is significant in this field.

Tom: Four points is huge. But what's really interesting to me is that they didn't just use any model. They tested a bunch of them, from small eight-billion-parameter models all the way up to a massive four hundred five-billion-parameter one. What did they learn from that comparison?

Jane: They learned that size matters, but it's not everything. In a zero-shot setting, GPT-4o was the best, hitting sixty-seven point five four exact match. But among the open-source models, they found that Llama three point one 70B was the sweet spot for their experiments. And here's the kicker — they managed to get that 70B model to outperform GPT-4o after fine-tuning, reaching seventy-five point seven three exact match on the reader component alone.

Tom: Wait, so a fine-tuned open-source model beat the closed-source giant? That's a big deal for accessibility. It means you don't need to pay for the most expensive API to get top-tier results if you're willing to put in the work to fine-tune.

Jane: Precisely. And they also showed that you can do knowledge distillation. They used the 70B model to generate chain-of-thought reasoning traces, and then used those traces to fine-tune the smaller 8B model. That smaller model got a boost, though not quite to the level of the 70B. It's a way to make smaller models punch above their weight.

Tom: So the summary is: break the problem into pieces, fine-tune on the specific task, and use bigger models to teach smaller ones. That's a recipe that could apply way beyond just question answering.

Jane: Absolutely. And the fact that they beat the previous state-of-the-art on HotpotQA — models like Beam Retrieval and PipNet — shows that this isn't just a fun academic exercise. It's a genuinely better way to build these systems.

Tom: I love that. But I'm curious about the selector part, because that's the part that usually gets overlooked. We'll get into that next.

Improvements Suggested: Tom: Welcome back. We're still on "Bactrainus: Optimizing Large Language Models for Multihop Complex Question Answering Tasks." Jane, we talked about the overall results, but let's zoom in on the improvements they suggest. What makes this paper more than just "we fine-tuned a model and it worked"?

Jane: Great question, Tom. The paper really digs into two specific improvement techniques. The first is chain of thought, which is essentially teaching the model to show its work. Instead of just giving the answer, the model learns to generate a step-by-step reasoning process. They found that when they used a 70B model to generate these reasoning traces and then fine-tuned an 8B model on them, the smaller model got better.

Tom: So it's like a teacher showing a student how to solve a math problem, and the student learns the method, not just the answer. But they also did something with question decomposition, right? That's the second technique.

Jane: Exactly. Question decomposition is about breaking a complex, multi-hop question into simpler sub-questions. For example, if the question is "Who won the Nobel Prize in Physics the year after Marie Curie's husband died?" you might first ask "When did Marie Curie's husband die?" and then "Who won the Nobel Prize in Physics that year?" They trained a separate model to do this decomposition, and then fed those sub-questions into the selector.

Tom: And did that actually help? Because in my experience, adding more moving parts can sometimes just add more places for things to go wrong.

Jane: It helped, but modestly. When they used sub-questions in the two-stage selector, the supporting fact F1 score went from eighty-nine point two one to eighty-nine point six three. It's not a massive jump, but it's consistent. And when you combine that with the chain of thought in the reader, the joint score went up to seventy-nine point seven zero, which is their best result.

Tom: So the improvements are real, but they're incremental. What I find more interesting is what they learned about the selector itself. They found that a single-stage selector — one that directly identifies both gold paragraphs and supporting sentences — was actually just as good as a more complicated two-stage paragraph-then-sentence selector. That's counterintuitive.

Jane: It is! You'd think splitting it into two focused tasks would help, but it didn't. The single-stage selector got an F1 of eighty-nine point two seven on supporting facts, and the two-stage got eighty-nine point two one. Basically a tie. The authors think it's because paragraph selection and sentence selection are so tightly coupled that separating them loses some useful context.

Tom: That's a really important finding for anyone building these systems. Sometimes the simpler architecture is just as good, and it's way easier to maintain. But the other thing I want to highlight is their analysis of how the reader depends on its input. They showed that if you give the model only the question with no supporting facts, it tanks — like twenty-one point six six exact match for the 8B model. But if you give it the gold paragraphs, it jumps to fifty-eight point two nine.

Jane: And that proves the model isn't just relying on memorized knowledge. It's actually reading and reasoning over the provided text. That's a good sanity check for the whole field. The model needs the right information, and it needs the selector to find that information accurately.

Tom: So the improvements are about teaching the reader to reason better and teaching the selector to find better evidence, but also knowing when to keep things simple. That's a great lesson. Now, let's bring in our expert panel to get their take on what this means for the real world.

Conclusion: Tom: We're wrapping up our discussion on "Bactrainus: Optimizing Large Language Models for Multihop Complex Question Answering Tasks." Jane, before we say goodbye to this paper, give me the final takeaway.

Jane: Tom, the takeaway is that modularity wins. By splitting the task into a selector and a reader, fine-tuning each one separately, and using techniques like chain of thought and question decomposition, they beat the previous state-of-the-art on HotpotQA. Their best model hit fifty-one point seven three exact match and seventy-nine point seven zero F1 on the joint task of finding supporting facts and answering correctly.

Tom: And they did it with open-source models. That's the part that gets me excited. You don't need a secret proprietary API to get these results. You can take Llama three point one 70B, fine-tune it with the right approach, and outperform systems that were specifically designed for this task.

Jane: Exactly. And the knowledge distillation part — using the 70B to teach the 8B — means that even smaller, cheaper models can be competitive. That has huge implications for deployment. You could run a decent multi-hop QA system on a single consumer GPU, which opens the door for smaller companies and researchers.

Tom: Now let's get our panel's final thoughts. Lu, what's the big-picture impact here?

Lu: This paper shows that the path to better reasoning isn't just about scaling up. It's about architecture and training strategy. The two-stage design is a template that could be applied to other complex tasks — legal document analysis, medical diagnosis from multiple reports, even scientific literature review. Anywhere you need to gather evidence from multiple sources and then reason over it.

Meng: From an engineering standpoint, I love that they published all their hyperparameters. The LoRA ranks, the learning rates, the batch sizes — it's all there. That means we can reproduce their results without guessing. And the fact that they used LoRA at all means the fine-tuning is actually feasible on modest hardware. That's practical.

Lalam: And from a cultural perspective, this paper democratizes advanced reasoning. When open-source models can match closed-source giants on a complex benchmark, it means the capability is no longer locked behind a paywall. That could lead to better educational tools, more accessible research assistants, and even citizen science projects where people build their own QA systems for niche domains.

Tom: Beautifully said. So we're saying goodbye to Bactrainus, but the ideas here — modular design, knowledge distillation, question decomposition — those are going to stick around. Jane, what's next on our reading list?

Jane: We've got a paper on retrieval-augmented generation for medical question answering coming up. Should be a good follow-up to this one.

Tom: Can't wait. Thanks for listening, everyone. We'll catch you on the next episode.

Iman Barati, Arash Ghafouri, Behrouz Minaei-Bidgoli

Iran University of Science & Technology

cs.CL, cs.AI

Submitted: 2026-08-18

Updated: 2026-08-19

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 45/100

Key concepts

Multi-hop Question Answering
A task where answering a question requires jumping across multiple pieces of text. Instead of finding an answer in one paragraph, a model must find a clue in one document, move to another document for a second clue, and then combine them to reach a conclusion.
Two-Stage Architecture
Named after the two humps of a Bactrian camel, this modular approach splits a task into two parts. A 'selector' first identifies the most relevant documents, and then a 'reader' uses those chosen documents to perform the actual reasoning required to answer the question.
Knowledge Distillation
A method where a larger, more powerful model is used to train a smaller, more efficient model. Researchers used a 70B parameter model to generate step-by-step reasoning traces, which were then used to fine-tune an 8B parameter model to improve its performance.

Terminology

Summary

Summary

This paper introduces Bactrainus, a two-stage selector-reader architecture designed to optimize large language models (LLMs) for multi-hop question answering (MHQA) tasks, evaluated on the HotpotQA dataset under the distractor setting. The authors state: “To tackle this problem, we have designed a two-stage selector-reader architecture, where each stage utilizes an independent LLM.” The central hypothesis is that “dividing the complex QA task into two independent sub-tasks can lead to improved performance,” allowing each component to focus on its respective objective.

The selector component is responsible for “identifying and extracting the supporting facts from among multiple candidate paragraphs (including the ‘gold paragraphs’ and distractors).” The reader component “utilizes the selector’s output (i.e., the chosen paragraphs or sentences) to answer the main question.” The authors also investigate two knowledge distillation techniques: chain of thought (CoT) in the reader and question decomposition in the selector. They note: “The results of the study show that the integration of large language models with these techniques can lead to up to a 4% improvement in F1 score for finding answers.”

Experimental Setup: The HotpotQA dataset is used in its distractor configuration, where “each question is accompanied by 10 candidate paragraphs, 2 of which are gold paragraphs.” Training samples are converted into the Alpaca format for supervised fine-tuning. All models are evaluated under a uniform configuration with “Temperature = 0.01” and “Top-p = 0.8” to produce focused, accurate outputs. Experiments were conducted on a server with 4 A100 GPUs (80 GB each).

Reader Component Evaluation: In zero-shot prompting, GPT-4o achieved the highest scores (EM = 67.54, F1 = 83.44), but was excluded from further experiments due to its closed-source nature and cost. Among open-source models, Llama3.1 8B Instruct performed best under 10B parameters (EM = 60.11, F1 = 74.52), and Llama3.1 70B Instruct performed best above 10B parameters (EM = 65.60, F1 = 80.04). These two models were selected for further study.

Impact of Supporting Facts Count: The authors analyzed performance across questions with two, three, or four or more supporting facts. Contrary to intuition, “a higher number of supporting facts does not necessarily yield lower performance.” However, the performance gap between smaller and larger models widens with more facts: “when there are two supporting facts, the difference in EM between Llama 3.1 8B Instruct and Llama 3.1 405B Instruct is only about 7%, but for four or more supporting facts, this gap increases to over 9%.”

Input Dependence: In a “Question Only” setting, performance dropped dramatically (e.g., Llama 3.1 8B: EM = 21.66, F1 = 29.76), indicating that “neither model possesses adequate internal knowledge to answer most HotpotQA questions.” Adding gold paragraphs improved scores, but adding distractor paragraphs caused noticeable declines. The authors conclude that “delivering concise, high-fidelity input (supporting facts or gold paragraphs) to the reader is crucial for success.”

Few-Shot and Chain of Thought: One-shot prompting yielded the best performance for both models (e.g., Llama 3.1 8B: EM = 63.24, F1 = 77.50; Llama 3.1 70B: EM = 68.18, F1 = 82.66). Adding more examples sometimes caused confusion. For the 8B model, CoT prompts slightly degraded performance, while for the 70B model, CoT provided modest benefits at higher shot counts. The best overall result remained one-shot without CoT.

Fine-tuning the Reader: Using LoRA, the authors fine-tuned Llama 3.1 Instruct models on HotpotQA. The best reader was Bactrainus Reader 70B, achieving EM = 75.73 and F1 = 90.01. Knowledge transfer via CoT generated by a 70B model slightly improved the 8B reader (EM = 74.19, F1 = 86.91), while CoT from an 8B model caused a minor drop (EM = 72.97, F1 = 85.62).

Selector Component: The authors fine-tuned Llama 3.1 Instruct 8B for selection tasks. The Single-stage Selector achieved strong paragraph-level performance (EM = 96.83, F1 = 98.37) but lower sentence-level accuracy (EM = 65.74, F1 = 89.27). The Paragraph Selector excelled at identifying gold paragraphs (EM = 96.64, F1 = 98.24). The Sentence Selector (with known paragraphs) achieved EM = 66.01, F1 = 89.93. A Two-stage Selector (paragraph + sentence) did not significantly outperform the single-stage approach. Adding sub-questions (generated by a larger LLM) provided a slight improvement in supporting-fact identification (EM = 65.93, F1 = 89.63).

Integrated System: Six integration scenarios were tested. The best configuration was Scenario 5 (two-stage selector with sub-questions, feeding supporting facts to the reader) using Bactrainus 70B as the reader, achieving:

  • Supporting Facts: EM = 65.93, F1 = 89.63

  • Answer: EM = 75.07, F1 = 89.01

  • Joint: EM = 51.73, F1 = 79.70

The all-in-one single-model approach (Scenario 1) performed worse (Answer EM = 71.24, F1 = 83.31; Joint EM = 47.93, F1 = 75.96), confirming that “dividing a complex QA task into simpler sub-tasks yields performance gains.”

Comparison to State-of-the-Art: Bactrainus 70B outperformed previous methods, including Beam Retrieval, PipNet, Smoothing R3, and FE2H on ALBERT, across all metrics. For example, Bactrainus 70B achieved Answer EM = 75.07 and Joint EM = 51.73, compared to Beam Retrieval’s Answer EM = 72.69 and Joint EM = 50.53.

Key Conclusions:

  1. LLMs can effectively serve as selectors, achieving near state-of-the-art paragraph identification.

  2. Larger models and CoT knowledge transfer improve reader performance.

  3. A modular selector-reader design outperforms a monolithic all-in-one model by about 2% in both supporting-fact and answer retrieval.

  4. The best configuration surpasses existing HotpotQA baselines, validating the utility of task decomposition and knowledge distillation.

Future Work includes scaling beyond 70B parameters, extending to other languages and domains, interactive and adaptive learning, enhanced interpretability, and multi-agent approaches for complex MHQA.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:


1. Modular Two-Stage Architecture (Selector-Reader)

  • Implementation: Replace monolithic end-to-end QA models with a two-stage pipeline: a fine-tuned LLM-based selector (identifies gold paragraphs and supporting sentences) followed by a fine-tuned LLM-based reader (generates the final answer from selected evidence).

  • Capability: The system will achieve 2% higher Exact Match (EM) and F1 on both supporting-fact identification and final answers compared to a single all-in-one model, as demonstrated in the paper (e.g., scenario 5 vs. scenario 1 in Table 11).

2. Fine-Tuning with LoRA on Instruction-Formatted Data

  • Implementation: Convert HotpotQA training samples into Alpaca-style instruction format. Fine-tune Llama 3.1 8B and 70B Instruct models using LoRA (rank 64, alpha 128, dropout 0.05) with a low temperature (0.01) and top-p (0.8) for inference.

  • Capability: The reader alone will achieve EM of 74.02 (8B) and 75.73 (70B) with F1 of 86.46 and 90.01 respectively, when given correct supporting facts—a significant jump from zero-shot (60.11 EM for 8B).

3. Chain-of-Thought (CoT) Knowledge Distillation from a Larger Teacher Model

  • Implementation: Generate step-by-step reasoning traces using Llama 3.1 70B (teacher) for training samples, then fine-tune the 8B student model to reproduce both the CoT and the final answer. Also, append CoT to prompts during inference for larger models.

  • Capability: The 8B reader with CoT from 70B achieves EM 74.19 (vs. 74.02 without CoT) and F1 86.91 (vs. 86.46). This provides a cost-effective way to boost smaller models' reasoning without scaling hardware.

4. Question Decomposition for Selector Enhancement

  • Implementation: Train a separate 8B model (question decomposer) to break multi-hop questions into simpler sub-questions. Feed these sub-questions along with the original question into the sentence selector to improve supporting-fact extraction.

  • Capability: Supporting-fact EM improves from 65.55 to 65.93 and F1 from 89.21 to 89.63 in the two-stage selector, leading to a small but consistent gain in final answer accuracy (e.g., scenario 5 vs. 3 in Table 11).

5. Optimal Few-Shot Prompting (One-Shot Without CoT)

  • Implementation: For zero-shot or few-shot inference (when fine-tuning is not feasible), use exactly one high-quality example (covering both bridge and comparison question types) without chain-of-thought reasoning.

  • Capability: This yields the best few-shot performance: EM 63.24 and F1 77.50 for the 8B model (vs. 60.11 zero-shot). Adding more examples or CoT degrades performance, so the system will avoid those configurations.

6. Input Truncation and Distractor Filtering

  • Implementation: Before feeding documents to the reader, filter out irrelevant paragraphs using a sentence-transformer (e.g., gte-large-en-v1.5) to select only the top-2 most similar distractors, or better, rely on the fine-tuned selector to output only supporting facts (not full paragraphs).

  • Capability: Prevents performance collapse: providing all 10 paragraphs drops EM to 45.21 (8B) and 46.50 (70B), whereas providing only supporting facts keeps EM above 74 (fine-tuned). This ensures the system remains accurate even with noisy retrieval.

7. Hyperparameter Standardization for Fair Comparison

  • Implementation: Use a uniform configuration across all models: temperature 0.01, top-p 0.8, cosine learning rate scheduler, warm-up ratio 0.03, and LoRA on QKVO + MLP weights. Only modify the system prompt when testing CoT or decomposition.

  • Capability: Ensures reproducible and comparable results, reducing the risk of overfitting to model-specific quirks and making the system robust across different LLM backbones.

  • Answer complex multi-hop questions (e.g., What is the capital of the country where the inventor of the telephone was born?) with 75.07 EM and 89.01 F1 on HotpotQA distractor setting—outperforming prior state-of-the-art (Beam Retrieval: 72.69 EM, 85.04 F1).

  • Identify supporting evidence with 65.93 EM and 89.63 F1, enabling explainable answers (the system can show which paragraphs and sentences it used).

  • Operate cost-effectively: Use an 8B model fine-tuned with CoT from a 70B teacher to achieve near-70B performance (74.19 EM vs. 75.73 EM) at a fraction of the inference cost.

  • Handle noisy or large input contexts: Automatically filter irrelevant documents and extract only the necessary evidence, preventing accuracy degradation when given 10+ paragraphs.

  • Adapt to new domains: The modular selector-reader design allows fine-tuning each component independently on new datasets (e.g., legal, medical) without retraining the entire system.

  • Provide step-by-step reasoning: The system can output chain-of-thought explanations, making its answers more interpretable and trustworthy for human review.

These improvements are directly derived from the paper's experimental results and are ready for implementation in production QA systems, retrieval-augmented generation (RAG) pipelines, or any application requiring multi-hop reasoning over multiple documents.

Abstract

In recent years, the use of large language models (LLMs) has significantly increased, and these models have demonstrated remarkable performance in a variety of general language tasks. However, the evaluation of their performance in domain-specific tasks, particularly those requiring deep natural language understanding, has received less attention. In this research, we evaluate the ability of large language models in performing domain-specific tasks, focusing on the multi-hop question answering (MHQA) problem using the HotpotQA dataset. This task, due to its requirement for reasoning and combining information from multiple textual sources, serves as a challenging benchmark for assessing the language comprehension capabilities of these models. To tackle this problem, we have designed a two-stage selector-reader architecture, where each stage utilizes an independent LLM. In addition, methods such as Chain of Thought (CoT) and question decomposition have been employed to investigate their impact on improving the model's performance. The results of the study show that the integration of large language models with these techniques can lead to up to a 4% improvement in F1 score for finding answers, providing evidence of the models' ability to handle domain-specific tasks and their understanding of complex language.

Sources

Related papers