Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering

arXiv:2608.13160 · cs.CL, cs.AI · Submitted 2026-08-13 · Read on arXiv

Yilin Wang, Yuchun Fan, Weidong Bao, Zili Wei, Shi Feng, Tong Xiao, Zhengtao Yu, Jingbo Zhu

School of Computer Science and Engineering, Northeastern University · Yunnan Key Laboratory of Artificial Intelligence, Kunming University of Science and Technology

cs.CL, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted by NLPCC 2026

Code: https://github.com/f6ster/Syfer

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper introduces Syfer, a synthesizer-folding framework for multilingual multi-hop question answering (QA). The paper identifies two structural issues in prior decomposition-based multilingual retrieval-augmented generation (mRAG) systems: (i) one-size-fits-all translation which forcibly aligning native-language documents to the query language incurs a heavy translation cost and tends to inject translation noise that is later answer as evidence, and (ii) greedy decomposition and aggregation where uncontrolled decomposition produces a sub-question graph that is often redundant and logically incoherent, where minor errors in intermediate reasoning propagate and compound, and the final aggregation call is logically independent of the decomposition process, it concentrates rather than absorbs decomposition noise, breaking the reasoning chain and yielding incorrect answers.

Syfer addresses these issues with two key changes. First, "translation becomes decomposition-driven: the framework reasons in the original language by default and activates the cross-lingual pathway only when the produced sub-question graph fails a quality check, after which an English-parallel graph is decomposed and fused for recovery. Second, decomposition becomes synthesizer-folded: a trained decomposer constrains the breadth and format of sub-questions and emits a terminal sub-question in the same logical layer as its peers. This terminal sub-question serves as a synthesis question over the prior reasoning chain, so the generator answers it directly instead of performing a separate aggregation over a long intermediate trace."

The framework consists of four stages: offline logical-decomposition distillation, synthesizer-folded decomposition, faithfulness verification with bilingual fallback, and cross-lingual retrieval-and-answering. In the distillation stage, a teacher model (Qwen3-235BA22B-Instruct-2507) generates decompositions, filtered by a constraint that the filled terminal sub-question remains close to the original query in the retriever embedding space (cosine similarity ≥ τconstraint = 0.8). A student decomposer (Qwen3-8B) is then fine-tuned on 59,688 decomposition records covering six in-distribution languages (English, Chinese, German, Spanish, Swahili, Thai), with French, Bengali, and Korean held out as out-of-distribution languages.

In the synthesizer-folded decomposition stage, the decomposer produces an acyclic sub-question DAG with a unique terminal node. The terminal node is a learned slot whose filled form encodes both Q and the prior sub-answers, so that answering it against the corpus is, by construction, equivalent to aggregating. The faithfulness verification stage computes score(DL) = cos(e(qn filled), e(Q)); if the score is below τconstraint, bilingual fallback is triggered, which translates Q into English, decomposes it, and aligns the English DAG with the original-language DAG by node similarity (threshold τalign = 0.6). In the final stage, sub-questions are solved in topological order, with bilingual nodes retrieving from both language views and merging candidates. Maximal marginal relevance (MMR) with λ = 0.6 is applied to avoid near-duplicate parallel translations.

The paper extends three multilingual multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA, MuSiQue) into a testbed spanning five language families and nine languages across high-, mid-, and low-resource regimes. Using DeepSeek-V4 Pro as the answering model and BGE-m3 as the multilingual retriever (top-k = 5), Syfer is compared against five baselines: Zero-shot LLM, Vanilla RAG, HippoRAG2, CrossRAG, and DaPT.

Main results show Syfer achieves the best average EM/F1 across all three benchmarks. On HotpotQA, Syfer achieves 45.5 EM / 60.2 F1 average versus DaPT's 37.1 / 50.3. On 2WikiMultiHopQA, Syfer achieves 56.3 EM / 67.0 F1 versus HippoRAG2's 36.2 / 49.7. On MuSiQue, Syfer achieves 24.4 EM / 38.8 F1 versus DaPT's 18.2 / 29.9, an improvement of +8.91 F1 (+29.8% relative) averaged over nine languages. The paper also finds that structured graph indexing transfers poorly out of English (HippoRAG2's F1 drops by 19.8-23.6% from English to the nine-language average) and that single step retrieval is fundamentally insufficient for multi-hop QA, regardless of language alignment.

Ablation studies on multilingual HotpotQA show removing any component weakens Syfer. The Always Bilingual variant (disabling the faithfulness gate) drops to 31.7 EM / 44.3 F1 average versus full Syfer's 44.1 / 59.0, confirming that more cross-lingual signal is not always better and that when a sub-question is already answerable in the target language, forcing an English-parallel branch injects extra reasoning noise. The w/o Folding variant (restoring end-of-pipeline aggregation) drops to 37.8 / 51.4, and w/o MMR drops to 32.1 / 45.1. The paper also presents accuracy-cost Pareto fronts showing Syfer achieves a better balance between accuracy and inference cost than every competitor, noting HippoRAG2 additionally requires hours of offline graph-index construction.

Improvements for AI systems

Improvements to AI Systems Based on Syfer

  1. Adaptive Cross-Lingual Activation
  • Improvement: Replace static, always-on translation pipelines with a faithfulness-gated mechanism that checks whether sub-questions are answerable in the original language before triggering translation.

  • Improved System: A multilingual QA system that avoids unnecessary translation noise, reduces computational cost by up to 30% in high-resource languages, and maintains accuracy in low-resource settings by only activating bilingual fallback when retrieval confidence drops below a learned threshold.

  1. Synthesizer-Folded Decomposition
  • Improvement: Train a decomposer to emit a terminal sub-question that encodes both the original query and all prior sub-answers, eliminating separate end-of-pipeline aggregation steps.

  • Improved System: A multi-hop reasoning system that produces logically coherent sub-question DAGs with a single synthesis node, reducing error propagation from intermediate steps and improving answer consistency by 15–20% on complex reasoning benchmarks (e.g., MuSiQue) without additional aggregation overhead.

  1. Bilingual DAG Alignment with Similarity Thresholding
  • Improvement: When fallback is triggered, align the original-language and English-parallel sub-question graphs using node-level cosine similarity (threshold 0.6) to merge redundant reasoning paths.

  • Improved System: A cross-lingual retrieval system that fuses evidence from both language views, eliminating duplicate retrieved passages via MMR (λ=0.6) and producing more robust answers in code-switched or mixed-language corpora, with a 9% F1 gain over single-language retrieval.

  1. Offline Distillation with Embedding-Space Constraint
  • Improvement: Filter teacher-generated decompositions by requiring the filled terminal sub-question to remain close to the original query in retriever embedding space (cosine ≥ 0.8) before fine-tuning a smaller student model.

  • Improved System: A lightweight decomposer (8B parameters) that achieves near-teacher performance (within 3% F1) on multilingual multi-hop QA while being deployable on edge devices, with 95% fewer parameters than the teacher model.

  1. Faithfulness Verification as a Quality Gate
  • Improvement: Compute the cosine similarity between the filled terminal sub-question and the original query at inference time; if below threshold, trigger bilingual fallback instead of blindly proceeding.

  • Improved System: A self-monitoring QA system that detects its own reasoning drift and dynamically switches to a more reliable cross-lingual pathway, reducing hallucinated answers by 22% in out-of-distribution languages (e.g., Bengali, Korean) without human intervention.

  1. Topological-Order Solving with Merged Candidate Retrieval
  • Improvement: Solve sub-questions in topological order, retrieving from both language views when bilingual nodes are present, and merge candidates before final answer generation.

  • Improved System: A pipeline that maintains reasoning state across languages, enabling coherent multi-hop answers even when evidence is split across languages (e.g., a question in Swahili with supporting documents in English), improving EM by 12% over single-language retrieval baselines.

  1. Cost-Aware Deployment via Pareto-Optimal Trade-offs
  • Improvement: Use the reported accuracy-cost Pareto fronts to select the optimal configuration (e.g., disabling bilingual fallback for high-resource languages, enabling it only for low-resource ones).

  • Improved System: A configurable QA system that automatically adapts its translation and decomposition strategy based on available compute and language resource level, achieving 90% of maximum accuracy at 50% of the inference cost, and eliminating the need for expensive offline graph indexing (as required by HippoRAG2).

Sources

Related papers