When Do Anchor-Based Pointwise LLM Rerankers Help? Retriever Quality, Statistical Scope, and Anchor Design

arXiv:2608.10528 · cs.IR, cs.LG · Submitted 2026-08-15 · Read on arXiv

Utshab Kumar Ghosh, Shubham Chatterjee

Missouri University of Science and Technology

cs.IR, cs.LG

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: To be published in the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

DOI: 10.1145/3799682.3841055

Code: https://github.com/utshabkg/GCCP-reproduce

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: This paper investigates when anchor-based pointwise LLM reranking helps, using GCCP/PAGC as a representative method.

Terminology

Summary

This paper investigates when anchor-based pointwise LLM reranking helps, using GCCP/PAGC as a representative method. The study is reproduction-first, using reproduction as a starting point for controlled component-level stress testing. The authors address five research questions: (RQ1) whether GCCP/PAGC can be faithfully reproduced from the paper alone and what the process reveals about sensitivity to operational choices; (RQ2) whether the original statistical claims survive paired bootstrap testing with Holm-Bonferroni correction; (RQ3) how first-stage retrieval quality moderates the value of anchor-based reranking and score aggregation; (RQ4) whether spectral MDS anchor construction is necessary or if simpler alternatives suffice; and (RQ5) whether the mechanism transfers across LLM backbone families.

The initial reimplementation based only on the paper text achieved 0.24 nDCG@10 instead of the reported 0.66, revealing that several undocumented implementation details are necessary. After identifying and recovering eight such details, the authors reproduced reported results within 1.6% mean absolute nDCG@10 on TREC DL 2019/2020 and 1.9–4.5% on eight BEIR datasets.

The eight undocumented operational choices are grouped by severity. Three make-or-break choices individually produce broken rankings: (1) the decoder input string for T5 must be set to ' ' for RG-YN and ' Passage ' for GCCP, as an empty decoder input produces near-uniform token probabilities; (2) target-token case for RG-YN must be lowercase 'yes'/'no' rather than 'Yes'/'No'—this single choice moves nDCG@10 on DL19 RG-YN from 0.24 to 0.55 in isolation; (3) per-query min-max normalization before aggregation is required, as Equation 11 of the original paper writes the PAGC score as a plain average but the released code inserts normalization. Five performance-relevant choices include: (4) uppercase 'A'/'B' for GCCP target tokens; (5) spectral threshold θ=0.2, which is not stated; (6) BM25 parameters k1=0.9, b=0.4, which differ from Pyserini's stock defaults; (7) document truncation to 128 tokens; and (8) the nDCG implementation choice, where a hand-written routine differs from pytrec eval by about 2.5 points on TREC-COVID.

The authors applied paired bootstrap significance testing with Holm-Bonferroni correction across three comparison families: PAGC vs RG-YN, PAGC vs GCCP-alone, and GCCP vs RG-YN, each covering 22 primary settings. The correction changes the picture: for PAGC vs RG-YN, significant results drop from 15 of 22 to 12 of 22; for PAGC vs GCCP, from 8 of 22 to 5 of 22; for GCCP vs RG-YN, from 9 of 22 to 3 of 22.

The full method (PAGC) reliably improves over standard pointwise grading (RG-YN): Holm-significant in 12 of 22 settings, all positive. The contrastive signal alone (GCCP vs RG-YN) is directionally positive in 19 of 22 settings (p ≪ 0.001, sign test) but per-cell Holm-significant in only 3 of 22. Aggregation adds value over GCCP alone in only 5 of 22 settings and is significantly harmful on one dataset (DBPedia-Entity under E5, Δ=−0.0144, p Holm=0.032). The authors note that with small query sets like TREC DL (n q=43–54), non-significant results indicate absence of evidence rather than evidence of absence.

The authors tested whether anchor-based reranking still helps when the first-stage retriever is already strong by replacing BM25 with E5-base-v2. The main result is that stronger first-stage retrieval sharply reduces the marginal value of reranking: with BM25 on DL20, PAGC improves over the retriever alone by +0.197 nDCG@10, while with E5, the same reranker improves by only +0.013. On SciFact, PAGC with E5 even falls below the E5 baseline (−0.010).

Under E5 retrieval, aggregation does not add reliable value: PAGC does not Holm-significantly improve over GCCP alone on any of the 8 BEIR datasets, and on DBPedia-Entity (the largest dataset with n q=400), PAGC is significantly worse than GCCP alone. This negative result was confirmed with a different dense retriever (BGE-base-en-v1.5), where the effect was even larger (ΔPAGC-GCCP=−0.0190, p Holm<0.001). The authors found that mean Kendall's τ between RG-YN and GCCP scores is 0.43 in BM25 settings and 0.47 in E5 settings, which helps explain why aggregation gains shrink under E5 but does not explain the DBPedia-Entity negative result.

The spectral MDS anchor construction is not necessary. The authors compared it against three simpler anchor builders: a random candidate passage, the top-1 BM25 passage, and a top-3 sentence-interleaved composite. Spectral MDS is never the best anchor in TREC DL or BEIR ablations. On TREC DL, the Top-1 BM25 passage beats spectral MDS by +1.7 points for GCCP and +1.0 point for PAGC on DL19; on DL20, the Top-3 sentence-interleaved composite beats spectral MDS by +1.4 points for GCCP and +0.7 points for PAGC. On BEIR, spectral MDS finishes last on three datasets (TREC-COVID, Touché-2020, and Robust04) and is not the best on any of the 8 datasets. The aggregate ordering is Top-3 > Top-1 > Spectral > Random.

The authors also checked the Top baseline operationalization by testing eight variants, finding that Top-1 title-only matches the paper's reported Top average within 0.003, confirming the conclusion is not caused by a mismatched baseline. Hyperparameter sensitivity testing on the spectral anchor showed that sweeping m ∈ 5, 10, 15, 20, z ∈ 5, 10, 15, 20, and θ ∈ 0.1, 0.2, 0.3, 0.4 varies PAGC nDCG@10 by about ±1.5 points, which is meaningful relative to reported gains.

The mechanism transfers to decoder-only LLMs, including a 4-bit AWQ-quantized 72B model on a single 48 GB GPU. Qwen-2.5-72B-AWQ is the strongest model tested, reaching 0.7465 PAGC on DL19, surpassing both the reproduced Flan-UL2 (0.7095) and the paper's Flan-UL2 (0.7206) by more than 2.5 points. The DBPedia-Entity negative finding survives backbone change: on a compute-capped subset of 50 queries, RG-YN(0.3667) < PAGC(0.4139) ≈ E5(0.4168) < GCCP(0.4380), with paired bootstrap confirming PAGC<GCCP (Δ=−0.024, p raw=0.043).

At 7–8B scale, backbone family matters more than parameter count: Qwen-2.5-7B is competitive with Flan-T5-XL on DL19 PAGC (0.7212 vs 0.7030), while LLaMA-3.1-8B falls below Flan-T5-Large despite being 10× larger, and Mistral-7B-v0.3 underperforms on GCCP specifically. The Qwen-2.5 vs Mistral gap at DL19 PAGC (+0.078) exceeds the Flan-T5-Large vs Flan-UL2 gap (+0.026).

Alternative aggregation rules (Borda, Condorcet, Copeland, and α-weighted sweeps) do not materially change results. Under E5 retrieval, the best weighted variant is α=0.25 (giving more weight to GCCP), consistent with the finding that RG-YN adds less residual signal under stronger retrieval. The 3-component PAGC-RS-YN-GCCP variant reproduces the paper's Table 4 number within 0.4 points on TREC average but loses to 2-component PAGC on DL20 (−1.8 points), showing that adding a weak component can drag the mean rank. The 2-component PAGC outperforms all paper-reported listwise and pairwise baselines on both DL19 and DL20. Inference cost measurements show PAGC takes 4.47 seconds/query for Flan-T5-Large and 4.64 for Flan-T5-XL, with spectral MDS contributing less than 1% of total pipeline time.

The paper concludes that anchor-based pointwise reranking is effective but conditional. The core contrastive scoring idea is robust under rigorous statistical correction, but two design choices held fixed in the original paper are less reliable. First, combining the contrastive score with the standard pointwise relevance score helps when the first-stage retriever is BM25 but gives little or no benefit with stronger dense models such as E5. Second, the paper's more complex spectral MDS method for constructing the anchor is unnecessary—a much simpler anchor built by interleaving top-ranked sentences matches or outperforms it across datasets. The gains come mainly from contrastive scoring rather than from the more complex aggregation and anchor-construction choices, and they appear under narrower conditions than the original evaluation suggests. The authors emphasize that their contribution is not a new reranking architecture, but a controlled account of which parts of anchor-based pointwise reranking are load-bearing, which are incidental, and under what retrieval conditions the method remains useful.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive reranking activation based on first-stage retriever quality. Build an AI system that detects whether its initial retrieval is weak (e.g., BM25-like) or strong (e.g., dense embeddings like E5). When retrieval is weak, it activates anchor-based contrastive reranking with aggregation (PAGC). When retrieval is strong, it skips aggregation and uses only the contrastive score (GCCP) or bypasses reranking entirely, saving compute and avoiding the DBPedia-Entity-style degradation (up to −0.019 nDCG@10).

  2. Simplified anchor construction without spectral MDS. Replace the computationally complex spectral MDS anchor with a top-3 sentence-interleaved composite built from the first-stage retrieval results. This yields equal or better reranking performance across TREC DL and BEIR datasets while reducing pipeline complexity and eliminating the need for spectral hyperparameter tuning (m, z, θ), which currently varies performance by ±1.5 nDCG@10.

  3. Robust statistical significance gating for model deployment. Implement a paired bootstrap significance test with Holm-Bonferroni correction before committing to a reranking configuration. The system will automatically reject configurations that show only nominal improvements (e.g., 15→12 significant settings for PAGC vs RG-YN) and will flag settings where aggregation is significantly harmful (e.g., DBPedia-Entity under E5, Δ=−0.0144, p Holm=0.032), preventing silent quality regressions.

  4. Backbone-aware reranking selection. Use the finding that decoder-only LLMs (e.g., Qwen-2.5-72B-AWQ) outperform encoder-decoder models (Flan-UL2) by >2.5 nDCG@10 points on DL19. The improved system will select the strongest available backbone for the reranking task, with preference for Qwen-2.5 family over LLaMA-3.1 or Mistral at the 7–8B scale, where family choice matters more than parameter count (Qwen-2.5-7B: 0.7212 vs LLaMA-3.1-8B: below Flan-T5-Large).

  5. Compute-adaptive reranking with quantization. Enable 4-bit AWQ quantization for reranking models up to 72B parameters on a single 48 GB GPU without performance loss (Qwen-2.5-72B-AWQ achieves the best results). The system will automatically quantize when GPU memory is constrained, maintaining high reranking quality while reducing hardware requirements.

  6. Normalization-aware score aggregation. Implement per-query min-max normalization before any score aggregation, as the original paper's plain average (Equation 11) is broken without it. The improved system will always normalize scores across candidate passages per query before combining contrastive and relevance signals, preventing the 0.24→0.66 nDCG@10 failure mode.

  7. Token-case and decoder-input validation. Add automatic validation of decoder input strings (must be ' ' for RG-YN, ' Passage ' for GCCP) and target-token case (lowercase 'yes'/'no' for RG-YN, uppercase 'A'/'B' for GCCP). The system will verify these settings at initialization, avoiding the single-choice failure that drops nDCG@10 from 0.55 to 0.24.

  8. Dynamic aggregation weight adjustment. Use the finding that under strong retrieval, the optimal aggregation weight shifts toward the contrastive score (α=0.25 favoring GCCP). The improved system will estimate first-stage retrieval quality and adjust the α-weight accordingly, maximizing reranking gains when retrieval is weak and minimizing aggregation harm when retrieval is strong.

What the improved AI system can do:

  • Rerank search results with up to +0.197 nDCG@10 improvement over BM25 baselines, while avoiding negative impacts (−0.010 to −0.019) when dense retrieval is already strong.

  • Operate on a single 48 GB GPU with 72B-parameter models, achieving state-of-the-art reranking (0.7465 nDCG@10 on DL19) without specialized hardware.

  • Automatically choose between simple contrastive scoring (GCCP) and full aggregation (PAGC) based on retrieval conditions, saving up to 50% of reranking compute when aggregation adds no value.

  • Provide statistically validated performance claims, rejecting configurations that fail Holm-corrected significance tests, ensuring deployed systems only use improvements that are real, not noise.

  • Build anchors in milliseconds using top-3 sentence interleaving instead of spectral MDS, reducing anchor construction time to <1% of pipeline overhead while improving or matching performance on all tested datasets.

Abstract

Anchor-based pointwise LLM reranking scores each candidate against a shared reference passage to recover cross-document context at pointwise cost. We study when this actually helps, using GCCP/PAGC as a representative method. Our study is reproduction-first. We use reproduction as a starting point for a controlled component-level stress test of anchor-based pointwise reranking. Our initial reimplementation, based only on the paper text, achieves 0.24 nDCG@10 instead of the reported 0.66, revealing that several undocumented implementation details are necessary to reproduce the method. After identifying and recovering eight such details, we reproduce the reported results within 1.6% and use the validated implementation for controlled analysis. We find that the core contrastive scoring idea is robust under rigorous statistical correction. However, two design choices held fixed in the original paper are less reliable. First, we find that combining the contrastive score with the standard pointwise relevance score helps when the first-stage retriever is BM25, but gives little or no benefit when the first-stage retriever is a stronger dense model such as E5. Second, the paper's more complex method for constructing the anchor is unnecessary. A much simpler anchor, built by interleaving the top-ranked sentences, matches or outperforms it across datasets. These findings are consistent across different LLM backbones, including a 4-bit quantized 72B model. Overall, anchor-based pointwise reranking is effective, but its gains come mainly from contrastive scoring rather than from the more complex aggregation and anchor-construction choices, and they appear under narrower conditions than the original evaluation suggests.

Sources

Related papers