Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

arXiv:2608.11891 · cs.CY, cs.AI, cs.HC · Submitted 2026-08-16 · Read on arXiv

Avinash Agarwal, Vridhi Jain

Unique Identification Authority of India

cs.CY, cs.AI, cs.HC

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 18 pages, 11 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models across eight capability

Terminology

Summary

This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability.

The study's central concern is distinguishing genuine capability gaps from gaps in public disclosure. The authors note that different organizations evaluate their models using different datasets, benchmark suites, evaluation harnesses, and reporting practices, and that a national capability assessment that does not account for this will be incomplete at best. At worst, it will mistake an evaluation-reporting gap for a capability gap.

The paper makes four contributions: (1) it is among the first structured, eight-domain, benchmark-based comparative assessments of Indian foundation models; (2) it proposes a three-tier working definition of an Indian-developed AI model distinguishing fully indigenous development (Tier 1), India-led development with global components (Tier 2), and India-adapted foreign base models (Tier 3); (3) it proposes an exploratory four-dimension Benchmark Maturity Index (BMI) scoring standardization, participation, independent verification, and national coverage at the capability domain level; and (4) it identifies eight cross-cutting properties of the benchmark ecosystem.

The methodology compares three groups: global frontier models (GPT-5.6, Claude Opus 5, Kimi K3, Qwen3.8-Max), global comparable-scale models with 12–50 billion active parameters (Inkling, Qwen3.6-27B, Nemotron 3 Super), and representative Indian models (Sarvam-105B, Sarvam-30B, Param2, Krutrim-2). The analysis relies entirely on publicly reported benchmark results, distinguishing developer-reported from independently verified scores.

Key findings across domains include:

General-purpose reasoning: Indian models achieve strong scores on MMLU (Sarvam-105B, 90.6) and MATH-500 (Sarvam-105B, 98.6), but these benchmarks are now widely regarded as saturated, and frontier developers no longer report them. On GPQA Diamond, Sarvam-105B (78.7) trails the frontier leader by roughly 15 points. On HLE with tools, Inkling (46.0) substantially outperforms Sarvam-105B (11.2). No Indian model reports scores on ARC-AGI-2 or ARC-AGI-3.

Coding and software engineering: Indian models report strong scores on foundational benchmarks (Sarvam-30B: HumanEval 92.1, MBPP 92.7, LiveCodeBench 71.7), but no Indian model reports a result on any of the four agentic or multi-step software engineering benchmarks in this review (SWE-bench Pro, DeepSWE, Terminal-Bench 2.1, Vibe Code Bench).

Agentic AI and computer use: BrowseComp is the only benchmark in this study where all three model tiers report results on the same benchmark. Sarvam-105B (49.5) trails the frontier leader Kimi K3 (91.2) and comparable-scale leader Inkling (77.1) by substantial margins. No Indian or comparable-scale model reports scores on OSWorld-Verified or OSWorld 2.0.

Cybersecurity: Reported results are confined to the frontier tier... No comparable-scale or Indian model reports any cybersecurity benchmark result. This is the domain with the least publicly available evidence.

Vision and image understanding: No Indian model reports a result on any vision-understanding benchmark reviewed here. Participation is limited even within the frontier tier.

Video and multimodal understanding: The originally identified benchmarks (VideoMMMU, MVBench) had no usable public results, requiring substitution with Video-MME and MMVU. India's representative model Varya is a video-generation system, not a video-understanding model, so it is marked not applicable rather than zero.

Scientific research: No comparable-scale or Indian model reports results on any of these benchmarks (ProtocolQA, PaperBench, LAB-Bench, ProteinGym). Frontier participation is limited but non-zero.

Indic language capability: The evaluation landscape is fragmented across organizations and model generations, with different models reporting results on different benchmark suites and relatively limited overlap. Sarvam-30B reports IndiVibe and MILU; Sarvam-105B reports IndiVibe; Krutrim-2 reports BharatBench; PARAM-1 reports SANSKRITI and MILU; PARAM-2 reports a broader set. No single Indic benchmark has achieved broad adoption across major Indian model developers as a common evaluation standard.

The paper proposes the Benchmark Maturity Index (BMI), scoring each domain on four dimensions (each 0–2): Standardization, Participation, Independent Verification, and National Coverage. Results show: General Purpose AI scores High (7/8), Coding & Software Engineering scores High (6/8), Agentic AI & Computer Use scores Moderate (4/8), while Cybersecurity, Vision & Image Understanding, Video & Multimodal Understanding, and Scientific Research all score Low (2/8). Indic Language AI scores Low (2/6, scaled). The BMI refines, and in two cases revises, the maturity judgments that a purely qualitative review would produce.

The paper identifies eight cross-cutting patterns: benchmark saturation (MMLU, MATH-500 no longer discriminate among top models); absence of a common benchmark set across providers; divergent benchmark participation between Indian and global models; benchmark scarcity in emerging domains; concentration of Indian benchmark reporting in a single organization (Sarvam AI reports the broadest coverage by a substantial margin); absence of publicly reported Indian benchmark results in four domains; benchmark silence not being evidence of capability absence; and sensitivity of reported scores to evaluation methodology.

The authors emphasize: "We cannot determine, from the available public evidence, whether these gaps reflect genuine capability deficits, strategic reporting choices, or the absence of evaluation infrastructure. This ambiguity is itself the central finding of the paper."

Policy implications include: closing apparent benchmark performance gaps is not solely a model-training problem; ecosystem-level claims should be disaggregated by organization; and publicly funded models should have minimum benchmark disclosure requirements, including mandatory reporting of scores, disclosure of evaluation conditions, independent third-party evaluation, periodic re-evaluation, and public availability of methodology documentation.

The paper acknowledges limitations: reliance on publicly reported scores without independent reproduction, snapshot nature as of August 2026, potential undocumented harness differences, exclusion of organizations that don't publish technical reports, lack of a complete candidate list, and the BMI's lack of external validation.

Improvements for AI systems

Improvements to AI Systems:

  1. Add a Benchmark Maturity Index (BMI) scoring module to AI evaluation pipelines. The improved system can automatically compute a four-dimension score (Standardization, Participation, Independent Verification, National Coverage) for any model's reported benchmarks, flagging domains where scores are unverified or under-participated, and adjusting capability claims accordingly.

  2. Implement a disclosure-aware capability comparator that separates genuine performance from reporting gaps. The improved system can, when comparing models across organizations, normalize for benchmark saturation (e.g., MMLU, MATH-500), require evidence of independent verification, and explicitly label no reported result as unknown rather than zero capability, preventing false conclusions.

  3. Build a benchmark participation recommender for model developers. Given a model's architecture and intended use cases, the improved system can suggest which under-reported benchmarks (e.g., SWE-bench Pro for coding, OSWorld-Verified for agentic AI, ProtocolQA for scientific research) to adopt, based on the paper's identified gaps, to improve cross-provider comparability.

  4. Add a tier-aware evaluation mode that classifies a model's development origin (Tier 1: fully indigenous, Tier 2: India-led with global components, Tier 3: adapted foreign base) and adjusts evaluation expectations and benchmark selection accordingly, preventing unfair comparisons across tiers.

  5. Create an evaluation-condition transparency checker that parses technical reports for missing details (e.g., harness versions, prompting strategies, tool-use settings). The improved system can flag reports that omit these conditions, and can re-run benchmarks under standardized conditions to reduce sensitivity to methodology.

  6. Develop a cross-domain benchmark gap detector that identifies domains where a model reports no results (e.g., cybersecurity, vision, video understanding) and automatically generates a prioritized list of benchmarks to run, based on the paper's finding that silence often reflects missing infrastructure, not missing capability.

  7. Integrate a saturation-aware scoring feature that down-weights or replaces saturated benchmarks (e.g., MMLU, MATH-500) with harder alternatives (e.g., ARC-AGI-2, HLE with tools) when evaluating frontier or comparable-scale models, ensuring that reported scores remain discriminative.

  8. Add a national ecosystem evaluator that, for any country or region, aggregates benchmark participation across its model developers, computes a BMI per domain, and outputs a policy-ready report highlighting infrastructure gaps (e.g., no Indian model reports cybersecurity results) to guide funding and standardization efforts.

  9. Implement a benchmark substitution engine that, when a target benchmark has no public results (e.g., VideoMMMU), automatically proposes validated substitutes (e.g., Video-MME) with similar difficulty and domain coverage, as done in the paper, to maintain comparability.

  10. Create an independent verification flag that tags every benchmark score as developer-reported or independently verified, and, when only developer-reported scores exist, automatically reduces the confidence interval for that capability claim in downstream applications (e.g., model selection, procurement).

Sources

Related papers