HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

arXiv:2608.10584 · cs.AI, cs.DL · Submitted 2026-08-12 · Read on arXiv

Xiaokang Qu, Yiting Lin

University of Science and Technology of China

cs.AI, cs.DL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: remove copyright information

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 50/100

The gist: HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment Abstract Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic

Terminology

Summary

HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

Abstract

Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing scholar assessment methods predominantly rely on bibliometric indicators and reputation proxies, while recent large language model (LLM)-based approaches mainly focus on evaluating individual research papers rather than comprehensively assessing scholars. We argue that scholar assessment should be formulated as an evidence-driven reasoning problem that jointly considers intrinsic research quality and externally verifiable scholarly behavior. To this end, we propose HexEval, an evidence-driven hexagonal framework for multidimensional scholar assessment. HexEval explicitly organizes scholar assessment into two complementary evidence layers. The intrinsic layer evaluates anonymized representative works along three dimensions, namely research rigor, methodological innovation, and scientific contribution, whereas the external layer characterizes scholars through knowledge translation, research coherence, and academic impact using heterogeneous evidence collected from GitHub, Lens, OpenAlex, and other publicly verifiable sources. Instead of producing opaque aggregate scores, HexEval preserves intermediate evidence, dimension-specific rationales, and verification signals throughout the evaluation process, enabling interpretable and auditable scholar profiles. Experiments across all six dimensions show dimension-dependent agreement with human or external reference criteria: structured calibration improves absolute agreement for intrinsic quality, while the external modules recover broad trajectory and ordinal impact signals. These results support evidence-driven reasoning over heterogeneous scholarly evidence as a promising paradigm for auditable AI-assisted scholar assessment, while exposing the coverage and attribution limitations of public scholarly data.

Introduction

Scholar assessment supports faculty recruitment, funding allocation, academic promotion, award nomination, and talent discovery. Existing methods predominantly rely on bibliometric indicators such as citation counts, publication numbers, h-index, field-normalized metrics, and publication venues. Although useful for measuring scholarly visibility and influence, these indicators mainly capture research outcomes rather than intrinsic research quality. They are also affected by field-specific citation practices, academic age, venue prestige, cumulative advantage, and reputation effects.

Recent large language models (LLMs) have enabled structured scientific-document understanding and paper-level quality assessment with encouraging agreement with human judgments. However, scholar assessment requires reasoning over broader evidence, including representative works, long-term research trajectories, knowledge translation, and scholarly impact. Paper-level content reasoning and scholar-level metric aggregation therefore remain largely disconnected.

We reformulate scholar assessment as a dual-layer evidence reasoning problem. The intrinsic layer evaluates whether representative research is rigorous, innovative, and scientifically valuable, while the external layer characterizes how research is translated, sustained, and recognized in the scholarly ecosystem. This separation distinguishes scientific merit from downstream influence and produces more interpretable evaluation results.

Based on this formulation, we propose HexEval, an evidence-driven hexagonal framework that represents each scholar through six dimensions. The intrinsic layer evaluates anonymized representative works in terms of research rigor, methodological innovation, and scientific contribution. The external layer evaluates knowledge translation, research coherence, and academic impact using publicly verifiable evidence from GitHub, Lens, and OpenAlex. Each dimension uses its own evidence source, scoring procedure, and evaluation protocol, while preserving intermediate evidence, rationales, and verification signals for auditing.

The main contributions are:

  • We formulate automated scholar assessment as a dual-layer evidence reasoning problem that separates intrinsic research quality from externally verifiable scholarly behavior.

  • We propose HexEval, a six-dimensional framework that independently evaluates research rigor, methodological innovation, scientific contribution, knowledge translation, research coherence, and academic impact while preserving auditable evidence.

  • We evaluate D1–D5 on public or curated reference data and operationalize D6 using the reproducible OpenAlex h-index, with explicit reporting of evidence coverage, attribution, sampling, and source limitations.

Related Work

Scholar-level assessment traditionally relies on publication counts, citations, the h-index, and field-normalized measures; later work adds topics, authorship, time, collaboration, venue, and expert interpretation. These indicators capture productivity and accumulated visibility more directly than intrinsic quality, and are affected by field practices, career length, coverage, authorship, and cumulative advantage. Their association with peer judgment is field-dependent; responsible-assessment guidelines therefore treat them as context rather than substitutes for qualitative evidence.

LLMs now assist peer review through critique generation, methodological diagnosis, score prediction, and review improvement. Their plausible outputs remain limited by long-document understanding, paper-specific criticism, technical error detection, and score reliability. Structured rubrics, retrieval, multi-stage reasoning, and agents improve consistency, but deployment evidence favors reviewer assistance over autonomous decisions. Direct quality estimation shows weak-to-moderate and field-dependent agreement with human judgments; repeated sampling, prompt design, and input selection materially affect results.

Evidence-grounded methods support judgments with inspectable information and sources. Retrieval-augmented generation connects LLMs to external knowledge, while evidentiality-guided generation models evidence relevance and support; retrieval alone, however, does not guarantee valid support. Scientific claim verification combines retrieval, support/refutation classification, and rationale extraction, as in SciFact and its extensions. Attribution-oriented methods such as RARR retain source links while revising unsupported claims. These works establish evidence–claim alignment and attribution as requirements for auditability. These methods mainly address QA, factual generation, or claim/document verification. Scholar assessment requires reasoning over works, career stages, artifacts, and impact channels; HexEval extends evidence-grounded reasoning to this setting with inspectable papers, software/patent records, trajectories, and citation traces.

HexEval Framework

HexEval represents a scholar by scientific outputs and externally observable scholarly traces. Let s denote a scholar, Ps the supplied representative works, and Is the identity metadata required by attribution-dependent dimensions. The framework produces H(s) = [D1(s), D2(s), D3(s), D4(s), D5(s), D6(s)]. The first three dimensions measure intrinsic research quality and the last three measure externally observable scholarly behavior. Intrinsic scores use anonymized works; external scores use identity-linked evidence. The two paths are therefore complementary but not substitutable.

Each dimension has its own input schema, prompt or deterministic computation, score scale, and evidence record. HexEval does not impose a universal weighted sum; it returns the profile and evidence package O(s) = H(s), E1(s),..., E5(s), E6(s), V(s), where Ei(s) contains inputs, structured outputs, rationales, and source metadata, and V(s) contains validation, coverage, and uncertainty information. Any downstream aggregation is application-specific and is not part of the default output.

For visualization only, a dimension score can be mapped from its native scale [li, ui] to a percentage. This affine mapping does not make the dimensions commensurate in a substantive sense and does not define a global scholar ranking.

Intrinsic Research Quality Assessment

The intrinsic pathway receives only anonymized papers. For dimension d ∈ 1, 2, 3 and paper p, the LLM returns a direct score qd,p, subdimension scores, rationale, and diagnostic evidence. Identity, institution, venue, citations, and author metadata are excluded. The scholar-level direct score is the mean over valid works. The three dimensions are not collapsed into one intrinsic score. With human labels, a dimension-specific Ridge calibrator is fitted for held-out validation. The calibrator is used only when this fitted mapping is available; otherwise, the direct mean is retained. Thus, benchmark calibration is not presented as an unsupervised scoring rule.

Representative-Work Anonymization

To reduce identity and reputation cues, HexEval maps each representative PDF p to an anonymized Markdown document. The PDF is converted to structured Markdown while retaining scientific content, including equations, tables, figures, and captions. Rule-based filters remove names, affiliations, emails, acknowledgments, funding, and revealing citation metadata. An LLM cleaning stage removes residual institutional, group, project, and acknowledgment cues. Methodological details, settings, formulations, results, limitations, and conclusions are retained. The same anonymized documents are supplied to D1–D3, which use independent prompts, rubrics, extraction procedures, and validation protocols.

To audit residual identity leakage, we sampled 100 representative works and used DeepSeek-V4-Flash as an automatic detector for explicit identity-revealing cues. The residual leakage rate decreased from 0.53 after Stage 1 (MinerU conversion) to 0.52 after Stage 2 (regex-based cleaning), and further to 0 in this sampled audit after Stage 3 (LLM anonymization). Anonymization mitigates but does not eliminate leakage: method names, datasets, benchmarks, writing style, or distinctive contributions may remain identifying. It is therefore a bias-mitigation mechanism, not a guarantee of identity-free evaluation.

D1: Research rigor. D1 evaluates whether claims are supported by sound methods and evidence. Its seven criteria are methodological validity, evidence adequacy, evaluation design, comparisons and controls, statistical or logical rigor, reproducibility and transparency, and limitation/claim calibration. The output contains criterion rationales and serious or minor weaknesses; unsupported claims reduce the relevant criterion rather than incur a reputation-based penalty.

D2: Methodological innovation. D2 evaluates originality in technical context. The main benchmark path receives extracted abstract, introduction, related-work/background, method, and conclusion sections from the anonymized paper, without identity or reputation cues. Its criteria are core originality, technical distinctiveness, nontriviality, and novelty-claim specificity. The evaluator distinguishes new mechanisms from new applications, tuning, implementation changes, and performance gains. An optional OpenAlex prior-work mode restricts candidates to works before the target year but is not used for the reported benchmark. The output includes the central method claim, novelty type, comparison rationale, and q2,p.

D3: Scientific contribution. D3 measures scientific significance and usefulness rather than method novelty alone. Its criteria are problem importance, contribution substance, result value, and generality/reusability. The output contains the main contribution, strengths, serious and minor weaknesses, rationale, and q3,p, allowing narrow but correct work to differ from broadly reusable work.

For all three dimensions, each saved score links to extracted paper content, subdimension values, and rationale; the scholar-level mean is therefore an aggregation of inspectable paper-level judgments.

External Scholarly Behavior Assessment

The external pathway operates at the scholar level because these dimensions describe observable scholarly behavior rather than the quality of an individual paper. The evidence sources, attribution checks, and scoring rules are kept separate for each dimension.

D4: Knowledge translation. D4 measures validated translation of research into reusable software and patented or otherwise documented intellectual property. Software evidence is collected from GitHub and public project records. A community project contributes only when its evidence confidence is strong or moderate, its repository is reachable, and the scholar's contributor attribution is verified. Personal repositories are restricted to owned, non-forked repositories with at least 100 stars. Patent candidates are deduplicated into families; only strong or moderate, non-review-required families receive positive weight. Weak, unverified, and review-required items remain in the audit trail but do not contribute to the score.

For a repository r, the software impact is Ir = 0.60L(starsr; 10000) + 0.30L(forksr; 3000) + 0.10ar, where ar = 1.0 for activity within two years, 0.7 for activity three to five years old, 0.4 for older or unknown activity, and 0.2 for archived repositories. A selected project receives gr = er wr (1 + 3Ir), where the implemented evidence weights are er = 1.0 for strong community evidence, 0.6 for moderate community evidence, and 0.45 for a qualifying personal repository. The type weights are wr = 1.20 for community projects, 1.00 for ordinary personal repositories, and 0.70 for personal repositories matching the low-value repository patterns. The top eight community projects and top ten personal repositories are retained, and Tsoft(s) = S(Gs; 12).

For a validated patent family f, the impact term is If = 0.50L(citationsf; 100) + 0.25L(familySizef; 10) + 0.25of, where of is the ownership score. The family score is gf = ef (1.5bf + 3If), where bf = 1.0 for a granted family and bf = 0.6 otherwise, while ef = 1.0 for strong evidence, 0.5 for moderate evidence, and 0 for weak or review-required evidence. The positive family-score sum is saturated as Tpat(s) = S(Ps; 8). The final score is D4(s) = 100[0.60Tsoft(s) + 0.40Tpat(s)]. This score represents validated translation evidence rather than complete individual contribution. Evidence strength, attribution status, validation status, and source identifiers are retained for auditing.

D5: Research coherence. D5 estimates research coherence from a sparse chronological sample of a scholar's fuller career trajectory. We reuse 110 preselected computer-science scholars and retrieve their author-matched OpenAlex publication records. The full-career package is divided into five chronological bins. Papers enter the primary coherence corpus only when year, title, and abstract are all available; excluded records remain documented in the package audit trail.

The adjudicated reference is constructed from the full-career packages. GLM and DeepSeek independently score the five coherence dimensions: thematic consistency, temporal continuity, main-thread clarity, related-branch integration, and low fragmentation. ChatGPT does not produce an independent score; it acts only as an anonymous judge of the two scorer outputs. The final reference value is the mean of the five final adjudicated dimension scores and is referred to as an adjudicated multi-LLM reference annotation, not as ground truth. A fixed 20/90 development/test split is used, and the test reference and sampling manifest are frozen before evaluation.

For the sparse evaluation, the same manifest supplies three papers per bin, or at most 15 papers per scholar, and each scholar is evaluated over five repeated samples. The evaluator receives only year, title, and abstract. If cs,j is the overall coherence score for repeat j, the reported prediction is D5(s) = (1/5)Σcs,j, with σ5(s) = SD(cs,1,..., cs,5).

D6: Academic impact. D6 is an OpenAlex-based bibliometric anchor rather than a newly proposed composite index or a separately labeled benchmark. For each resolved scholar, we retrieve the OpenAlex author record and the author-matched work records. The canonical D6 value is the author-level h-index in authors.summary stats.h index. If that field is unavailable, the implementation uses a documented h-index fallback computed from the retrieved valid works. Other OpenAlex fields are retained as auditable evidence but are not combined into an additional score. These fields include total citations, i10-index, works count, two-year mean citedness, FWCI, citation-normalized percentiles, recent citations and works, top works, pagination status, and the retrieval timestamp. Thus, D6 provides a reproducible and updateable citation-based impact signal, while D1–D5 capture non-bibliometric properties that h-index cannot represent.

Experimental

Evaluation Dataset Construction

Because the dimensions use different evidence sources and validation roles, D1–D3 rely on human judgments, D4 uses curated scholarly evidence, D5 uses adjudicated multi-LLM references, and D6 uses the OpenAlex h-index without a separately constructed benchmark.

Intrinsic Quality Dataset: We use public OpenReview reviews with dimension-level human judgments. D1 uses NeurIPS soundness annotations, D2 uses ICLR 2022 technical novelty and significance scores sampled by novelty quartiles, and D3 uses contribution-related annotations. Each dimension contains 300 samples.

External Scholarly Behavior Dataset: D4–D5 are evaluated at scholar level because they characterize observable scholarly behavior rather than paper quality. D4 groups scholars into high, middle, and low translation levels using awards and publicly documented software/IP outcomes (90 scholars). D5 uses author-matched OpenAlex publication corpora for preselected computer-science scholars (110 scholars). Full-career evidence packages are scored independently by GLM-5.2 and DeepSeek-V4-Flash, and GPT-5.5 acts as an anonymous adjudicator of their outputs to produce the final coherence references. Evaluation uses sparse chronological publication samples.

Experimental Setting

Local models are served with vLLM. Representative PDFs are converted to structured Markdown using MinerU, processed by rule-based filters, and anonymized using Qwen2.5-72B-Instruct-AWQ with temperature 0 and a maximum output length of 1,200 tokens. All primary LLM-based predictions for D1–D3 and D5 use Qwen3.6-27B with a 32K context window. D1 uses temperature 0, D2 and D3 use temperature 0.1, and D5 uses temperature 0. The maximum output lengths are 3,000 tokens for D1, 1,800 tokens for D2 and D3, and 2,048 tokens for D5. Ridge calibration is trained only on the calibration split and evaluated on held-out test data. Model versions, prompts, retrieval dates, random seeds, and sampling manifests are recorded for reproducibility.

Baselines: For D1–D3, we compare HexEval with direct overall scoring, chain-of-thought prompting, self-reflection prompting, and the unweighted mean of structured subdimension scores. All methods use the same evaluator model, anonymized paper inputs, and dimension-specific scoring scales. HexEval additionally applies Ridge calibration trained only on the calibration split. For D4, we compare the full score with single-indicator and count-based baselines derived from the same verified software and patent evidence. For D5, we compare HexEval-D5 with TF–IDF, SPECTER2, and direct zero-shot LLM scoring. All D5 methods use the same frozen five-bin sampling manifest, three papers per bin, and five repeated samples.

Results

Intrinsic Research Quality Evaluation

We evaluate the three intrinsic dimensions by comparing model scores with human peer-review judgments. D1, D2, and D3 use reviewer-averaged rigor, technical novelty, and contribution scores as reference labels, respectively. Each dimension contains 100 held-out test papers disjoint from the calibration set. We fix the evaluator model to Qwen3.6-27B and compare direct overall scoring, chain-of-thought reasoning, self-reflection, simple averaging of structured subdimension scores, and HexEval with Ridge calibration. We report Spearman's rank correlation (ρ), mean absolute error (MAE), and the proportion of predictions within 0.5 points of the human score (Acc@0.5).

D1: Research Rigor. HexEval performs strongest on rigor. Ridge calibration achieves the highest rank correlation (ρ =.467) and the lowest MAE (.361), improving over direct overall scoring on both measures. Direct scoring obtains the highest Acc@0.5 (.73), but its larger negative bias indicates less calibrated absolute scoring. Overall, structured rigor decomposition plus learned aggregation improves the consistency and scale alignment of soundness assessment.

D2: Methodological Innovation. Innovation remains the most difficult intrinsic dimension. Self-reflection gives the highest rank correlation (ρ =.436), and CoT gives the highest Acc@0.5 (.70). HexEval achieves the lowest MAE (.378), indicating better score-scale calibration, but it does not improve ordinal ranking. This suggests that calibration reduces systematic score error, while relative novelty ordering is still sensitive to model judgment.

D3: Scientific Contribution. For contribution, HexEval substantially improves absolute agreement: MAE drops from.516 under direct scoring to.295, and Acc@0.5 increases from.54 to.86. However, direct overall scoring still gives the highest rank correlation (ρ =.291). Thus, structured calibration is effective for aligning contribution scores with the human scale, but the relative ordering of papers remains challenging.

Overall, HexEval most clearly improves score calibration and absolute agreement across intrinsic dimensions. Ranking gains are strongest for D1, whereas D2 and D3 show that scale calibration and ordinal ranking are not always improved by the same mechanism.

External Scholarly Behavior Evaluation

We evaluate D4 against external knowledge-translation tiers and D5 against the frozen adjudicated coherence reference. D6 is implemented as a source-backed bibliometric indicator rather than evaluated against a separately constructed label set.

D4: Knowledge Translation. At the main thresholds of 30 and 50, the fixed D4 score achieves.611 accuracy and.601 Macro-F1 on the balanced three-level benchmark. The full D4 score obtains the highest High–Low AUC (.930), exceeding GitHub max stars (.894) and raw software-plus-patent counts (.893), the two strongest baselines. Its F1@k and Acc.@k values are both.867, tying several simpler indicators. The principal advantage of the complete D4 formulation therefore lies in separating high- and low-translation scholars across the full ranking, rather than in improving top-k retrieval alone. The continuous ranking is independent of the reporting thresholds, whereas the derived three-level classification remains sensitive to the selected cut points. Overall, these results support the joint use of verified software and patent/IP evidence, while retaining evidence attribution, validation status, and source identifiers for auditing.

D5: Research Coherence. On the frozen 90-scholar test set, HexEval-D5 achieves ρ =.743, τ =.620, MAE =.266, and Acc.@0.5 =.878. Although direct zero-shot scoring yields marginally higher rank correlations (ρ =.746, τ =.622), HexEval-D5 substantially improves absolute agreement, reducing MAE from.576 to.266 and increasing Acc.@0.5 from.489 to.878. TF–IDF and SPECTER2 obtain lower rank correlations, suggesting that structured evaluation better aligns sparse publication samples with the full-career coherence reference.

D6: Academic Impact. D6 uses the OpenAlex author-level h-index as its canonical impact indicator. Total citations, i10-index, works count, two-year mean citedness, yearly citation counts, FWCI, citation-normalized percentiles, recent activity, top works, and retrieval and author-resolution metadata are retained as supporting evidence and are not combined into a new composite score. Because D6 is a source-backed operational indicator rather than a separately labeled benchmark, no D6 baseline-comparison or ablation table is reported.

Case Study

We illustrate the complete HexEval pipeline using an anonymized scholar, denoted as Scholar A. Six recent representative papers are anonymized and evaluated independently by D1–D3, while D4–D6 use identity-resolved GitHub, Lens, and OpenAlex evidence. Identifying information is removed from the reported case study, while the underlying identity links are retained only for evidence attribution. The case study illustrates the reporting format and evidence flow rather than population-level accuracy. D5 estimates research coherence from five repeated chronological samples of the author-resolved OpenAlex publication corpus. D6 uses the OpenAlex author-level h-index as its primary impact indicator, while other bibliometric fields are retained only as supporting evidence. Both dimensions are interpreted together with evidence coverage and attribution status.

Discussion

HexEval is intended as an evidence-grounded decision-support framework rather than a replacement for expert judgment. Its main advantage is not a single superior ranking, but the separation of heterogeneous signals that conventional metrics collapse: D1–D3 characterize the quality of representative research, whereas D4–D6 describe knowledge translation, research coherence, and bibliometric impact. This separation improves interpretability, but the six dimensions should not be treated as interchangeable or mechanically averaged. The framework remains limited by the completeness of public records, author disambiguation, and source-specific biases. GitHub and patent evidence may underrepresent some disciplines, OpenAlex coverage varies across fields, and the h-index remains sensitive to career age and citation practices. HexEval should therefore support, rather than determine, high-stakes assessment decisions.

Improvements for AI systems

Based on this paper, I can improve AI systems in the following specific ways:

1. Dual-Layer Evidence Reasoning Architecture

  • Implement a system that separates intrinsic quality assessment (anonymized work evaluation) from external behavioral assessment (identity-linked evidence)

  • The improved system can evaluate both the inherent merit of outputs and the verifiable external impact, preventing conflation of quality with reputation

2. Structured Multi-Dimension Scoring with Calibration

  • Replace single-score outputs with dimension-specific scoring (rigor, innovation, contribution, translation, coherence, impact)

  • Apply learned calibration (Ridge regression) trained on human labels to align model scores with human scales

  • The improved system can produce scores with better absolute agreement (MAE reduced from.516 to.295 for contribution) rather than just relative rankings

3. Evidence-Preserving Output with Audit Trails

  • Maintain intermediate evidence, rationales, verification signals, and source metadata for every score

  • The improved system can provide fully auditable assessments where every claim traces to specific evidence, enabling human verification and reducing blind trust in opaque outputs

4. Anonymization Pipeline for Bias Mitigation

  • Implement a three-stage anonymization process (structured conversion → rule-based filtering → LLM cleaning) that removes identity cues while preserving scientific content

  • The improved system can reduce identity leakage from 53% to 0% in sampled audits, enabling fairer evaluation of work independent of author reputation

5. Heterogeneous Evidence Integration with Confidence Weighting

  • Combine evidence from multiple sources (GitHub, patents, OpenAlex) with explicit confidence levels, attribution checks, and validation status

  • The improved system can distinguish verified from unverified evidence, weighting only strong/moderate evidence while retaining weak evidence in audit trails

6. Sparse Chronological Sampling for Trajectory Assessment

  • Sample limited data points across time bins with repeated sampling to estimate stability (σ5)

  • The improved system can assess long-term coherence from sparse data with high absolute agreement (MAE.266 vs.576 for direct scoring) while quantifying uncertainty

7. Dimension-Specific Prompt and Rubric Design

  • Use independent prompts, rubrics, and extraction procedures for each dimension rather than a generic evaluation prompt

  • The improved system can evaluate different aspects (methodological validity vs. novelty vs. contribution) with tailored criteria, avoiding cross-dimension contamination

8. Explicit Coverage and Limitation Reporting

  • Automatically report evidence coverage, attribution limitations, sampling constraints, and source biases alongside scores

  • The improved system can flag when public data underrepresents certain fields or when author disambiguation is uncertain, preventing overconfident conclusions

9. Non-Aggregative Profile Output

  • Return a multidimensional profile with separate scores rather than a single composite index

  • The improved system can support application-specific aggregation downstream without imposing a universal weighting scheme, preserving interpretability

10. Reproducible Operationalization of Impact Metrics

  • Use source-backed indicators (e.g., OpenAlex h-index) with documented fallbacks and retrieval timestamps

  • The improved system can produce updateable, reproducible impact signals that are transparent about their provenance and limitations

Sources

Related papers