SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension
Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
MDPI
cs.DL, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: 14 pages, 5 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 57/100
The gist: The paper introduces SCALE (Scientific Concept Aggregation via LLMs and Embeddings), a framework that extends the OpenAlex taxonomy with a new level of scientific Concepts positioned below the
Terminology
Summary
The paper introduces SCALE (Scientific Concept Aggregation via LLMs and Embeddings), a framework that extends the OpenAlex taxonomy with a new level of scientific Concepts positioned below the existing Topics level. The authors motivate this work by noting that "the increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science. They observe that while author keywords offer greater specificity,
their fragmentation, redundancy, and terminological variability limit their use as stable units of knowledge organization."
The core idea is to treat keywords not as isolated descriptors but to organize semantically related terms into coherent and interpretable conceptual units and integrate them within the existing disciplinary hierarchy.
The framework combines scientific text embeddings, large language models, and graph-based community detection to construct this additional layer, producing approximately 114,000 Scientific Concepts connected to OpenAlex Topics.
The paper situates SCALE across four lines of research:
-
Science mapping: Methods like VOSviewer reveal structures in bibliographic data but
are usually designed to reveal patterns, clusters, or relationships in bibliographic data rather than to create persistent conceptual units that a hierarchical classification system can incorporate.
-
Semantic representation: Transformer-based models such as SPECTER embed scientific documents as dense vectors, but
these models primarily represent documents or concepts in a semantic space rather than construct explicit structures for knowledge organization.
-
Topic modeling: Approaches like LDA and BERTopic uncover latent thematic structure, but
the resulting topics typically remain corpus-specific, can vary across different runs, and do not align with any pre-existing classification hierarchy.
-
Knowledge Organization Systems: The authors contrast prior taxonomy construction methods (TaxoGen, TaxoCom, TaxoAdapt) with SCALE: "Unlike these approaches, our framework does not construct or recursively complete an entire corpus-specific hierarchy. Instead, it preserves the four-level OpenAlex taxonomy and introduces a controlled fifth level of fine-grained Concepts below Topics through LLM-based granularity and field classification, field-specific semantic graphs, and Leiden clustering."
The paper also discusses the Computer Science Ontology (CSO) and Klink-2 as examples of combining automated methods with expert knowledge. Regarding OpenAlex specifically, the authors note that OpenAlex deprecated its earlier Concepts in favor of the Topics hierarchy, supplemented by Keywords (a fixed, closed set of 26,000+ terms, ten per topic), which do not constitute a hierarchical level of stable, reusable concept units with an explicit path to subfields, fields, and domains.
The framework processes approximately three million distinct author keywords from nearly two million MDPI scientific articles through four main steps:
Each author keyword is classified into one of three categories—high-level, Concept-level, or low-level—using OpenAI models. "High-level terms include existing OpenAlex 4 levels, examples such as Computer Science, Physics, Artificial Intelligence, and Epidemiology. Low-level terms are those too specific or contextual to serve as stable taxonomy units. Concept-level terms represent the target granularity of the fifth level, including methods, techniques, research themes, and specialized application areas." Only Concept-level keywords are retained.
Rather than building a single global graph, the authors construct field-specific semantic graphs because author keywords do not always carry a single stable meaning across all fields.
Each candidate keyword is assigned to one or more of the 26 OpenAlex Fields using an LLM. Each keyword is then enriched with a short field-specific definition generated by an LLM, and the keyword and its definition are then jointly encoded with SPECTER2, producing a contextualized semantic representation that captures both the keyword itself and its disciplinary meaning.
Semantic edges are created via HNSW nearest-neighbor search, with two keyword nodes linked only when they appear in each other's nearest-neighbor lists, and their semantic similarity exceeds a predefined threshold.
The Leiden algorithm identifies semantic communities within each field-specific graph. Clustering parameters, especially the Leiden resolution parameter, are tuned via grid search using two manually annotated sample sets: must-link pairs, which should be placed in the same cluster, and should-not-link pairs, which should be separated into different clusters.
After clustering, for each cluster, we prompt an LLM with the constituent candidate keywords and their field-specific definitions to generate a representative Concept name and a short description,
thereby formalizing each community as a named and described Concept.
Concept names and descriptions are encoded with SPECTER2 and compared against OpenAlex Topic embeddings via cosine similarity to narrow the candidate set. Then GPT-4o-mini evaluates this limited candidate set and selects the most appropriate Topic based on the Concept and disciplinary context.
Concepts can be attached to multiple Topics when a Concept matches more than one Topic above a similarity threshold,
allowing interdisciplinary themes to appear across multiple disciplinary paths.
An incremental maintenance procedure identifies emerging keywords based on the gain of its recent frequency percentile relative to its historical percentile.
Emerging candidates undergo similarity comparisons with existing Concepts and granularity filtering, then are attached to OpenAlex Topics and tagged with a version label.
The grid search tested the number of nearest neighbors, minimum similarity threshold, and Leiden resolution parameter. Based on 68 must-link and 68 should-not-link pairs curated by the authors, the final configuration used k=50, a minimum similarity threshold of 0.85, and a Leiden resolution parameter of 0.6, achieving 69.5% must-link precision and 87.5% should-not-link precision
on these challenging boundary cases.
The pipeline produced 113,892 Concepts. After hierarchical attachment, at least one Concept was associated with 98.8% of the 4,516 OpenAlex Topics (4,461 Topics).
The Domain-level distribution is: Physical Sciences with 65,615 Concepts (57.6%), Social Sciences with 18,814 (16.5%), Health Sciences with 15,437 (13.6%), and Life Sciences with 14,026 (12.3%).
Three access channels are provided: (1) Taxonomy Explorer, a prototype web application for hierarchical navigation with radial visualization, search, filtering, and export functions; (2) Atlantis, a public semantic map of the Concept layer in low-dimensional embedding space supporting 2D/3D views and paper retrieval; (3) a versioned open dataset on Hugging Face containing Domains, Fields, Subfields, Topics, and Concepts tables with stable identifiers, released under CC0 1.0.
The taxonomy has been deployed in a production paper classification pipeline for the MDPI corpus of nearly two million papers. 94% of Concepts were assigned to at least one paper. Papers received a median of 5 Concepts (IQR 5–6).
A human evaluation with 22 domain experts across 11 fields (110 papers; 220 paper×evaluator units; 3,958 Concept judgments)
compared two tagging configurations based on GPT-5.4 mini and Qwen3. Across all units, mean precision@5 was 88.9% for GPT-5.4 mini and 88.7% for Qwen3. Precision@1 was 94.5% and 95.9%, respectively.
Per-field precision@5 ranged from 75.0% to 97.0% for GPT-5.4 mini and 79.0% to 98.0% for Qwen3. Inter-rater agreement between two experts per field showed observed agreement of 0.76, with low Cohen's κ (0.19) but moderate prevalence-adjusted coefficients (PABAK=0.52, Gwet AC1=0.66).
The authors identify the main scientific contribution as the introduction of Concepts as a new unit for organizing scholarly knowledge,
consolidating semantically related candidate terms into named units located within explicit disciplinary paths.
SCALE differs from OpenAlex Keywords because "OpenAlex Keywords are derived from Topics and assigned to individual works, whereas SCALE Concepts introduce a separate, finer-grained structural level that is constructed bottom-up from author keywords and integrated below the existing Topic hierarchy."
The paper also discusses the potential for ontological extension.
The authors note that in the current implementation, Concepts are hierarchical taxonomic units rather than formally defined ontological entities,
and future work could decompose Concepts into finer ontological entities and use citation patterns, co-occurrence signals, publication metadata, and other semantic evidence to infer explicit, interpretable relations among them.
They suggest that existing knowledge graph infrastructure like MarmotGraph within EBRAINS could provide a broad scholarly layer to extend such graph-based representations beyond neuroscience and toward a more general organization of scientific knowledge.
The authors acknowledge that the must-link and should-not-link Concept pairs used for parameter selection were curated by the authors and may therefore introduce subjective bias,
that the quality and conceptual coherence of sampled keyword clustering results would benefit from independent human evaluation,
and that the current distribution of Concepts across Fields and Domains is uneven,
potentially reflecting disciplinary breadth differences, with future work needed to investigate methods to mitigate this imbalance and improve coverage of underrepresented areas.
Improvements for AI systems
-
Fine-grained scientific concept discovery: The AI system can automatically aggregate millions of author keywords into 114,000 coherent Scientific Concepts, filtering out overly broad or overly specific terms and preserving only stable, Concept-level units of knowledge organization.
-
Field-aware semantic disambiguation: The system can generate field-specific definitions for each keyword and jointly encode keyword-plus-definition with SPECTER2, enabling semantically ambiguous terms to be represented differently across 26 scientific fields rather than as a single global embedding.
-
Graph-based taxonomy construction via Leiden clustering: It can build field-specific semantic graphs using mutual nearest-neighbor search and similarity thresholds, then apply Leiden community detection to group semantically related keywords into interpretable clusters, with resolution parameters tuned against human must-link/should-not-link constraints.
-
LLM-driven concept naming and description: Given a cluster of candidate keywords and their definitions, the system can prompt an LLM to generate a representative Concept name and short description, producing formalized, reusable conceptual units rather than raw keyword clusters.
-
Hierarchical attachment to existing taxonomies: The system can attach each new Concept to one or more OpenAlex Topics by combining SPECTER2 cosine similarity retrieval with LLM-based reranking, enabling interdisciplinary Concepts to appear in multiple disciplinary paths while preserving the existing four-level hierarchy.
-
Scalable production classification: The improved system can assign Concepts to papers at scale, tagging each paper with a median of 5 Concepts, and achieve 89% precision@5 and 95% precision@1 in expert evaluation across 22 domain experts and 11 fields.
-
Incremental taxonomy maintenance: It can detect emerging keywords by comparing recent frequency percentiles against historical baselines, filter them for Concept-level granularity, attach them to existing Topics, and version them—allowing the taxonomy to stay current without full recomputation.
-
Interactive taxonomy navigation and exploration: The system can power web-based explorers with radial hierarchical visualization, 2D/3D semantic maps, search, filtering, and export functions, enabling researchers to navigate Concepts from Domains down to specific scientific specializations.
-
Interdisciplinary linking through multi-path attachment: Because a Concept can be attached to multiple Topics above a similarity threshold, the system can explicitly represent interdisciplinary themes under multiple disciplinary routes, allowing paper retrieval and browsing across field boundaries.
-
Human-in-the-loop parameter optimization: The system can tune graph clustering hyperparameters (nearest neighbors, similarity threshold, Leiden resolution) against human-annotated must-link and should-not-link pairs, improving cluster coherence and separation on challenging boundary cases.
Sources
- OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Qwen3 Technical Report