vSV-ViT: Variable-size SuperVertex Vision Transformer for Cortical Surface Learning in Alzheimer's Disease

arXiv:2605.26514 · cs.CV, cs.AI, cs.LG · Submitted 2026-05-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "vSV-ViT: Variable-size SuperVertex Vision Transformer for Cortical Surface Learning in Alzheimer's Disease".

Jane: Confirming Alzheimer’s disease (AD) typically relies on positron emission tomography (PET), which remains costly and invasive, motivating the use of structural MRI-based prescreening.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So Jane, this paper is called "vSV-ViT: Variable-size SuperVertex Vision Transformer for Cortical Surface Learning in Alzheimer's Disease," and it's tackling the whole problem of how to use structural MRI data to screen for AD before we go into those more invasive PET scans.

Jane: It sounds like they are addressing a major hurdle in using deep learning on brain surfaces, specifically the issue where standard methods create overlapping patches that mix different regions together.

Lu: The authors are Geonwoo Baek and Ikbeom Jang from Hankuk University of Foreign Studies, and their focus is clearly on solving that geometric problem inherent to brain data structures.

Meng: From an engineering standpoint, it’s interesting how they are trying to handle the variability of different cortical regions when making these patches.

Lalam: I see a lot of potential here for improving how we interpret complex visual data across different scales, which could really enhance our internal understanding capabilities.

The paper's summary: Tom: Essentially, the core idea is introducing a novel method called CSV-ViT, which uses these variable-sized patches to avoid mixing regions and include unwanted stuff like the medial wall.

Jane: That sounds much more precise than fixed patching; instead of using uniform patches across the whole cortex, they partition it into Cortical Supervertices or CSVs constrained by specific Regions of Interest or ROIs.

Lu: The methodology involves a multi-stage process: first registering surfaces to ico6, then partitioning based on the Desikan–Killiany atlas while actively trying to prevent ROI mixing through specific reassignment rules for small fragments.

Meng: So, they’re building these connected supervertices within each ROI up to a certain size limit L or H, which is a clever way to maintain regional specificity.

Lalam: This partitioning process seems vital because it directly tackles the problem of region-specific representations that uniform patches often lose; it’s about making sure the AI focuses on what matters locally.

The paper's improvements: Tom: What really stands out is how they tackle the limitations of previous surface models, like SiT, by ensuring ROI preservation and preventing vertex duplication at patch boundaries.

Jane: They specifically aim to ensure that the patches are variable in size and shape, which is necessary because cortical regions have such different sizes and geometries.

Lu: The improvements include a global planning stage where they search for feasible bounds L and H, focusing on reducing ROI-wise imbalance by minimizing a specific metric involving the average vertex count per ROI.

Meng: From a practical perspective, that balancing step is critical; ensuring the supervertex size stays within those bounds while keeping connectivity is something we have to model carefully in any deployment.

Lalam: The paper shows that this entire system—the ROI-preserving tokenization, the vertex-based partitioning, and the variable-sized patches—all contribute positively to performance across different AD tasks like diagnosis (Dx), amyloid positivity (Aβ), and tau positivity.

Conclusion: Tom: So to wrap up, the vSV-ViT framework successfully generates ROI-preserving CSVs with variable sizes, which feeds into a mask-aware embedding in the Vision Transformer to classify AD status using only T1 MRI data like cortical thickness and curvature.

Jane: It gives clinicians a way to get a prescreening tool for AD pathology prediction, which could be really useful before we move on to more expensive tests like PET or CSF analysis.

Lu: The implication is that we can start building models that are inherently better at understanding the non-Euclidean topology of the brain by treating it as a graph with supervertices.

Meng: From an engineering standpoint, this means if we can reliably generate these CSVs and feed them into a ViT, we have a robust pipeline for triage that doesn't rely on external confirmation methods immediately.

Lalam: I think what’s most exciting is how this advance in structural learning could be applied across other complex medical imaging modalities, potentially improving the cultural understanding of how we model human anatomy in AI systems.

Geonwoo Baek, Ikbeom Jang

Department of Computer Science and Engineering, Hankuk University of Foreign Studies

cs.CV, cs.AI, cs.LG

Submitted: 2026-05-26

Updated: 2026-10-04

Comments: Accepted to ACCV 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Confirming Alzheimer’s disease (AD) typically relies on positron emission tomography (PET), which remains costly and invasive, motivating the use of structural MRI-based prescreening.

Key concepts

Cortical Supervertices (CSVs)
These are non-overlapping clusters of vertices on the brain's surface, constrained by specific Regions of Interest (ROIs). The partitioning process ensures that these supervertices maintain connectivity within their assigned ROI while strictly avoiding overlap between different ROIs, creating a structured representation of the cortex.
CSV-ViT Architecture
This is a Vision Transformer model designed to process the cortical surface data. Instead of using fixed patches, it ingests variable-sized CSVs. A mask-aware embedding mechanism then converts these variable-sized vertex groups into single tokens, allowing the model to learn relationships across different cortical regions effectively.
ROI-Preserving Partitioning
This is a multi-stage process used to create the CSVs. It involves registering surfaces, partitioning based on an atlas, and then refining the structure by reassigning small fragments to adjacent ROIs. This ensures that the resulting supervertices respect regional boundaries and do not mix features from different areas.
Mask-Aware Patch Embedding
This is a specific technique within CSV-ViT where a binary mask is generated based on the CSV map. This mask is used to zero out padded entries in the input tensor before embedding. This allows the model to focus its attention specifically on the features contained within each individual, variable-sized supervertex.

Terminology

Summary

Confirming Alzheimer’s disease (AD) typically relies on positron emission tomography (PET), which remains costly and invasive, motivating the use of structural MRI-based prescreening. The proposed CSV-ViT framework addresses challenges in deep learning on non-Euclidean manifolds by introducing a novel cortical surface tokenization that performs ROI-preserving, vertexbased, variable-sized patch partitioning, demonstrating consistent improvements over recent surface-based models for classifying AD diagnosis and pathology status.

The gist

We propose an ROI-preserving cortical partition that avoids ROI mixing while ensuring non-overlapping vertices, and introduce CSV-ViT, a mask-aware embedding that ingests variable-sized CSV patches.

Partitioning for CSV-ViT

The partitioning process involves several stages to generate the Cortical Supervertices (CSVs). The procedure begins by registering FreeSurfer cortical surfaces to fsaverage6 and then partitioning each hemisphere into ROI-constrained cortical supervertices (CSVs) using the Desikan–Killiany atlas. To mitigate ROI mixing caused by atlas artifacts, "minor disconnected fragments (<10% of an ROI) are reassigned to the adjacent ROI with the largest boundary contact; non-cortical labels (e.g., the medial wall) are excluded." The mesh is modeled as a graph G = (V, E) with 1-ring adjacency to construct connected CSVs within each ROI under a global size bound L ≤ SV ≤ H.

Global Planning and Seed Initialization

The global planning stage searches for feasible bounds (L, H) and ROI-wise counts where P r Kr = Ktotal and constraints on the number of vertices are met. Among feasible allocations, the selection process aims to reduce ROI-wise imbalance by minimizing P r (nr/Kr −s¯) 2, where s¯ is the average vertex count per ROI. Seed initialization handles multiple connected components within an ROI by distributing Kr among them, initialized via farthestpoint sampling (FPS) on vertex direction vectors.

Supervertex Growing and Balancing

The supervertex growing stage involves defining near-equal integer quotas for each ROI component and grow SVs by absorbing unassigned 1-ring boundary vertices, which preserves connectivity. If an SV fails to meet its quota, seeds are updated based on the current assignment. The final step is supervertex balancing, which enforces the size bound L ≤ SV ≤ H by transferring boundary vertices between adjacent SVs on an SV adjacency graph while preserving connectivity. This results in 642 ROI-preserving CSVs per hemisphere (cortex-only, lossless, non-overlapping).

CSV-ViT Architecture and Embedding

CSV-ViT represents a subject by variable-sized CSVs. Given per-vertex cortical features on the ico6 mesh with C channels and a precomputed CSV map, the input is an tensor X ∈ R B×C×N×Vmax, where N is the number of CSVs (both hemispheres) and Vmax=69. A binary mask M ∈ 0, 1 N×Vmax is created from the CSV map to zero-out padded entries before embedding. The model uses a mask-aware patch embedding where vertices within each CSV are flattened, and a single linear layer is applied to obtain a single token per CSV: ti = Linearvec(xi), i = 1,..., N (Equation 2).

Classification Performance

The framework was evaluated on three binary classifications: diagnosis (Dx), amyloid positivity (Aβ), and tau positivity, using cortical thickness (CT) and curvature (Curv) as inputs. CSV-ViT achieved higher classification performance than recent surface-based models across all tasks. For instance, in the Dx task, CSV-ViT achieved an AUROC of 0.852±0.014 for CT and 0.764±0.025 for Curv, outperforming SiT (AUROC 0.809±0.028) and DiffusionNet (AUROC 0.821±0.047). The incremental ablation study confirmed that all three components—ROI-preserving tokenization, vertex-based partitioning, and variable-sized patches—contribute positively to performance, with the largest drop observed when removing ROI preservation for the harder A and T tasks.

Conclusion

We proposed an ROI-preserving cortical surface partitioning method that generates variable-sized CSV patches and designed CSV-ViT, a Vision Transformer that can ingest such variable-sized patches via padding and a mask-aware embedding. Our patching produces ROI-preserving and non-overlapping CSVs, enabling inter- and intra-regional attention across cortical regions. This approach supports MRI-based prediction of AD status prior to PET or CSF confirmation.

Improvements for AI systems

Here are specific improvements to existing AI systems based on the proposed CSV-ViT framework:

  1. The improvement is a novel method for feature extraction from structural MRI data, specifically by introducing an ROI-preserving cortical surface tokenization technique that generates variable-sized Cortical Supervertices (CSVs).

  2. The improved AI system, CSV-ViT, can perform accurate binary classification of Alzheimer’s Disease (AD) related pathologies—specifically diagnosis (Dx), amyloid positivity (Aβ), and tau positivity—using only T1-weighted MRI data.

  3. It can function as a prescreening tool by predicting AD-related status prior to costly and invasive PET or CSF confirmation, offering a non-invasive alternative for clinical triage.

  4. The system is specifically designed to overcome the limitations of traditional surface models (like SiT or MS-SiT) by ensuring ROI preservation, preventing vertex duplication at patch boundaries, and excluding non-cortical regions (like the medial wall).

  5. It utilizes a variable-size patch tokenization strategy enabled by padding and a mask-aware embedding mechanism within the Vision Transformer architecture, allowing it to effectively model cortical features across regions of varying geometry.

  6. The improved system can be validated through targeted ablation studies, confirming that components like ROI preservation, vertex-based partitioning, and variable-sized patches are all critical for superior performance across the three pathology classifications (Dx, Aβ+, Tau+).

Sources

Related papers