P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization

arXiv:2608.07549 · cs.CV, cs.AI · Submitted 2026-08-01 · Read on arXiv

Zhenhong Sun, Xibin Song, Haozhe Liu, Yifu Wang, Senbo Wang, Huadong Mo, Daoyi Dong, Hongdong Li, Pan Ji

Australian National University · Vertex Lab · University of New South Wales · University of Technology Sydney

cs.CV, cs.AI

Submitted: 2026-08-01

Updated: 2026-08-11

Code: https://github.com/ashawkey/cubvh

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: "field-centric volumetric sampling" (e.g., grid SDFs) and "edge-intersection surface sampling" (e.g., Dual Contouring).

Terminology

Summary

Summary

The paper introduces P2Voxel, a framework for 3D mesh tokenization that converts triangle meshes into compact, structured, and learnable tokens. The authors frame mesh tokenization as a geometric sampling problem and propose a new paradigm called local surface evidence sampling, which identifies the minimal geometric evidence inside each active voxel that is sufficient for deterministic surface recovery. This contrasts with existing approaches: field-centric volumetric sampling (e.g., grid SDFs) and edge-intersection surface sampling (e.g., Dual Contouring). The paper states: "Beyond field-centric volumetric sampling and edge-intersection surface sampling, we try to formulate mesh tokenization as local surface evidence sampling: identifying the minimal geometric evidence inside each active voxel that is sufficient for deterministic surface recovery."

The framework is built on three key innovations, each grounded in an explicit assumption:

  1. Pivot Voxelization (based on the Local Planarity assumption): Each active voxel is represented by a compact pivot token p, s, where p in R cubed is a pivot point on the local surface patch and s in-1, +1 is a binary inside/outside orientation sign. The paper explains: "Under the Local Planarity assumption, Pivot Voxelization represents each active voxel with a surface pivot and an orientation sign, providing minimal local evidence that can induce the corner values required for deterministic reconstruction." Given the pivot token and voxel size h, the voxel center c is recovered by snapping p to its containing voxel, the oriented normal is computed as n = s c-p overc-p 2, and an ideal plane = x in R cubed (x-p) n = 0 induces signed distances at the eight voxel corners. These corner values are used by Sparse Marching Cubes (Sparse-MC) for deterministic reconstruction. To ensure consistency across neighboring voxels, the paper enforces a vertex-consistent scalar field by averaging predictions from incident active voxels.

  2. Pyramid Pivot Voxelization (based on the Spatial Complexity assumption): This exploits the spatial non-uniformity of real surfaces by allocating finer pivot tokens to geometrically complex regions while keeping smooth regions coarse and compact. The method partitions the normalized space into macro-blocks and assigns each block a complexity score S B(g) based on curvature magnitude, curvature variation, curvature gradient, normal deviation, and sample count. This score is discretized into a sampling map M: g r in R = r,, r, where R is a predefined set of sampling resolutions. Pivot voxelization is then performed at each resolution, and tokens are aggregated according to the sampling map. For reconstruction, Pyramid Resolution Lifting lifts all tokens to the maximum resolution r by subdividing each original voxel into finer voxels and evaluating their corner SDF samples, enabling unified Sparse-MC reconstruction.

  3. Pyramid VAE (based on the Block Reconstructability assumption): The paper argues that "mesh recovery does not require learning the entire high-resolution voxelized shape as one monolithic object, because each pyramid block contains sufficient local surface evidence for deterministic reconstruction within its spatial extent. The Pyramid VAE learns compact multi-resolution latent codes over locally reconstructable pivot blocks," avoiding the need to model a dense global field. The encoder uses hierarchical 3D CNN branches to aggregate block-wise geometric features from coarse to fine levels, mapping them into a shared pyramid latent representation z = z n n=1 N. The decoder mirrors this hierarchy to reconstruct pivot-token grids at multiple resolutions, guided by the sampling map M. The paper notes that the pyramid representation separates where to allocate resolution from what surface evidence to store.

The paper formulates the problem as: given a watertight mesh M, find a token set T and a deterministic decoder D such that = D(T) and T is minimized subject to about M.

Experiments are conducted on three test sets (400 shapes each) from ABO, Objaverse, and Wild datasets. The authors benchmark against SDF-based baselines (Vecset, Dora, Hunyuan3D-2.1) and Dual Contouring baselines (FaithC, TRELLIS 2), focusing strictly on sampling strategy and reconstruction quality. Metrics include Chamfer Distance (L1/L2), Earth Mover's Distance (EMD), and F-score (τ = 0.002).

Key quantitative results from Table 1 (at 512 resolution): On ABO, Pivot-512 achieves CDL1 of 2.1166, CDL2 of 0.0110, EMD of 3.3364, and F-score of 0.4889 with 880,134 samples, outperforming SDF-based baselines and rivaling FaithC (Dim 18) which uses 820,445 samples. Pyramid-1024 reduces sampling count to 385,151 with only a small accuracy reduction (CDL1 of 2.1583). On Objaverse, Pivot-512 achieves CDL1 of 2.0362, CDL2 of 0.0105, EMD of 3.2096, and F-score of 0.5311. On Wild, Pivot-512 achieves CDL1 of 1.5090, CDL2 of 0.0058, EMD of 2.5602, and F-score of 0.7370.

Table 2 evaluates the Pyramid VAE: Pivot-512 achieves the best or second-best results on most metrics across datasets, with CDL1 of 2.1741 on ABO, 2.1071 on Objaverse, and 1.6202 on Wild. The paper notes the VAE is trained only on the ABO training split with limited data, without using Objaverse or Wild for training, yet still demonstrates cross-dataset generalization potential.

Ablation studies show: (1) Increasing pivot sampling resolution from 64 to 1024 consistently improves quality (CDL1 drops from 4.8159 to 1.6463, F-score rises from 0.1937 to 0.6139) while reducing artifact rates (TopoC from 14.85% to 0.22%, HighCurv from 17.39% to 0.87%), with diminishing returns beyond 512. (2) Relaxing sampling-map cut ratios activates more fine voxels (116,992 to 207,885) and improves quality (CDL1 from 2.3154 to 2.2229, F-score from 0.4234 to 0.4547), with additional primitives selectively allocated to geometrically demanding regions.

The paper concludes: "By turning dense mesh geometry into compact, locally reconstructable pyramid pivot blocks, P2Voxel provides a concrete token space for learning-based 3D representation. Future work will further explore its use in latent modeling and diffusion-based shape generation."

Improvements for AI systems

Based on the P2Voxel paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace dense SDF grids or edge-state encodings with the Pivot Voxelization scheme—each active voxel stores only a 4D token (3D pivot point + 1D orientation sign), reducing per-voxel memory from 18+ dimensions (as in FaithC/TRELLIS 2) to just 4 dimensions.

What the improved system can do:

  • Generate 3D shapes with 2–6× fewer tokens than current SDF-based or dual-contouring methods while maintaining comparable or better reconstruction fidelity (CDL1: 2.1166 vs 2.1205 for FaithC on ABO).

  • Train autoregressive or diffusion transformers on 3D meshes with substantially lower memory footprint (e.g., 385k tokens vs 880k tokens for the same quality).

  • Support direct mesh generation without post-processing steps like Poisson reconstruction.

Improvement: Implement the block-wise complexity scoring function SB(g) = c̄(g) + λvar·σc(g) + λgrad·‖∇c̄(g)‖ + λn·η̄(g)·N(g) to dynamically allocate sampling resolutions (e.g., 128/256/512/1024) based on local curvature, normal variation, and sample density.

Improvement: Use the pivot token p, s to analytically compute corner SDF values dδ = (vδ − p)Tn for Sparse Marching Cubes, eliminating the need to store or learn dense distance fields.

Improvement: Train a Pyramid VAE that encodes/decodes local pivot blocks (e.g., 323) independently rather than the entire 5123 or 10243 global field, using hierarchical 3D CNNs with multi-stage training (weights 5:3:2 for resolutions 32/16/8).

Improvement: Replace brute-force uniform voxelization with the two-stage pipeline: (1) compute per-block complexity scores from surface samples, (2) quantile-bin scores into allocation maps, then perform pivot voxelization only in active blocks.

Improvement: Use the empirical artifact curves (TopoC, HighCurv, All ratios) to automatically select the minimum resolution that keeps artifact rates below a user-defined threshold (e.g., 1% at 512 resolution).


Summary of system capabilities enabled by these improvements:

  • A 3D generation system that produces high-fidelity meshes with 2–6× fewer tokens than current state-of-the-art, trainable on a single RTX 4090.

  • A real-time 3D reconstruction pipeline that adaptively samples geometry based on local complexity, achieving near-lossless quality with minimal memory.

  • A scalable 3D asset processing framework that handles 10243 resolution in seconds, suitable for large-scale dataset preparation and interactive editing.

Abstract

Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable tokens. Beyond field-centric volumetric sampling and edge-intersection surface sampling, we retarget mesh tokenization as local surface evidence sampling: identifying the minimal geometric evidence inside each active voxel that is sufficient for deterministic surface recovery. To this end, we introduce P2Voxel, a pyramid pivot voxelization framework for compact and reconstruction-aware mesh tokenization. P2Voxel is built on three key innovations. Under the Local Planarity assumption, Pivot Voxelization represents each active voxel with a surface pivot and an orientation sign, providing minimal local evidence that can induce the corner values required for deterministic reconstruction. Under the Spatial Complexity assumption, Pyramid Pivot Voxelization exploits the spatial non-uniformity of real surfaces by allocating finer pivot tokens to geometrically complex regions while keeping smooth regions coarse and compact. Under the Block Reconstructability assumption, a Pyramid VAE learns compact multi-resolution latent codes over locally reconstructable pivot blocks, avoiding the need to model the entire high-resolution voxelized shape as a dense global field. Together, these designs convert meshes into compact, structured, and learnable pyramid pivot tokens, enabling efficient mesh reconstruction for downstream 3D tasks.

Sources

Related papers