KaiNinja: Extending Native 3D Generators to the Part Level
cs.GR, cs.AI, cs.CV
Submitted: 2026-09-14
Updated: 2026-09-15
Comments: Project page: https://alaya-lab.github.io/KaiNinja Code: https://github.com/AlayaLab/KaiNinja
Code: https://github.com/AlayaLab/KaiNinja
Project page: https://alaya-lab.github.io/KaiNinja
License: http://creativecommons.org/licenses/by/4.0/
The gist: Native 3D generators turn one image into a single mesh.
Terminology
Abstract
Native 3D generators turn one image into a single mesh. TRELLIS.2 and its peers deliver high-fidelity non-watertight geometry with materials, but the output is one fused object, while downstream work such as editing, rigging and simulation operates on part-level assets. A naive idea is to run a 3D segmentation network on the fused mesh that TRELLIS.2 generates, but such pipelines are slow and bounded by the accuracy of the segmentation. We want a simple way to extend an existing native 3D generator to the part level. But we face a critical problem: the O-Voxel grid stores one sheet of surface per voxel, so a single volume cannot represent the interface where two parts touch, at any resolution. We introduce a dual-volume representation to solve this problem and put forward KaiNinja, a part-level extension of TRELLIS.2 built on a dual-volume form of its O-Voxel representation. KaiNinja keeps the generation speed and quality of TRELLIS.2 while extending it to the part level, with no mask or segmenter in the pipeline. Its training data come from sources of many kinds, including CAD models and assets authored by an LLM-driven agent; to our knowledge it is the first 3D generative model trained on agent-authored part data. Surprisingly, we also find that whole-object fidelity improves over the same backbone fine-tuned on the same dataset. Against part generation pipelines of different paradigms, it lowers whole-object Chamfer distance by 40% and raises strict part F-score by 16%.
Sources
- AutoPartGen: Autogressive 3D Part Generation and Discovery
- Classifier-Free Diffusion Guidance
- CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner
- TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
- Infinite Mobility: Scalable High-Fidelity Synthesis of Articulated Objects via Procedural Generation
- PAct: Part-Decomposed Single-View Articulated Object Generation
- LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow
- P3-SAM: Native 3D Part Segmentation
- Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
- DINOv3
- Segment Any Mesh
- Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- Native and Compact Structured Latents for 3D Generation
- InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
- Meshy T2: Fast Native Mesh Generation with Flow Matching
- PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image
- X-Part: high fidelity and structure coherent shape decomposition
- HoloPart: Generative 3D Part Amodal Segmentation
- SAMPart3D: Segment Any Part in 3D Objects
Related papers
- SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
- CADReasoner: Iterative Program Editing for CAD Reverse Engineering
- QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
- DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches
- MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles
- MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering