HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
cs.CV, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://mathmagic-official.github.io/twinkle3d
Terminology
Sources
- PolyDiff: Generating 3D Polygonal Meshes with Diffusion Models
- Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
- PP-OCR: A Practical Ultra Lightweight OCR System
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation
- Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale
- MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
- UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement
- Shap-E: Generating Conditional 3D Implicit Functions
- Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
- Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
- TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
- Point-E: A System for Generating 3D Point Clouds from Complex Prompts
- SAM 3D: 3Dfy Anything in Images
- Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
- Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
- TripoSR: Fast 3D Object Reconstruction from a Single Image
- AssetGen: Deployable 3D Asset Generation at Interactive Speed
- iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
- LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models