JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
cs.CV, cs.AI
Submitted: 2026-06-25
Updated: 2026-09-27
Code: https://github.com/black-forest-labs/flux
Terminology
Sources
- Text-to-image Diffusion Models in Generative AI: A Survey
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Scalable Diffusion Models with Transformers
- CogView: Mastering Text-to-Image Generation via Transformers
- Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
- Long-CLIP: Unlocking the Long-Text Capability of CLIP
- Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
- LoRA: Low-Rank Adaptation of Large Language Models
- Q-Diffusion: Quantizing Diffusion Models
- Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- LCM-LoRA: A Universal Stable-Diffusion Acceleration Module
- Bose condensation of upper-branch exciton-polaritons in a transferrable microcavity
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models