Draw This First

arXiv:2608.12064 · cs.CV, cs.LG · Submitted 2026-08-12 · Read on arXiv

Dazhi Zhong, Rowan Bradbury, Grant Davis

Krea.ai · Wand Technologies · Bradbury Group

cs.CV, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/Bradbury-Group/bbml

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The paper "Draw This First" presents a method for generating ordered vector sketches from text descriptions or images, where the drawing order can be specified via text instructions.

Terminology

Summary

The paper Draw This First presents a method for generating ordered vector sketches from text descriptions or images, where the drawing order can be specified via text instructions. The core idea is to invert the typical sketch generation formulation: instead of predicting strokes in sequence, the model predicts a 2D field that defines the order in which strokes are drawn.

Method Overview:

The approach uses a pretrained latent flow-matching transformer (Qwen-Image-Edit-2509) as the image prior. The model predicts an intermediate representation, while a trained VAE decoder predicts an order field, stroke mask, and stroke segmentation. The predicted segmentation is vectorized into polylines and sorted by the field, producing an ordered vector sketch.

Key Components:

  1. Dataset: A commissioned dataset of 47,318 handmade artist drawings, averaging 77.5 strokes and 6,864 points per drawing. Each drawing is annotated with a two-level bounding-box hierarchy (region then subject).

  2. Order-as-color encoding: The drawing order is encoded into an HSV color space. Global arc length (drawing progress) is encoded in the Hue channel (H = a · 342/360), and within-stroke arc is encoded in the Value channel (V = 1 − u/2). This creates an intermediate image representation that the diffusion model can generate within its trained image latent space.

  3. Order-native decoder: The VAE decoder is finetuned with a new pyramid head that emits 10 channels: global arc field, foreground mask, and an 8-dimensional instance embedding. The decoder is trained with a stroke-weighted L1 loss, perceptual loss, and a push-pull loss for instance embeddings.

  4. Permutation training: The model is trained with permuted stroke orders, with captions programmatically generated to match the permuted order. This teaches the model to follow text instructions specifying arbitrary order. Within-region order is preserved while varying between-region order.

  5. Vectorization: HDBSCAN clusters foreground embeddings to segmentations. Clusters are sorted by mean predicted arc, and a nearest-neighbour walk traces the foreground pixels into vector strokes, followed by smoothing and Ramer-Douglas-Peucker simplification.

Results:

  • Decoder ceiling: The order-native head cuts arc L1 error by 9–19× relative to reading hue from the frozen decoder, with arc Spearman ≥ 0.994 and mask IoU ≥ 0.97 on all five datasets. The deployment ceiling reaches matched Kendall τ of 0.91–0.94 on multi-stroke datasets.

  • Order control via language: Deleting the caption drops in-domain order to the level of geometry-only baselines (τ 0.096 vs. 0.105). Stating the recorded order lifts out-of-domain adherence (τ 0.373), and reversing the instruction inverts it (τ −0.267). The caption, not the condition image, is the order channel.

  • Image prior retention: The model retains open-vocabulary drawing ability. On 50 QuickDraw categories never seen as vectors, CLIP ViT-B/32 recognition achieves Top-1 0.70 (guidance 5.0) and 0.75 (7.5), Top-5 0.84.

  • Precision of order control: Region-level instructions are followed reliably (τ 0.838), but adherence falls to 0.46 at part level and 0.44 at pass level. On the external ControlSketch-Part dataset, τ reaches 0.78 as stated and 0.73 reversed, from an uninstructed baseline of 0.36.

Limitations:

  • Without instruction, the model has no verifiable human-like global order prior.

  • The order of objects in the caption text influences draw order, and generalization to instructed order conflicting with order of appearance is reduced.

  • Below named units there is no control: within-unit order agreement is near zero under every instruction.

  • Recovered paths run 1.7–2.5× the true stroke count, causing fragmentation.

  • Recorded time, pressure, tilt, and brush width are dropped; predicted paths are 1 px unweighted centerlines parameterized by cumulative arc length rather than time.

Conclusion:

The system generates ordered vector sketches by rendering draw order into color, generating that image with a pretrained diffusion transformer, and reading order back out with an order-native decoder into replayable polyline paths. Text is the order channel: captions carry most in-domain order, recorded-order instructions improve it, and reversed instructions invert it, all without changing geometry.

Improvements for AI systems

Improvements to AI systems:

  1. Unified representation for sequential generation tasks: Use the order-as-color encoding (HSV hue for global progress, value for within-stroke progress) to convert any sequential output (e.g., handwriting, route planning, assembly steps) into a single 2D image that a pretrained diffusion model can generate. This decouples what to draw from in what order, enabling text-controlled ordering without retraining the base generative model.

  2. Text-conditioned ordering in multimodal systems: Integrate the permutation-training scheme (programmatic captions matching permuted orders) to teach any vision-language model to follow explicit ordering instructions (e.g., draw the eyes first, then the nose) while preserving geometric fidelity. This improves controllability in image editing, diagram generation, and instructional content creation.

  3. Order-native decoder for fine-grained temporal supervision: Replace standard decoders with a pyramid head emitting global arc field + instance embeddings + mask, trained with stroke-weighted L1 and push-pull losses. This yields sub-stroke-level temporal precision (Spearman ≥0.994) and can be reused for any task requiring pixel-level sequencing, such as video frame prediction, robotic trajectory planning, or surgical step segmentation.

  4. Fragmentation-aware vectorization pipeline: Adopt the HDBSCAN + nearest-neighbor walk + RDP simplification to convert dense segmentation maps into clean polylines. Improve it by adding a stroke-merging post-processor that uses arc-field continuity to merge fragmented segments (reducing the 1.7–2.5× stroke count inflation), enabling more compact and human-readable vector outputs for CAD, animation, and SVG generation.

  5. Language-driven order inversion for interactive systems: Use the finding that captions are the sole order channel (not the condition image) to build an interactive drawing assistant where users can say reverse the drawing order or draw background last and see the stroke sequence invert in real time, without re-generating geometry. This enables rapid prototyping of animations, stroke-based tutorials, and accessibility tools for visual learners.

  6. Hierarchical order control with explicit fallback: Leverage the two-level bounding-box hierarchy (region → subject) to implement a system that accepts instructions at multiple granularities (region, part, pass). When finer-level instructions fail (τ drops to 0.44), the system can automatically fall back to the coarser level and inform the user of the achievable control, improving reliability in human-AI collaborative sketching tools.

  7. Open-vocabulary sequential generation: Combine the image prior retention (CLIP Top-1 0.75 on unseen categories) with the order field to enable zero-shot generation of ordered sketches for novel object classes—e.g., a user types draw a teapot, handle last and the system produces a plausible ordered vector sketch even if no teapot vectors were in the training set. This extends to other domains like floorplan design or molecular structure drawing.

  8. Calibrated confidence for order instructions: Use the measured adherence gap (region τ=0.838, part τ=0.46, pass τ=0.44) to build a confidence estimator that predicts whether a given instruction will be followed. The system can then either rephrase the instruction, request a coarser order, or flag uncertainty—improving trust in autonomous drawing and design tools.

  9. Time-parameterized output for animation: Replace the dropped temporal metadata (time, pressure, tilt) with the predicted cumulative arc length as a pseudo-time parameter. This allows the system to output stroke sequences that can be replayed at variable speeds, enabling automatic generation of drawing process videos from a single text prompt or image—useful for educational content and explainer videos.

  10. Cross-domain transfer of ordering priors: The decoder’s order field generalizes across five datasets (including external ControlSketch-Part). This suggests the same architecture can be fine-tuned for ordering in non-sketch domains—e.g., predicting the order of surgical instrument usage from a static OR image, or the assembly sequence of furniture from a product photo—by re-encoding the target sequential process into the HSV order space.

Abstract

We invert the typical formulation of sketch generation: instead of drawing strokes in order, we predict a 2D field that defines the order in which strokes are drawn. We use a pretrained latent flow-matching transformer to supply the image prior to predict an intermediate representation, while training the VAE's decoder to predict the order field, stroke mask, and stroke segmentation. We vectorize the predicted segmentation into polylines and sort them by the field, producing an ordered vector sketch. Our model can predict an ordered vector sketch from a text description or derender an image into ordered vectors; for either, it follows text instructions specifying the order of drawing.

Sources

Related papers