PACE: Precise AI Cinematic Expression
cs.CV, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-18
Comments: 36 pages, 7 figures, 3 tables. Code: https://github.com/StudioPiLabs/pace-core
Code: https://github.com/StudioPiLabs/pace-core
License: http://creativecommons.org/licenses/by/4.0/
The gist: Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands.
Terminology
Abstract
Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director's words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: https://github.com/StudioPiLabs/pace-core
Sources
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
- Teaching language models to support answers with verified quotes
- Structured Prediction as Translation between Augmented Natural Languages
- A Unifying Scheme for Extractive Content Selection Tasks
- Reasoning about Actions and State Changes by Injecting Commonsense Knowledge
- Tiered Reasoning for Intuitive Physics: Toward Verifiable Commonsense Language Understanding
- STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
- From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
- STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- Automated Movie Generation via Multi-Agent CoT Planning
- FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
- Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- LAMP: Language-Assisted Motion Planning for Controllable Video Generation
- Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models