PPTBench: Can Coding Agents Reconstruct the Visual World through Structured, Editable Slides
cs.CL, cs.AI
Submitted: 2026-08-31
Updated: 2026-09-28
Project page: https://lab.einsia.ai/pptbench
Terminology
Sources
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
- DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
- Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation
- AutoPresent: Designing Structured Visuals from Scratch
- PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion
- Paper2SysArch: Structure-Constrained System Architecture Generation from Scientific Papers
- SciFig: Towards Automating Editable Figure Generation for Scientific Papers
- COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
- DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
- Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering