Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng, Jin Jiang, Junsheng Zhang, Wenhao Wang
Vast Intelligence Lab · University of Technology Sydney
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 24 pages, 10 figures
Code: https://github.com/Spark-To-Paper-Skills/spark-to-paper-skills
Project page: https://victorchen96.github.io/auto_research/auto_research_survey.pdf
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: Spark-to-Paper is an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or
Terminology
Summary
Spark-to-Paper is an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. The system separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode called the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs 8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.
Improvements for AI systems
Improvements to AI systems:
-
Separate judgment from execution – AI systems should explicitly split tasks into model-based reasoning (e.g.,
what claim is supported?
) and deterministic operations (e.g.,run this code, check this output
). This reduces hallucination and allows verifiable steps to be executed and validated without model drift. -
Pre-register evidence before claims – AI systems should require a plan of required evidence (experiments, metrics, citations) before generating results or claims. This prevents post-hoc rationalization and forces the system to abandon or revise claims if evidence fails.
-
Add a Self-Refutation Loop guard – AI systems should detect when repeated attempts to validate a hypothesis keep failing, and automatically trigger a re-framing of the research objective or a stop condition. This prevents wasted compute and infinite loops of negative results.
-
Integrate deterministic integrity checks with self-critique – AI systems should combine hard checks (e.g., citation existence, figure file validity, code execution success) with model-based self-review. This dual-layer catches both factual errors and logical inconsistencies.
-
Generate editable, programmatic artifacts – AI systems should output vector figures and diagrams as code (e.g., matplotlib, TikZ) rather than static images. This enables post-hoc human or automated editing, improving reproducibility and trust.
-
Use adversarial review as a separate validation stage – AI systems should include a dedicated adversarial reviewer that tries to falsify claims, not just a self-critic. This improves precision of accepted claims (74% in the paper) and reduces overconfidence.
-
Bound token/cost/time budgets explicitly – AI systems should track and cap resource usage per task (e.g., 11.9M tokens, 8.1, 3.2 hours) to make long-horizon generation feasible and predictable, and to trigger early termination if budgets are exceeded.
What the improved AI system can do:
-
Generate a full research manuscript from a topic with 99.5% citation validity and 96.4% editable figures.
-
Detect fabrication in drafts with 92% accuracy (vs. 14% for single-pass), by using integrity checks and self-critique.
-
Automatically revise or abandon claims when experimental evidence contradicts the original objective, avoiding endless failed iterations.
-
Produce code-based figures that humans can edit directly, enabling collaborative refinement.
-
Operate entirely within a coding assistant (e.g., VS Code, Jupyter) without needing a separate agent platform, making it lightweight and composable.
-
Run end-to-end research generation in 3 hours at 8 per manuscript, making it practical for large-scale hypothesis testing or literature exploration.
Sources
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Robin: A multi-agent system for automating scientific discovery
- Survey of Hallucination in Natural Language Generation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Self-Refine: Iterative Refinement with Self-Feedback
- Kosmos: An AI Scientist for Autonomous Discovery
- WebGPT: Browser-assisted question-answering with human feedback
- Generative Agents: Interactive Simulacra of Human Behavior
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Agent Laboratory: Using LLM Agents as Research Assistants
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Retrieval Augmentation Reduces Hallucination in Conversation
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- CycleResearcher: Improving Automated Research via Automated Review
- Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering