Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

arXiv:2608.11924 · cs.CL · Submitted 2026-08-12 · Read on arXiv

Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng, Jin Jiang, Junsheng Zhang, Wenhao Wang

Vast Intelligence Lab · University of Technology Sydney

cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 24 pages, 10 figures

Code: https://github.com/Spark-To-Paper-Skills/spark-to-paper-skills

Project page: https://victorchen96.github.io/auto_research/auto_research_survey.pdf

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: Spark-to-Paper is an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or

Terminology

Summary

Spark-to-Paper is an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service. The system separates model-based judgment from deterministic operations that can be directly executed and checked. It further separates experiment planning from reporting, so that required evidence is specified before results are observed and manuscript claims are revised according to measured outcomes. To improve reliability over long research trajectories, the system combines deterministic integrity checks with self-critique and bounds a failure mode called the Self-Refutation Loop, in which repeated experiments continue to reject the original research objective. Spark-to-Paper also produces editable vector figures through programmatic plotting for experimental results and code-based reconstruction for generated method diagrams. Across eight controlled research topics, Spark-to-Paper achieves 99.5% citation validity and 96.4% figure editability. A controlled ablation increases fabrication detection from 14% for a single-pass draft to 92% with the full integrity and review stack, while adversarial review achieves 74% precision. The full system uses 11.9M tokens, costs 8.1 per manuscript, and requires 3.2 hours on average. These results show that end-to-end research paper generation can be implemented as a lightweight, composable workflow inside existing coding assistants while keeping experimental evidence central to how claims are accepted, revised, or abandoned.

Improvements for AI systems

Improvements to AI systems:

  1. Separate judgment from execution – AI systems should explicitly split tasks into model-based reasoning (e.g., what claim is supported?) and deterministic operations (e.g., run this code, check this output). This reduces hallucination and allows verifiable steps to be executed and validated without model drift.

  2. Pre-register evidence before claims – AI systems should require a plan of required evidence (experiments, metrics, citations) before generating results or claims. This prevents post-hoc rationalization and forces the system to abandon or revise claims if evidence fails.

  3. Add a Self-Refutation Loop guard – AI systems should detect when repeated attempts to validate a hypothesis keep failing, and automatically trigger a re-framing of the research objective or a stop condition. This prevents wasted compute and infinite loops of negative results.

  4. Integrate deterministic integrity checks with self-critique – AI systems should combine hard checks (e.g., citation existence, figure file validity, code execution success) with model-based self-review. This dual-layer catches both factual errors and logical inconsistencies.

  5. Generate editable, programmatic artifacts – AI systems should output vector figures and diagrams as code (e.g., matplotlib, TikZ) rather than static images. This enables post-hoc human or automated editing, improving reproducibility and trust.

  6. Use adversarial review as a separate validation stage – AI systems should include a dedicated adversarial reviewer that tries to falsify claims, not just a self-critic. This improves precision of accepted claims (74% in the paper) and reduces overconfidence.

  7. Bound token/cost/time budgets explicitly – AI systems should track and cap resource usage per task (e.g., 11.9M tokens, 8.1, 3.2 hours) to make long-horizon generation feasible and predictable, and to trigger early termination if budgets are exceeded.

What the improved AI system can do:

  • Generate a full research manuscript from a topic with 99.5% citation validity and 96.4% editable figures.

  • Detect fabrication in drafts with 92% accuracy (vs. 14% for single-pass), by using integrity checks and self-critique.

  • Automatically revise or abandon claims when experimental evidence contradicts the original objective, avoiding endless failed iterations.

  • Produce code-based figures that humans can edit directly, enabling collaborative refinement.

  • Operate entirely within a coding assistant (e.g., VS Code, Jupyter) without needing a separate agent platform, making it lightweight and composable.

  • Run end-to-end research generation in 3 hours at 8 per manuscript, making it practical for large-scale hypothesis testing or literature exploration.

Sources

Related papers