Engineering Systems for Data Analysis Using Interactive Structured Inductive Programming
cs.AI, cs.SE
Submitted: 2025-03-18
Updated: 2026-03-08
Comments: Accepted for publication in the 38th International Conference on Advanced Information Systems Engineering (CAiSE 2026)
Journal ref: Advanced Information Systems Engineering (CAiSE 2026), LNCS 16558, 249-266 (2026)
DOI: 10.1007/978-3-032-28110-4_14
Project page: https://shraddhasurana.github.io/dhaani
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Engineering information systems for scientific data analysis presents significant challenges: complex workflows requiring exploration of large solution spaces, close collaboration with domain
Terminology
Abstract
Engineering information systems for scientific data analysis presents significant challenges: complex workflows requiring exploration of large solution spaces, close collaboration with domain specialists, and the need for maintainable, interpretable implementations. Traditional manual development is time-consuming, while "No Code" approaches using large language models (LLMs) often produce unreliable systems. We present iProg, a tool implementing Interactive Structured Inductive Programming. iProg employs a variant of a '2-way Intelligibility' communication protocol to constrain collaborative system construction by a human and an LLM. Specifically, given a natural-language description of the overall data analysis task, iProg uses an LLM to first identify an appropriate decomposition of the problem into a declarative representation, expressed as a Data Flow Diagram (DFD). In a second phase, iProg then uses an LLM to generate code for each DFD process. In both stages, human feedback, mediated through the constructs provided by the communication protocol, is used to verify LLMs' outputs. We evaluate iProg extensively on two published scientific collaborations (astrophysics and biochemistry), demonstrating that it is possible to identify appropriate system decompositions and construct end-to-end information systems with better performance, higher code quality, and order-of-magnitude faster development compared to Low Code/No Code alternatives. The tool is available at: https://shraddhasurana.github.io/dhaani/
Sources
- Program Synthesis with Large Language Models
- A Model for Intelligible Interaction Between Agents That Predict and Explain
- Evaluating Large Language Models Trained on Code
- Large Language Models for Software Engineering: Survey and Open Problems
- Structured Prompting: Scaling In-Context Learning to 1,000 Examples
- Unlocking Structured Thinking in Language Models with Cognitive Prompting
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Towards a Grounded Dialog Model for Explainable Artificial Intelligence
- GPT-4 Technical Report
- Program Synthesis using Inductive Logic Programming for the Abstraction and Reasoning Corpus
- Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
- UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection