Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
cs.SE, cs.AI
Submitted: 2025-08-06
Updated: 2026-09-07
Comments: The paper has been accepted for publication in ACM Transactions on Software Engineering and Methodology (TOSEM)
License: http://creativecommons.org/licenses/by/4.0/
The gist: The application of Large Language Models (LLMs) is growing in the productive completion of Software Engineering tasks.
Terminology
Abstract
The application of Large Language Models (LLMs) is growing in the productive completion of Software Engineering tasks. Yet, studies investigating productive prompting techniques often employed a limited problem space, focusing primarily on well-known prompting patterns and targeting function-level SE practices. We identify significant gaps in real-world workflows that involve complexities beyond class-level (e.g., multi-class dependencies) and different features that can impact Human-LLM Interaction (HLI) processes in code generation. To address these issues, we designed an experiment to comprehensively analyze HLI features related to code generation productivity. Our study presents two project-level benchmark tasks that extend beyond function-level evaluations. We conducted a user study with 36 participants from diverse backgrounds, asking them to solve the assigned tasks by interacting with the GPT assistant using specific prompting patterns. We also examined participants' experiences and behavioral features during interactions by analyzing screen recordings and GPT chat logs. Our empirical investigation, based on statistical guidance, revealed (1) that three out of 15 HLI features emerged as consistently supported factors for productivity; (2) five primary guidelines for enhancing productivity for HLI processes; and (3) a taxonomy of 29 runtime and logic errors that can occur during HLI processes, along with suggested mitigation plans.
Sources
- Evaluating Large Language Models Trained on Code
- ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
- Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
- Is ChatGPT the Ultimate Programming Assistant -- How far is it?
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties