Evaluating and Preventing Security Smells in AI-Generated Ansible Code
cs.SE, cs.AI, cs.CR
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/stackblitz-labs/bolt.diy
Project page: https://livecodebench.github.io
Terminology
Sources
- Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
- CodeLMSec Benchmark: Systematically Evaluating and Finding Security Vulnerabilities in Black-Box Code Language Models
- A Framework for Measuring the Quality of Infrastructure-as-Code Scripts
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties