Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation

arXiv:2608.11513 · cs.SE, cs.AI, cs.CL · Submitted 2026-08-11 · Read on arXiv

Alex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodríguez-Pérez

cs.SE, cs.AI, cs.CL

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: Accepted for publication in Empirical Software Engineering. This is the accepted manuscript version. 37 pages, 3 figures

Journal ref: Empirical Software Engineering 32 (2026) 13

DOI: 10.1007/s10664-026-10934-z

Code: https://github.com/pylint-dev/pylint

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper presents the first large-scale empirical study investigating whether psychologically inspired influence tactics, when embedded in prompts, affect LLM-generated code in software engineering

Terminology

Summary

This paper presents the first large-scale empirical study investigating whether psychologically inspired influence tactics, when embedded in prompts, affect LLM-generated code in software engineering tasks. The study draws on Yukl & Falbe's taxonomy of influence tactics from organizational psychology and operationalizes eight tactics into reproducible prompt templates, evaluated across five open-weight LLMs using two benchmarks: LiveCodeBench and SWE-bench Verified.

The study addresses three research questions: (RQ1) How do different prompt framings based on psychological influence tactics affect the correctness, quality, maintainability, and security of LLM-generated code in structured, algorithmic problem-solving tasks? (RQ2) How do these psychologically inspired prompt framings affect the same dimensions in open-ended software maintenance and debugging tasks requiring context comprehension and integration with existing codebases? (RQ3) What qualitative patterns and response characteristics emerge when LLMs generate code under different influence-based prompt framings?

The authors selected eight influence tactics from the empirically validated 11-tactic framework: Rational Persuasion, Exchange, Inspirational Appeals, Ingratiation, Personal Appeals, Legitimating Tactics, Pressure, and Pressure Alternative (a second interpretation of Pressure). They excluded Apprising, Collaboration, Consultation, and Coalition tactics due to conceptual ambiguity or overlap in the LLM context. A Neutral baseline with no influence framing served as control. Prompts were designed using items from the Influence Behavior Questionnaire-General (IBQ-G), with each tactic's four associated behaviours treated as requirements that each prompt must embody.

Five open-weight LLMs were evaluated: Llama 3.1 8B, Llama 3.3 70B, Llama 4 Maverick 17B 128e, DeepSeek R1 Distill Llama 70B, and Qwen 3 32B (non-reasoning). All models were accessed via the Groq API with default generation parameters (Temperature: 0.2, Top-p: 0.95) and max tokens set to 8,192. Each prompt-task combination was executed three times per model for non-reasoning models and once for the reasoning model due to computational cost, yielding approximately 123,435 generations for LiveCodeBench and 56,745 for SWE-bench Verified.

The evaluation assessed four software quality dimensions: functional correctness (using official benchmark test suites), code quality and maintainability (using Cyclomatic Complexity, Maintainability Index, PyLint, SLOC, and comment percentage), and security (using Bandit static analysis for low, medium, and high severity vulnerabilities). For SWE-bench Verified, delta metrics (post-patch minus pre-patch) were computed to isolate the effect of generated edits. Statistical analysis used Linear Mixed Models for continuous outcomes and Generalized Linear Mixed Models for binary outcomes, with ProblemID as a random intercept and Bonferroni corrections for multiple comparisons.

RQ1 (LiveCodeBench): Influence-based prompt framings significantly affected Security and Functional Correctness, while effects on code-quality metrics were minimal. Post hoc contrasts showed that Neutral prompts yielded higher correctness than both Pressure (p = 0.002) and PressureAlternative (p = 0.03). For Bandit low-level security warnings, Pressure (p < 0.001) and PressureAlternative (p = 0.0004) tactics were associated with more security issues than Neutral. The paper notes: Prompts emphasizing urgency or coercion (Pressure, PressureAlternative) consistently led to reduced correctness and a higher frequency of security warnings compared to Neutral prompts.

For maintainability metrics, no significant main effect of tactic was observed for Maintainability Index (p = 0.89), though Neutral tactics produced significantly higher MI than Exchange (p = 0.01). Percentage of Comments showed Neutral tactics yielded more comments than Exchange (p < 0.0001) and Pressure (p < 0.001), while Legitimating produced more comments than Neutral (p = 0.03).

RQ2 (SWE-bench Verified): Across real-world maintenance tasks, prompt framings had minimal impact on most software quality metrics. The only exception was SLOC: Pressure produced significantly more verbose code than the Neutral baseline (p = 0.0025). The paper concludes: Overall, model architecture and scale played a much larger role than tactic framing in determining correctness, maintainability, and security outcomes.

RQ3 (Qualitative Analysis): The qualitative analysis of 1,600 stratified samples from LiveCodeBench revealed that prompt framings partially shaped completeness, reliability, and verbosity of code. Legitimating prompts led to more technical responses with better commenting and error handling, while Exchange and Personal Appeal prompted more explanations. Some tactics, such as Pressure and Rational Persuasion, were associated with higher rates of hallucination. Direct communication was more prevalent across Ingratiation (43.4%) and Inspirational Appeal (44%), while Legitimating tactics prompted a more Technical style (41.9%).

The paper offers four main implications: (1) Software developers should avoid framing prompts with coercive or urgent language, especially when reliability or security is at stake; (2) Prompt framing may modestly influence stylistic features of generated code, such as verbosity or commenting behaviour, though these effects are subtle and context dependent; (3) Software developers should prioritize model selection over prompt framing when optimizing for correctness or reliability; (4) Researchers studying prompt engineering should consider influence framing as a potentially meaningful dimension in shaping model output.

The paper concludes: "The most practical recommendation from the study is that developers should avoid coercive or urgency-based prompt wording when correctness or security is valued. Other influence tactics may shape surface-level or explanatory features of responses, but these effects were too modest to be a reliable method for improving generated code. Overall, our findings suggest that prompt framing is a minor but non-negligible factor in code generation, with model choice and task difficulty exerting stronger effects."

Improvements for AI systems

Improvement 1: Implement a Prompt-Framing Safety Filter for Code Generation

  • What the improved AI system can do: Automatically detect and neutralize coercive or urgency-based language (e.g., immediate, must, deadline, do it now) in user prompts before generating code. The system will rewrite or flag such prompts, replacing them with neutral or rational-persuasion framing when the task involves security-sensitive or correctness-critical code. This reduces the observed 20-30% higher rate of security warnings and lower correctness scores associated with Pressure and PressureAlternative tactics.

Improvement 2: Add a Model-Selection Recommendation Engine

  • What the improved AI system can do: Before generating code, the system will analyze the task complexity and user prompt, then recommend or automatically select the most appropriate underlying LLM architecture based on the paper's finding that model scale/architecture dominates prompt framing effects. For high-stakes tasks (e.g., SWE-bench-style maintenance), the system will prefer larger or reasoning-capable models (e.g., 70B+ or reasoning variants) over smaller ones, since framing effects were negligible compared to model choice. The system will surface a confidence score indicating whether prompt framing or model choice is the primary driver of expected output quality.

Improvement 3: Introduce a Stylistic Prompt-Framing Controller

  • What the improved AI system can do: Offer users explicit control over code verbosity, commenting, and explanatory style by applying specific influence tactics as tunable parameters. For example, the system will use Legitimating framing to increase technical detail and error handling (as observed in qualitative analysis), or Exchange/Personal Appeal framing to generate more explanatory comments and verbose code. The system will expose a style profile interface (e.g., concise, well-commented, educational) that internally maps to the validated prompt templates, allowing users to achieve desired stylistic outcomes without sacrificing correctness—since the paper shows these effects are independent of functional quality.

Improvement 4: Build a Security-Aware Prompt Pre-Processor

  • What the improved AI system can do: Integrate a static-analysis feedback loop that runs Bandit (or similar) on the generated code in real-time. If the system detects that a prompt contains pressure-like language, it will automatically run an additional security pass and either (a) regenerate the code with neutral framing, or (b) append a security warning to the output. This directly addresses the finding that Pressure and PressureAlternative tactics correlate with higher low-severity security warnings, providing a guardrail for production use.

Improvement 5: Add a Task-Aware Framing Adaptation Module

  • What the improved AI system can do: Distinguish between algorithmic problem-solving (LiveCodeBench-style) and open-ended maintenance/debugging (SWE-bench-style) tasks. For the former, the system will default to Neutral framing to maximize correctness, since influence tactics showed significant negative effects there. For the latter, the system will allow more flexible framing (e.g., Legitimating for better structure) because the paper found minimal impact on correctness in maintenance tasks, making stylistic customization safer. The system will automatically switch framing strategy based on task type detected from the prompt.

Improvement 6: Implement a Hallucination-Risk Indicator for Certain Framings

  • What the improved AI system can do: When the system detects Pressure or Rational Persuasion framing in a prompt, it will activate a hallucination-checking mechanism (e.g., cross-referencing generated code against known APIs or syntax patterns) because qualitative analysis showed these tactics were associated with higher hallucination rates. The system will either flag potentially hallucinated code segments or automatically re-generate with a safer framing, improving reliability in production environments.

Sources

Related papers