Prompts in the Wild: A Large Analyzed Collection of Transactional Prompts in Code

arXiv:2608.12905 · cs.CL · Submitted 2026-08-13 · Read on arXiv

Victoria Basmov, Yoav Goldberg, Reut Tsarfaty

Bar-Ilan University · Allen Institute for Artificial Intelligence

cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Journal ref: Proc. of the 20th Linguistic Annotation Workshop (LAW XX), pp. 257-308, 2026

DOI: 10.18653/v1/2026.law-main.19

Code: https://github.com/OnlpLab/transactionalPromptsCollection

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: " Summary This paper argues that prompts, the instructions given to large language models (LLMs), should be treated as "first-class objects for empirical scientific and linguistic investigation"

Terminology

Summary

"

Summary

This paper argues that prompts, the instructions given to large language models (LLMs), should be treated as first-class objects for empirical scientific and linguistic investigation rather than informal artifacts. To facilitate this, the authors introduce a large-scale dataset, a structured ontology, an empirical analysis, and a user interface.

1. Contributions and Motivation

The paper's main contributions are:

(i) a large-scale dataset of 57.5K transactional prompts gathered from GitHub;

(ii) a structured prompt-ontology that captures the primary prompt features and components;

(iii) empirical analysis of the structured prompts, highlighting patterns in the way programmers use LLMs; and

(iv) a user interface for browsing and searching prompts by their properties and components to support further research.

The authors focus on transactional prompts, which are defined as reproducible, parameterized tasks, as part of a larger software-based workflow. This is contrasted with casual, one-off interactive prompts. The paper notes that transactional prompts are predominantly single-turn and are distinct from agentic usage, which involves iterative tool-use loops and is beyond the scope of this work.

2. Data Collection

The dataset was collected from public GitHub repositories by searching for files that invoke the chat.completion.create API or the PromptTemplate constructor from the LangChain package. The authors used static analysis to resolve variable assignments and function calls to extract the actual prompt text. After filtering and deduplication, they obtained 57,640 unique prompts (36,916 from chat.completion.create and 20,724 from PromptTemplate). The paper acknowledges a potential bias in this collection method, as it excludes projects using other APIs or programming languages.

3. The Prompt Ontology

To enable quantitative analysis, the authors define an ontology that captures the primary features of prompts. The ontology is grounded in inherent prompt properties, prior literature, and manual inspection. It includes the following dimensions:

  • Languages: Detected languages in the prompt text and explicit language mentions.

  • Task and domain: Coarse and fine-grained categories for the intended task and application domain.

  • Input characteristics: Identifies the high-level instructions, the question/task, and supporting context, and whether each is hard-coded or a variable input. It also captures the language, structure, and modality of input slots.

  • Output characteristics: Annotates the expected output's modality, structure, language, and answer paradigm.

  • Prompt structure: Represents the prompt as a sequence of role-messages (System, User, Assistant). Each message is broken down into a sequence of individual instructions, each associated with one of 42 semantic kinds (e.g., role specification, constraint or restriction). Instructions are marked as negative or central.

  • Prompting techniques: Extracts a list of 12 prompting techniques (e.g., role assignment, structured output, chain of thought) with supporting evidence spans.

  • Meta-data: Includes ID, GitHub URL, timestamp, and full prompt text.

4. Annotation and Quality Control

The annotation was performed using an LLM-based process with gpt-4.1. The authors conducted a manual error analysis on 100 randomly selected data points to assess annotation quality. The results show generally strong performance across most fields, with many categories exceeding 90% accuracy (e.g., Prompt Language 93.0%, Domain 90.6%, Prompting Techniques 98.5%). However, some fields were more challenging, particularly Output Type (60.4%) and Directions Text (69.4%). The main error categories were hallucinations, confusions between related categories, omissions, and segmentation errors, often occurring when indirect inference was required.

5. Empirical Analysis

The analysis of the structured prompts revealed several key findings:

  • Language and Modality: English is overwhelmingly dominant, accounting for 84.66% of identifiable cases. The dataset includes prompts in 62 languages. Only 6.3% of prompts are multilingual, and 8.19% are entirely non-English. Explicit language mentions cover a much broader range of 151 languages and dialects. Text is the dominant input (77.82%) and output (over 97%) modality, with images being the most common non-text modality.

  • Domain: 57.38% of prompts have an identifiable domain, resulting in 77 distinct domains with a Zipf-like distribution. The top domains are education & instruction, software development, and business & commerce.

  • Structure and Semantics: Prompts are broken down into an average of 6.85 instruction blocks. The most frequent instruction kinds include Input context placeholder (16%), Constraint or restriction (11%), and Output content requirements (9.7%). Only 18.2% of instructions are central to the task, while the rest are meta-instructions. Roughly 31% of prompts contain at least one negative instruction. Constraints represent 33.3% of all instruction blocks, with some prompts containing up to 155 distinct rules.

  • Messages: The majority of prompts (64%) have two messages, with system-user being the most common sequence (96% of those). This suggests developers use the System role as a static configuration layer and the User role as the dynamic input.

  • Tasks and Inputs: NLP tasks prevail, with question answering, general text generation, information extraction, and summarization covering over 48% of the data. 97.9% of prompts include a question or main task, and 89.2% include a supporting context. The question/task is expected as input 70% of the time, while context is provided as input 94% of the time. 89.25% of all cases are grounded in a provided context.

  • Prompting Techniques: Role assignment is the most frequently used technique, accounting for over 45% of all instances. Other common techniques include structured output (12.9%), decomposition via prompt (10.34%), and sections (7.83%).

6. User Interface and Related Work

The paper provides a web-based UI for exploring the dataset, allowing users to filter by ontology fields, search semantically, and inspect individual prompts. The related work section distinguishes this dataset from user-prompt datasets like LMSYS-Chat-1M and WildChat, and from other transactional prompt datasets like PromptSet, which the authors argue has less rigorous extraction and cleanup.

Improvements for AI systems

Improvements to AI Systems:

  1. Dynamic Prompt Structuring Engine – Implement an AI that automatically decomposes a user’s high-level goal into a multi-message prompt (System/User/Assistant) with explicit instruction blocks (role, constraints, output format, context placeholders), mirroring the paper’s ontology. The system can generate prompts with an average of 6.85 instruction blocks, prioritizing central instructions and minimizing negative phrasing.

  2. Constraint-Aware Generation – Build an AI that detects and enforces constraints (33.3% of instruction blocks in the dataset) by parsing user requirements into explicit, machine-checkable rules. The system can handle up to 155 distinct constraints per prompt, validate outputs against them, and flag violations in real-time.

  3. Multilingual and Multimodal Prompt Adaptation – Develop an AI that automatically translates and localizes prompts across 62 languages (with support for 151 dialects) while preserving instruction semantics. It can also convert text-only prompts into multimodal variants (e.g., adding image input slots) based on the task domain, since only 6.3% of prompts are currently multilingual and text dominates.

  4. Role-Based Configuration Layer – Create an AI that separates static system-level instructions (e.g., persona, tone, global rules) from dynamic user-level inputs, following the 96% system-user pattern. The system can cache and reuse the system role across multiple tasks, reducing token overhead and improving consistency.

  5. Prompt Technique Recommender – Implement an AI that analyzes a task description and recommends the most effective prompting techniques (role assignment, structured output, chain-of-thought, decomposition) based on the paper’s frequency data. For example, it would suggest role assignment for 45% of tasks, structured output for 12.9%, and decomposition for 10.34%, then generate the corresponding prompt.

  6. Grounded Context Injection – Build an AI that ensures 89.25% of generated prompts are grounded in provided context by automatically extracting relevant supporting information (from files, databases, or user history) and inserting it as context placeholders. The system can distinguish between hard-coded context and variable inputs, and dynamically fill slots at runtime.

  7. Negative Instruction Filter – Develop an AI that detects and reframes negative instructions (present in 31% of prompts) into positive, actionable directives. It can also identify when a negative instruction is essential (e.g., “do not include citations”) and preserve it, while converting ambiguous negatives into explicit constraints.

  8. Domain-Specific Prompt Templates – Create an AI that generates domain-aware prompts for 77 identified domains (education, software dev, business) by using fine-grained task categories. The system can auto-select the top domain (e.g., education & instruction) and tailor instruction kinds (e.g., “output content requirements” for summarization tasks) accordingly.

  9. Output Structure Predictor – Implement an AI that predicts the expected output modality, structure, and answer paradigm (e.g., JSON, markdown, code, table) from the prompt’s instruction blocks. It can then format its responses to match, improving parseability and reducing post-processing needs.

  10. Prompt Quality Auditor – Build an AI that evaluates existing prompts against the paper’s ontology, scoring them on completeness (e.g., missing context, vague constraints, lack of role specification). It can suggest edits to add missing instruction kinds, reduce ambiguity, and align with best practices from the empirical analysis (e.g., adding a system role, specifying output structure).

What the Improved AI System Can Do:

  • Automatically generate production-grade transactional prompts for software workflows, with explicit roles, constraints, and grounded context.

  • Adapt prompts across languages, modalities, and domains without manual rewriting.

  • Enforce complex rule sets and validate outputs against them in real-time.

  • Recommend and apply the most effective prompting techniques based on task type.

  • Audit and improve existing prompts to increase reliability and reduce errors.

  • Support both single-turn and multi-turn workflows by separating static configuration from dynamic inputs.

Abstract

The behavior of contemporary generative Large Language Models (LLMs) is directly shaped by prompts, unstructured texts that describe the desired output and model behavior. In this paper we argue that prompts are linguistic objects that merit investigation in their own right. To this end, we collect 57.5K unique samples of prompts from GitHub. Specifically, we focus on transactional prompts: reproducible natural language instructions that are integrated into software. To enable the empirical, quantitative study of prompts, we introduce a structured ontology, capturing the properties of prompts as well as their formal and semantic components. Based on this ontology, we transform prompts from unstructured raw texts into richly structured linguistic objects. Analysis of these structured data reveals significant diversity of usage patterns across languages, domains, tasks, and modalities, in a typical Zipf-like distribution where some clearly prevail and others, more diverse, appear in the long tail. To validate the reliability of the ontology-based annotation of the prompts, we perform a comprehensive error analysis across all fields, providing a detailed assessment of annotation quality. We release the dataset together with a browsing and exploration interface (https://github.com/OnlpLab/transactionalPromptsCollection).

Sources

Related papers