IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

arXiv:2608.10920 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst, Signe Riemer-Sørensen, Tobias Herb, Meeyoung Cha, Daniel Thilo Schroeder

King's College London · Max Planck Institute for Security and Privacy · BI Norwegian Business School · University of Oslo · SINTEF Digital · Korea Advanced Institute of Science and Technology (KAIST)

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/DISARMFoundation/DISARMframeworks

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: IO Factory is an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes.

Terminology

Summary

IO Factory is an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now extends beyond persuasive text from individual language models to AI swarms—persistent groups of coordinated agents that adapt to platform feedback and disguise organized campaigns as ordinary social interaction. Because such campaigns cannot be identified from isolated messages alone, they must be analyzed across a continuous spectrum of planning, platform action, exposure, interpretation, measurement, and adaptation. IO Factory represents this process inside a controlled simulated platform, linking actor roles, platform actions, exposure records, structured model-based evaluations, and configured changes in the simulated population. The architecture is implemented and evaluated across configurations of up to 100,000 agents. The results show that IO Factory executes campaign timelines at scale and produces inspectable evidence of exposure and measured movement in configured belief variables. By recording the actors, objectives, action constraints, exposure paths, and measurement rules used in each run, IO Factory supports reproducible research and red-team analysis of coordinated influence.

The framework is organized into three planes: a control plane that steers experiment execution, a simulation plane in which platform activity unfolds, and an evaluation plane that turns this activity into measurements and evidence. The control plane contains the run controller, which orchestrates execution across all other components, initializes the experiment from a configuration and seed, advances discrete simulation steps, schedules actor actions, evaluates transition conditions, advances the campaign lifecycle when they are met, and records checkpoints and run metadata. The simulation plane contains the environment, the actors, and the study models. The environment is the complete simulated platform holding public platform state, including the social graph, content, and discovery surfaces. Actors are stateful agents that observe the environment through a limited platform view and act on it through validated actions. Three actor categories are defined: civilians (ordinary simulated users forming the measured audience), IO operators (controlled intervention actors with public accounts and private campaign guidance), and the manager (a non-public coordination component issuing bounded directives). The study models govern this activity and define what is studied, comprising a campaign model for objectives and lifecycle phases, an action model for permitted behavior, and an exposure model determining which civilian encounters enter measurement. The evaluation plane receives what the simulation plane produces, with a measurement layer processing civilian encounters with visible content, recording exposure events, obtaining structured judge readings, and applying construct-update rules, while the evidence layer preserves links among action records, exposure records, state trajectories, metrics, and matched comparison artifacts.

The campaign lifecycle gives simulated campaign time an operational structure, using a ten-phase reference framework: reconnaissance, narrative design, infrastructure, content production, laundering, integration, amplification, absorption, adaptation, and evaluation. Each phase gates which actor categories can be selected, which actions they may propose, and which phase-specific records are available. Narrative formation translates configured construct definitions into intermediate campaign objects, but these cannot directly change civilian construct state—civilian movement can occur only after a later IO-operator action is validated, committed to the platform, made visible to a civilian, recorded as an exposure, interpreted by the judge, and passed through the construct-update rule. Actor execution follows an observe, select, validate, commit, and record cycle that is phase-gated, with manager directives entering as private guidance for eligible IO operators.

The measurement pipeline connects platform activity to possible movement in simulated civilian state. A construct definition is the study-specific description of what is being measured, such as trust in a source or support for a policy. An exposure event is the measurement record created when a civilian encounters a platform object through a configured visibility or discovery path. An LLM judge evaluates exposed content against the construct definition, returning structured readings for relevance, stance, confidence, and persuasiveness—these are simulator measurements, not ground truth. The simulator-scale movement is computed as Δi,c,e = pe τi,s(e) ui ri,e λc de,c γe,c, where pe is judged persuasiveness, τi,s(e) is civilian trust in the source, ui is a civilian-level update weight, ri,e is the repetition weight, λc is the construct-specific update rate, de,c is the direction, and γe,c is judge confidence. The civilian state is updated by clipping the sum of the old value and the movement to the configured scale. Repeated exposure uses a configurable decay term ri,e = max(ρmin, ρni,e). The comparison protocol uses matched baseline runs without IO operators or manager guidance, computing directional lift as Ls,c = qc[(x̄active s,c,T − x̄active s,c,0) − (x̄baseline s,c,T − x̄baseline s,c,0)], where qc aligns the result with the target direction.

The reported runs use Gemma 4 31B on H200 nodes, with 13 matched simulation replicates of 10,000 civilians per condition and 1,000 IO operators in active runs. Two main empirical designs are reported: a two-construct increase-target design targeting trust in Russia and support for eating insects, and a single-construct decrease-target design targeting lower trust in public institutions. All three primary endpoints move in the target direction. In the two-construct design, directional lift is 0.132 for support for eating insects and 0.130 for trust in Russia. In the single-construct design, trust in public institutions declines more than baseline, yielding sign-adjusted lift of 0.336. All three estimates are significant after Holm correction with p < 0.001. The two increase-target constructs had comparable processed-exposure counts (about 495k ± 71k) but different update rates (29.3% for support for eating insects and 33.5% for trust in Russia). The single-construct design had more processed exposures (597k ± 66k), a higher update rate of 90.4%, and a larger mean update size. Final active-condition agenda share was 0.039 ± 0.009 in the two-construct design and 0.426 ± 0.005 in the public-institutions design. The active reach fraction is about 0.494 in both main designs, meaning roughly half of the modeled civilians had an interaction record involving an IO-operator source.

The results should be interpreted within three boundaries. First, the reported runs use modeled civilians, configurable platform rules, a simulated exposure graph, and construct movement on a simulator scale—these results should not be read as estimates of real-world persuasion. Second, LLM-based measurement makes large-scale construct measurement feasible, but LLM judges are not ground truth; their outputs depend on construct definitions, prompts, model versions, serving settings, and context. Third, the reported results are model-dependent, as LLMs are used for narrative design, action selection, content generation, manager assessment, and judge measurement. Future work should develop IO Factory as shared research infrastructure, including benchmark scenarios, systematic variation in graph structure and recommender behavior, defensive scenarios as first-class interventions, and distinguishing human-facing exposure from machine-facing exposure. The longer-term goal is a shared representational vocabulary for AI-enabled influence research, allowing the community to state what was represented, what changed, who was measured, which records count as evidence, and where interpretation stops.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Campaign lifecycle enforcement for multi-agent coordination – Implement a phase-gated execution loop (reconnaissance → narrative design → infrastructure → content production → laundering → integration → amplification → absorption → adaptation → evaluation) where each agent’s action proposals are validated against the current phase, preventing premature or out-of-order influence attempts. The improved system can run coordinated influence campaigns with strict operational sequencing and traceable state transitions.

  2. Bounded directive issuance for hierarchical agent control – Add a non-public manager agent that issues bounded directives (e.g., target constructs, action constraints, phase-specific goals) to operator agents, while civilians remain unaware of the coordination. The improved system can simulate multi-level command structures where operators receive private guidance without exposing campaign intent to the measured population.

  3. Exposure-driven construct update with decay and trust weighting – Integrate the movement formula Δi,c,e = pe · τi,s(e) · ui · ri,e · λc · de,c · γe,c into agent belief systems, including source-trust modulation, repetition decay (ri,e = max(ρmin, ρni,e)), and judge-confidence scaling. The improved system can model how repeated exposure to persuasive content shifts agent beliefs non-linearly, with diminishing returns and source credibility effects.

  4. Matched baseline comparison for causal lift estimation – Implement the directional lift metric Ls,c = qc[(x̄active − x̄baseline)] using paired simulation runs with identical seeds and configurations, differing only in the presence of IO operators. The improved system can quantify the marginal effect of coordinated influence versus organic drift, enabling red-team analysis of campaign effectiveness.

  5. LLM judge as a structured measurement layer – Replace binary or scalar persuasion scoring with a judge that returns structured readings (relevance, stance, confidence, persuasiveness) against a configurable construct definition, then applies construct-specific update rates. The improved system can measure belief movement on a continuous scale with interpretable sub-scores, allowing fine-grained attribution of which content properties drive change.

  6. Scalable agent-environment interaction with validated action constraints – Build a simulation plane where agents observe only a limited platform view (social graph, content, discovery surfaces) and can only commit actions that pass validation against an action model. The improved system can run up to 100,000 agents with realistic platform constraints, preventing unrealistic omnipotence and enabling reproducible experiments on influence dynamics.

  7. Automated red-team campaign generation with inspectable evidence trails – Generate campaign timelines (narrative, content, amplification schedules) that produce full audit trails linking actor actions, exposure events, judge readings, and belief changes. The improved system can automatically design and execute influence campaigns for defensive testing, producing evidence chains that identify which specific messages, at which exposure frequencies, moved which beliefs.

  8. Phase-aware metric recording and checkpointing – Record phase-specific metrics (e.g., agenda share, reach fraction, processed exposures) at each lifecycle stage, with checkpoints and run metadata for reproducibility. The improved system can compare campaign trajectories across phases, identifying where interventions are most effective and enabling adaptive campaign steering based on intermediate measurements.

Sources

Related papers