Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

arXiv:2608.11727 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang

ByteDance Seed · Tsinghua University · Peking University

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 26 pages, 7 figures, 8 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: Harness-IF is a benchmark that turns operational instruction following into a rule-level measurement problem.

Terminology

Summary

Harness-IF is a benchmark that turns operational instruction following into a rule-level measurement problem. Its 642-rule library is instantiated as 60 realistic multi-turn coding items scoring 256 distinct rules, and every run yields a verdict per applicable rule rather than one task-level outcome. The benchmark scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

To separate compliance from coincidence, the paper introduces Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1–85.9% and AP-Acc 66.1–78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.

The evaluation has three complementary components. The main coding panel measures rule-level compliance and prior alignment across 12 models. E0 isolates surface precedence under four counterbalanced conflicts and nine model builds. The non-coding extension tests breadth on 40 cases using a domain-appropriate case-macro metric.

The contributions are: a benchmark that scores rules, not tasks—a 642-rule library of which 302 rules are placed on the five configurable instruction surfaces of a deployed coding agent across 60 realistic multi-turn items, and 256 receive execution-grounded verdicts, one per applicable rule per run; a metric that controls for unprompted defaults—AP-Acc scores only rules labeled as opposing the unprompted default, separating instruction following from coincidence; and evidence that all 12 models perform worse on against-prior rules under like-for-like and common-support analyses, while rule families and modalities exhibit distinct difficulty and failure signatures; E0 further reveals a robust pooled surface ordering that is inconsistent with simple prompt-depth accounts.

The main coding panel evaluates 12 frontier model builds on 60 items over three rounds: 2,160 agent–item–round runs and 40,104 rule-level verdict rows, one per applicable constraint per run. Deterministic checks (regex, AST, cross-file, command-output) cover 13.3% of eligible verdicts; a GPT-5.2 judge scores rubric constraints by three-vote majority and adjudicates hybrid checks, so 86.8% of rows involve the judge. The like-for-like binary analysis contains 37,616 eligible verdicts, including 19,449 against-prior verdicts.

Every model is less successful on the against-prior subset, so aggregate scores overstate compliance where a rule departs from model defaults. The overstatement is model-specific: it ranges from 3.6 to 7.4 points, a twofold spread. Claude-Opus-4.7 leads all four columns, so prior control does not change the top-ranked build, but it exchanges three adjacent rank pairs (2–3, 4–5, and 11–12). Accuracy spans 13.7 points across the cohort, with a standard deviation of 4.2 points. Models largely agree on which rules are hard: correlating each model’s per-rule pass-rate vector against the cohort mean, over the 242 rules every model attempted, gives 0.57–0.89 (mean 0.80).

Against-prior rules expose a consistent compliance gap. Every evaluated model scores lower on the against-prior subset. Under the like-for-like binary definition, the mean Acc–AP-Acc gap is 5.81 points across 37,616 eligible verdicts and remains positive for all 12 models. A stricter common-support analysis retains 2,430 of 3,342 (item, round, rule) observation keys (72.7%) on which every model produced a clean pass/fail outcome; item-clustered 95% intervals for the paired gap remain positive model by model.

Failures concentrate on rules that demand action. Shortfall rules absorb 77.1% of the 8,440 failures, against 20.8% for overstep rules and 2.1% for soft preferences. The asymmetry reflects exposure rather than propensity: the panel contains 27,306 shortfall instances against 8,443 overstep instances, and the two classes fail at similar rates (23.8% versus 20.8%). Most compliance failures in realistic coding work are therefore omissions of demanded behavior. Failure mass is unevenly distributed across families. Output control and workflow together account for 53.9% of all failures (27.6% and 26.3%), followed by code style (15.7%), conditional logic (12.4%), tool use (9.5%), and quantitative limits (8.6%).

Commands and output-control rules are hardest. Commanding rules are hardest by modality (76.0%), against 79.4% for numeric bounds, 79.7% for conditional rules, and 90.6% for preferences. Family differences are narrower: output control scores 70.9% against 82.6% for quantitative limits, an 11.7-point spread; it is the lowest-scoring family for 11 of the 12 builds, and no build clears 79% on it.

Surface precedence does not follow prompt depth. E0 provides a controlled, scoped test of surface precedence. System prompts, project files, and user instructions tie exactly for the best mean rank: each has a model-level rank sum of 20, hence 20/9 = 2.22.... They precede tool descriptions at 3.78 and skill descriptions at 4.56. A pooled Bradley–Terry analysis places SP, PF, and UI above TD, with SD last. The ordering survives separate direction fits, equal-cell weighting, all leave-one-pair/model-out fits, and four conservative assignments of the 27 errors. A crossed model-and-pair bootstrap preserves the complete ordering in 9,652/10,000 resamples; only 6/9 individual-build fits reproduce it exactly, so it is interpreted as a pooled cross-build tendency, not a universal hierarchy.

Reliability: the prior-alignment gap stays positive for all 12 models in the fully deterministic subset, averaging 13.09 points over 5,013 eligible verdicts. Test–retest ICC is 0.725 across models and 0.599 across agent–item cells, and an earlier human-reference audit, run under a five-vote rather than the released three-vote configuration, showed 69.0% agreement (κ = 0.515 over 919 rows). Verdicts are far less stable under a judge swap (62.1% agreement, κ = 0.163 on 116 paired clean verdicts). Since 86.8% of verdict rows involve the judge, this is the dominant uncertainty in the measurement.

The instrument transfers beyond code, the ranking does not. A separate exploratory panel scores 40 non-coding cases across five domains over 1,428 valid trajectories from the same 12 builds. Case-macro pass rates span 65.8–84.8%, led by GPT-5.5 rather than the coding leader; that panel’s metric and population differ and are never pooled with the coding results.

The benchmark is prepared for public release as a self-contained package: the 642-rule constraint library in YAML, the 60 assembled coding items with their scenario fixtures and ground-truth scoring scripts, the 2,160-record verdict panel behind every number reported here, and the evaluation and analysis code. A single script recomputes each displayed figure and table from the shipped verdict records, so the results in this paper can be reproduced offline, without model API access and without re-running any agent.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Add a prior-alignment calibration layer to instruction-following models.
  • The improved system explicitly separates rules that align with its default behavior from those that oppose it.

  • It computes a model-specific AP-Acc (Against-Prior Accuracy) score during deployment, not just aggregate accuracy.

  • It can report a confidence interval for compliance on novel instructions, flagging when a high aggregate score hides a 3.6–7.4 point drop on against-prior rules.

  1. Implement a rule-level, execution-grounded compliance monitor instead of task-level success.
  • The improved system tracks each applicable rule as a separate verdict (pass/fail) from actual tool outputs, file states, and command results.

  • It can identify which specific rule failed (e.g., must not modify imports vs. must add error handling) and why, enabling targeted retraining or runtime correction.

  • It distinguishes shortfall (omitted action, 77.1% of failures) from overstep (forbidden action, 20.8%) and soft preference (2.1%), so it can prioritize missing behaviors over excess ones.

  1. Add a surface-precedence engine that does not assume prompt depth.
  • The improved system learns a pooled precedence order (system prompt > project file > user instruction > tool description > skill description) from counterbalanced conflicts, rather than assuming deeper prompts win.

  • It can resolve conflicting instructions by consulting this learned order, and it can flag cases where a tool description contradicts a user instruction—since the latter wins in the pooled data.

  • It can also detect when its own precedence behavior deviates from the pooled tendency, allowing per-deployment adjustment.

  1. Build a rule-difficulty predictor and failure-family classifier.
  • The improved system predicts which rules are hardest based on family (output control at 70.9% vs. quantitative limits at 82.6%) and modality (commanding at 76.0% vs. preferences at 90.6%).

  • It can allocate more reasoning steps, self-checks, or verification passes to output-control and workflow rules (53.9% of all failures).

  • It can automatically generate extra test cases for shortfall-prone rules, since omissions dominate failures.

  1. Add a judge-stability and uncertainty-aware evaluation mode.
  • The improved system, when using a language-model judge for rubric scoring, reports agreement under judge swaps (κ=0.163) and human-reference audits (κ=0.515).

  • It can flag verdicts with low inter-judge stability and either re-run with a different judge or mark them as low-confidence.

  • It can switch to deterministic checks (regex, AST, cross-file, command-output) for 13.3% of rules to reduce judge dependence, and it can estimate the margin of error on any reported compliance score.

  1. Enable offline, reproducible compliance audits.
  • The improved system ships with a self-contained rule library (YAML), scenario fixtures, and verdict records, so any deployment can recompute AP-Acc, per-rule pass rates, and surface-precedence rankings without API access.

  • It can run a prior-withheld probe (re-running tasks with a rule removed) to measure how much of its compliance is coincidental versus genuine instruction following.

  • It can generate a per-model overstatement margin (the Acc–AP-Acc gap) and adjust its reported compliance claims accordingly.

  1. Add a cross-domain transfer warning.
  • The improved system knows that its coding-panel ranking does not transfer to non-coding tasks (e.g., GPT-5.5 leads non-coding, not the coding leader).

  • It can refuse to extrapolate compliance scores from code to other domains, and instead run a separate domain-appropriate metric (e.g., case-macro) before making claims.

  • It can flag when a model’s rank changes across domains, preventing overconfident deployment decisions based on a single benchmark.

  1. Implement a failure-mass rebalancing mechanism.
  • The improved system, seeing that output control and workflow cause 53.9% of failures, can dynamically increase its verification effort on those families (e.g., re-reading the output spec, checking file writes, validating command sequences).

  • It can also detect when it is about to omit a demanded action (shortfall) and trigger a did I do X? self-check, since shortfalls are 3.7× more common than oversteps.

  • It can use the per-rule pass-rate correlation (mean 0.80 across models) to identify universally hard rules and pre-train on those specific patterns.

What the improved AI system can do:

  • Report both aggregate accuracy and AP-Acc, with a clear statement of how much its compliance is inflated by prior defaults.

  • Pinpoint exactly which rule failed, in which family and modality, and whether it was an omission or a forbidden action.

  • Resolve conflicting instructions using a learned surface precedence order, not a depth heuristic.

  • Self-assess the reliability of its own judge-based verdicts and switch to deterministic checks when needed.

  • Reproduce all its compliance numbers offline, from shipped verdict records, for any auditor.

  • Avoid overclaiming cross-domain ability, and instead provide domain-specific metrics.

  • Focus its reasoning and verification resources on the rule families that are most failure-prone, reducing omissions and oversteps in real coding tasks.

Abstract

When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.

Sources

Related papers