Do Judges Behave Like Algorithms?

arXiv:2608.10400 · cs.LG · Submitted 2026-08-12 · Read on arXiv

Riya Manchanda, Eric Chen, Chloe Zhu, Cynthia Rudin, Brandon Garrett, Songman Kang

Duke University · Sungkyunkwan University

cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted at the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)

Code: https://github.com/eychen2/AreJudgesAlgorithms

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper investigates whether judges already behave like algorithms in their decision-making, using data from misdemeanor bail hearings in Harris County, Texas.

Terminology

Summary

This paper investigates whether judges already behave like algorithms in their decision-making, using data from misdemeanor bail hearings in Harris County, Texas. The authors analyze over 22,000 unique cases decided by 21 magistrates following the 2019 O’Donnell Consent Decree, which gave magistrates greater discretion in bail decisions.

The study is structured around four research questions. First, the authors identify unexplainable cases—decisions that cannot be predicted by any good model in the Rashomon set (the set of all near-optimal sparse decision trees). These cases amount to 2%-33% of each judge's cases. Manual analysis of a sample revealed that 100% of these unexplainable cases arose from either data-entry errors, missing information about charges outside Harris County, or case-specific qualitative details absent from structured data.

Second, the authors test whether each judge can be described by a simple algorithm. Using GOSDT to train globally optimal sparse decision trees with depth ≤5, they find that most judges can be approximated by small, interpretable formulas using features such as prior revocations, pending cases, and defendant age. For 16 of 21 judges, at least 85% of explainable decisions can be captured algorithmically, with accuracies between 75% and 100%.

Third, the authors examine whether judges consider the same variables. Using a bootstrapped feature selection approach that averages variable importance over the Rashomon set, they find three generally important variables across judges: age, warrant count, and pre-defined flags of specific crime types. However, the degree of importance and choice of additional variables differ substantially between judges, indicating that judges act like different algorithms.

Fourth, the authors test consistency between judges. Cross-judge evaluation shows that models trained on one judge's decisions perform poorly on other judges' cases, revealing pronounced heterogeneity. Rule mining identifies structured disagreement patterns, such as Judge 9 and Judge 4 disagreeing in 94% of cases involving individuals aged 24 or younger. A triplet consistency experiment shows judges are highly self-consistent (76% agreement on similar cases) but only modestly consistent with each other (55% agreement, close to random guessing).

The paper concludes with a paradox: judges do behave algorithmically much of the time, but their algorithms differ from judge to judge, leading to inconsistency and unequal treatment across defendants. The authors suggest that identifying and comparing these implicit rule sets could enable data-driven judicial training, standardized guidelines, and mechanisms to promote fairness, while noting that better data collection could reduce the need for standards-based decision-making in cases where missing or erroneous data currently require discretion.

Improvements for AI systems

Improvements to AI Systems:

  1. Rashomon-set-based explainability for high-stakes decisions: Build an AI system that, instead of outputting a single prediction, computes the full Rashomon set of near-optimal sparse decision trees for a given judge or policy. The system can flag unexplainable cases (those not covered by any model in the set) and automatically categorize them into data-entry errors, missing external data, or qualitative case-specific details—enabling real-time data quality audits and targeted human review.

  2. Personalized algorithmic auditing and calibration: Train a system that learns a globally optimal sparse decision tree (depth ≤5) per decision-maker (e.g., judge, loan officer, clinician). The AI can then quantify the percentage of a decision-maker's explainable decisions captured by the model, detect when a decision-maker deviates from their own implicit algorithm, and provide interpretable feedback on which features (e.g., prior revocations, pending cases, age) drive their decisions—supporting individualized training and consistency checks.

  3. Cross-decision-maker consistency detection and fairness alerts: Develop an AI system that trains separate models for each decision-maker and performs cross-evaluation to measure inter-rater heterogeneity. Using rule mining, the system can automatically discover structured disagreement patterns (e.g., Judge A and Judge B disagree in 94% of cases with defendants aged ≤24) and generate real-time alerts when a new decision would likely conflict with the majority or with a standardized guideline—promoting equal treatment.

  4. Triplet-based self-consistency scoring and drift monitoring: Implement an AI module that constructs triplet tests (pairs of similar cases) from historical decisions to compute a self-consistency score for each decision-maker. The system can continuously monitor this score over time, flagging significant drift or fatigue-induced inconsistency, and trigger refresher training or workload adjustments when self-consistency drops below a threshold.

  5. Data-driven standardization and guideline generation: Use the identified implicit rule sets from multiple decision-makers to synthesize a consensus algorithm—e.g., a sparse decision tree that maximizes agreement across all judges while minimizing unexplained variance. The improved AI can then propose standardized decision guidelines, highlight which missing data fields (e.g., out-of-county charges) most often force discretionary judgment, and recommend new structured data collection to reduce reliance on case-specific qualitative details.

  6. Interpretable feature importance aggregation over the Rashomon set: Replace single-model feature importance with a bootstrapped, Rashomon-set-averaged importance measure. The improved AI system can output stable, robust rankings of variables (e.g., age, warrant count, crime-type flags) for each decision-maker, and automatically detect when different decision-makers rely on different variables—enabling comparative audits and targeted policy interventions.

What the improved AI system can do:

  • Automatically audit any human decision-making process (legal, medical, hiring) to identify unexplainable cases and their root causes.

  • Provide each decision-maker with a transparent, simple model of their own behavior, enabling self-reflection and training.

  • Detect systemic inconsistency across decision-makers in real time, flagging potential bias or unequal treatment.

  • Generate evidence-based, standardized decision rules that balance interpretability with fairness, while pinpointing data gaps that need better collection.

  • Continuously monitor decision-maker consistency and drift, ensuring accountability and supporting data-driven policy updates.

Abstract

What if judges already behave like algorithms? As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have debated whether judges should be allowed to rely on them. Instead, we ask whether judges follow predictable, algorithmic-like rules already. If judges already follow consistent, formula-like rules based on discrete and static factors such as criminal history, age, and charge type, then judicial behavior may be improved. However, if judges rely on individualized information that cannot be identified through court data, then standards-based decision-making may be more challenging to understand or improve. This work explores these questions by studying judicial decision-making in misdemeanor bail hearings in Harris County, Texas. Using available court data, we investigate whether magistrate judges follow what resembles an algorithm; whether they consider the same variables in their decision-making; and whether they are consistent with themselves and with each other. To do this, we train machine learning models for each judge, measure variable importance metrics to determine important variables for each judge's decision-making, and analyze outcomes of similar cases for judges. Our results reveal that these judges generally behave algorithmically: their decisions can be captured by small, interpretable formulas. However, in some cases, judges differ substantially, leading to surprising inconsistency and unequal treatment across similar defendants. Identifying cases where algorithms do not explain judicial decision-making can improve the justice system by focusing attention on decisions where individualized standards, rather than rules, better explains outcomes.

Related papers