Imagine standing before a judge, knowing that an algorithm has generated a score predicting your likelihood of committing another crime.
Would you trust that score?
As artificial intelligence becomes more embedded in public decision-making, this is no longer hypothetical. From healthcare and finance to policing and courts, algorithms increasingly influence high-stakes decisions.
One prominent example is the Correctional Offender Management Profiling for Alternative Sanctions (COMPAS), a risk assessment tool used to estimate recidivism. Developed by Northpointe Institute for Public Management, Inc., now known as Equivant, COMPAS was designed to provide “decisional support” for correctional decisions involving placement, offender management, and treatment planning.
The tool came under significant scrutiny in State v. Loomis,[1] where the Wisconsin Supreme Court considered whether its use at sentencing raised due process concerns. A central issue was transparency. Northpointe treated the algorithm’s internal weighting as proprietary, limiting the defendant’s ability to understand how individual factors produced the final risk score.
Northpointe also sought to file an amicus brief addressing COMPAS’s history, accuracy, and efficacy, but the Court denied the request even while relying on the company’s public materials. Justice Abrahamson criticized that tension in a separate opinion.
The controversy surrounding COMPAS therefore raises a broader question that extends well beyond one tool: Can algorithms assist the justice system without compromising transparency, fairness, and judicial responsibility?
The Promise of Algorithmic Decision-Making
The rationale for risk-assessment tools like COMPAS predates the tool itself. In 2007, the Conference of Chief Justices endorsed sentencing practices informed by empirical research on recidivism, reflecting a broader movement toward evidence-based decision-making in criminal justice. The American Bar Association similarly cautioned that placing low-risk offenders alongside medium- and high-risk populations could actually increase the likelihood of reoffending. The point was not to displace judicial discretion, but to improve the quality of information available to judges when exercising it.
Wisconsin courts had expressed a similar concern even before State v. Loomis. In State v. Gallion,[2] the Wisconsin Supreme Court warned against making “high consequence conclusions about human nature that seem to be intuitively correct at the moment” and emphasized the importance of giving sentencing judges “more complete information upfront.” The underlying concern was straightforward. Intuition, experience, and professional judgment are valuable, but they are not immune from incomplete information, cognitive limitations, or inconsistent decision-making.
Viewed in that context, a risk-assessment score does not necessarily compete with careful judicial judgment. Instead, it may serve as an additional source of information in a process often shaped by crowded dockets, incomplete records, and limited time to evaluate a defendant’s history and circumstances. That distinction is important. In Malenchik v. State,[3] the Indiana Supreme Court treated actuarial risk-assessment tools as supplemental aids that may inform, but not replace, judicial decision-making.
The promise of algorithmic decision-making lies in using data to inform, not replace, human judgment. Its value, however, depends on reliability, transparency, and appropriate safeguards.
State v. Loomis: A Defining Moment
These issues came to the forefront in State v. Loomis, a landmark decision of the Wisconsin Supreme Court. Eric Loomis challenged the use of a COMPAS risk assessment at his sentencing on three grounds:
i) that the tool’s proprietary nature prevented him from testing the accuracy of his scoring,
ii) that reliance on group-based statistical data defeated his right to an individualized sentence, and
iii) that the tool’s use of gender as a factor was constitutionally impermissible.
The Court rejected each argument but did not simply wave the algorithm through. Instead, it adopted a more nuanced position, distinguishing between a sentencing court considering a COMPAS score and relying on it. Chief Justice Roggensack’s concurrence said this distinction was the key to the whole ruling. If a judge merely considers the score, it just informs their thinking, but if the judge relies on it, the score ends up making the decision for them. The Court further observed that any presentence report using a COMPAS score has to come with a written warning about the tool’s limits, that the company won’t disclose how it weighs the factors, that it measures risk for groups rather than for the specific individual in front of the judge, and that emerging research had raised concerns that the tool tends to flag minority defendants as high-risk more often than it should. The Court didn’t try to settle whether that claim was actually true. Instead, it treated the possibility of racial bias as something a judge needs to be told about, not as a reason to bar the tool altogether.
Throughout, the Court stressed that sentencing has to stay a matter of the judge’s own independent judgment. A risk score can never be the deciding factor in guilt, sentence severity, or supervision terms. The ruling left behind a lasting message that AI may inform judicial decision-making, but it cannot replace judicial responsibility. Though Loomis itself leaves open whether disclosure and discretion, without any mechanism to verify how judges actually weigh a warned-of risk of bias, are enough to make that principle more than aspirational.
The Scoring Method Under Scrutiny
The record in Loomis itself supplies a more concrete basis for caution than any general claim about empathy. At the post-conviction hearing, the defense’s expert testified that a sentencing judge has no way of knowing what population a COMPAS score is even measured against. The Wisconsin Supreme Court’s own list of required cautions conceded that no Wisconsin-specific validation study existed at the time. The opacity, in other words, isn’t just about withheld source code. It’s about not knowing what the number was normed to in the first place.
That concern has since found empirical support. A 2018 Dartmouth study[4] found that COMPAS’s 137-factor scoring model produced accuracy statistically indistinguishable from a two-variable formula using only age and prior convictions, and from predictions made by anonymous online respondents given a two-line case summary. Whether that finding fully holds up is contested. A 2018 rejoinder[5] argued that the comparison undersells COMPAS relative to trained evaluators, but the exchange itself makes the point. The accuracy of a “score” is an empirical question that can be tested and disputed, not a property that follows automatically from having more inputs. The rejoinder also points out that COMPAS outperformed the laypeople by a statistically significant margin, even if narrowly, and that decades of prior research already show structured or actuarial assessments beating unaided human judgment. Their conclusion is not that COMPAS is beyond criticism, but that abandoning structured risk tools in favor of unstructured human judgment would likely increase both inaccuracy and bias, not reduce it.
A post-Loomis illustration, though not a binding precedent, is an Indiana case, Boes v. State,[6] where the sentencing judge cited the defendant’s IRAS (Indiana Risk Assessment System) score as an aggravator but openly discounted it, saying, “I don’t give it a lot of weight because I don’t know how to predict the future,” and grounded the sentence primarily in the defendant’s criminal history and probation violations instead. The Court of Appeals held this was consistent with the rule in Malenchik v. State,[7] that a score may supplement, but never substitute for, a judge’s own reasoning.
Where This Leaves Us Algorithms can support judicial decision-making, but they cannot replace the reasoning and accountability justice requires. Until their accuracy and fairness can be fully tested, their role should remain limited to inform judgment, not determine it.
[1] 2016 WI 68, 371 Wis. 2d 235, 881 N.W.2d 749 (Wis. 2016).
[2] 2004 WI 42, 270 Wis. 2d 535, 678 N.W.2d 197
[3] 928 N.E.2d 564 (Ind. 2010)
[4] https://home.dartmouth.edu/news/2018/01/court-software-may-be-no-more-accurate-web-survey-takers-predicting-criminal-risk
[5] https://www.uscourts.gov/sites/default/files/82_2_8_0.pdf
[6] 139 N.E.3d 759 (Ind. Ct. App. 2019) (mem. dec.)
[7] 928 N.E.2d 564 (Ind. 2010)
For more practical litigation drafting tips, pleading and legal research insights, follow LawCompany.