On Agreeing with AI
Exhibit A: an AI stroke-rehabilitation study
The phenomenon of humans and AI agreeing with each other has already received abundant commentary under the header of sycophancy, which is what happens when AI appears to bend its judgment to support the ideas of whomever it is currently talking to. However, a paper recently crossed my desk that got me thinking about what is arguably a much bigger problem with human-AI agreement. Like most big issues in AI, this one seems to be more of a human problem than a technical one. The problem is that we are using human-AI agreement to evaluate when AI is appropriate for use in decision-making.
To understand exactly why this is a problem from an evaluation perspective, let’s walk through a recent study as an example. Consider Evaluating Equity in AI-Supported Functional Assessment: Agreement Between Clinician Judgment and Digital Metrics in Stroke Rehabilitation, published this year in Evaluation & the Health Professions. The paper asks: when an AI system measures how well a stroke patient moves, and a clinician measures the same thing by eye, how far do the two agree, and what accounts for the places where they part company?
Before getting into study, I’ll remind the reader that my cards are on the table. I’ve argued that Evaluation is the Future of AI and that the center of gravity in this technology is moving away from getting systems to function at all and toward defining what functioning well would mean, and that the discipline of evaluation is being conscripted into that project whether or not it is ready (mostly not). I’ve proposed standards for what a defensible evaluation looks like: it posits non-arbitrary criteria and standards, it gathers performance data relevant to those criteria, it compares the two probabilistically so as to estimate how far the evaluand departs from the standard and in which direction. The paper I’m writing about here gives us a chance to watch our own trade performed by people who inherited a different set of habits, and to speculate about how much of the conclusion those habits fix before any data are gathered.1
Studying the AI stroke treatment system
The design is a cross-sectional agreement study. Ninety adults with chronic stroke, all at least six months past the event and receiving outpatient rehabilitation, each performed a standardized battery of motor tasks: upper-limb reaching, sit-to-stand transitions, and level-ground walking at a comfortable pace. Ninety-six licensed physical and occupational therapists rated those performances. Added together, the authors report a sample of 186 “participants.”2
Two measurements were taken of each performance. The first came from an AI pipeline: inertial sensors on the wrists, ankles, sternum, and lumbar spine, together with a depth camera feeding data to pose-estimation and biomechanical software, producing features such as speed, smoothness, physical symmetry, and consistency. The second came from a clinician, who scored the same task immediately afterward on a set of study-specific items, a zero-to-ten rating anchored to observable speed, quality, physical symmetry, and adaptability.
The construct at the center of the paper is what the authors call the “AI-clinician discrepancy score,” the AI metric minus the clinician rating for the same task. To make that subtraction possible, they rescaled the AI outputs onto the clinician’s range. A positive score meant the machine judged the patient more capable than the clinician did.
From there the analysis runs along three tracks. First, agreement, measured with intraclass correlation coefficients, the ICC(2,1) two-way random-effects model with absolute agreement. The authors supplement the ICCs with two precision statistics: the standard error of measurement, and the minimal detectable change (MDC) at the 95% level.3 Second, group comparisons: does agreement differ by clinician discipline and years of experience? Third, a multivariate regression predicting the discrepancy score from clinician characteristics (discipline, experience) and patient characteristics (age, time since stroke, motor severity, cognition).
The headline results look tidy. Agreement was moderate to good, with ICCs from 0.68 to 0.79. The strongest agreement was for task execution speed (ICC of 0.79); the weakest was for lower-limb symmetry (ICC of 0.68). Across every domain the AI scored patients slightly higher than the clinicians did, with mean differences of roughly 1.8 to 2.9 points and small-to-moderate effect sizes (Cohen’s d from 0.29 to 0.50). Physical therapists agreed with the AI a little more than occupational therapists did (0.75 against 0.70). More experienced clinicians agreed more than less experienced ones (0.77 against 0.68). The regression, which accounted for about 28% of the variance, found larger discrepancies with older patients and longer time since stroke, and smaller discrepancies with greater motor impairment. Clinicians leaned on their own judgment more than on the AI, especially for the mildly impaired.
The conclusion is the reasonable, unsurprising one that nearly every study of this kind reaches: AI-generated metrics show meaningful concordance with clinician judgment and are best treated as a complementary tool that requires contextualized interpretation. This is of course generally good advice, but it’s also 1) a trivial outcome from a peer-reviewed study, and 2) not a valid evaluative conclusion given the design.

The seduction of a sensible conclusion
The study has performance data, two streams of it, the AI’s and the clinician’s. It has a comparison mechanism, the discrepancy score and the ICCs. What it lacks is a criterion and a standard: a defensible statement of what good functional assessment would be, against which either measurement could be judged. In place of that, the discrepancy score smuggles in an implicit standard through its own arithmetic. Discrepancy in this study is defined as AI minus clinician, so the clinician’s judgment is the ground truth, and the AI is scored on how far it strays from that truth.
The AI is judged wrong to the extent that it diverges from the clinician; the clinician is right because we placed the clinician on the correct side of the subtraction. No reality outside the two measurements is ever consulted. We are asked to certify one ruler by holding it against a second ruler, without asking whether either ruler is any good, or good for what. The study then goes on to demonstrate, in its own results, that the second ruler changes length depending on where you take it from! The study shows that clinician judgment varies by discipline. It varies by experience. Occupational and physical therapists disagree with the machine by different amounts, which means they disagree with each other, which means “the clinician rating” is not a fixed quantity that depends on who happens to be holding the clipboard today. The benchmark, looked at closely, will not hold still but is not represented as a random variable either.
The authors have, in their own reference list, the study that should have upset the whole enterprise. They cite Sylolypavan and colleagues in passing to note that clinicians vary in how they read a case. Sylolypavan’s team had eleven intensive-care consultants at a single Glasgow hospital independently label the same sixty patients on a five-point severity scale, then built a separate classifier from each consultant’s labels. Even on the training data, the eleven experts agreed only fairly (at a Fleiss’ kappa of 0.383). When the resulting eleven classifiers were turned loose on an external dataset, their pairwise agreement fell to a minimal level (average Cohen’s kappa of 0.255). The experts disagreed more about who was ready for discharge (kappa of 0.174) than about who was going to die (kappa of 0.267). Essentially, expert clinical judgment, on a well-defined task with only sixty cases, carries so much noise that no single expert’s labels can stand in for the truth.
In the Sylolypavan study, the team went looking for a “super expert” whose judgment could serve as the gold standard and found no support for it.
Internal skill at reproducing one’s own labels correlated only weakly, and in one condition negatively, with external accuracy.
Second, the obvious fix, a majority vote across all eleven, produced a worse model than a vote restricted to the experts whose judgments were internally learnable in the first place (with an kappa of 0.254 against 0.438). The lesson is that a defensible standard can sometimes be assembled out of noisy expert judgments, but only after you decide, on principled grounds, whose judgments to keep and whose to set aside. That decision is itself an evaluation problem of the Matryoshka variety, and it is precisely the decision that we assume away by anointing whichever clinician happened to be in the room.4
What would a symmetric version look like? We would begin by defining the criterion, what functional recovery consists of for this population, in terms a patient would recognize as mattering, and a standard for it, drawn either from prior programs or from the practical demands of daily life. It would then treat the AI and the clinician as two fallible measurements of that latent construct, and ask how each relates to the standard rather than only how they relate to each other. The clinician would stop being the yardstick and would take up the more honest role of a second imperfect instrument aimed at the thing we care about.5
Outputs, again
Outputs are the first-order products of running the thing. Outcomes are the reason you ran it. AI evaluation has been banging its head on the low ceiling of its own expectations because it keeps measuring impressive outputs and mistaking them for outcomes, certifying systems that “perform well on benchmarks” and then turn out to be junk in the world.
Agreement is an output. The outcome that would justify any of this is whether AI-supported assessment leads to better decisions and better recoveries: shorter time to functional independence, fewer missed impairments, therapy dosed more appropriately, scarce clinician time spent where it helps. On that question the study is mute, and to its credit it says so plainly in the limitations. Michael Scriven called this problem the fallacy of statistical surrogation, the use of a mere correlate in place of the criterion of merit. Agreement with a clinician is a correlate. The criterion of merit is whether the patient gets better.
The point that I think evaluations be positioned to press is that the most valuable behavior of an AI assessment tool may be exactly where it disagrees with clinicians. A machine that only echoes what the therapist already sees is an expensive redundancy. Its potential payoff is in catching what the eye misses, the thing a busy clinician on the fifteenth patient of the day does not register. An evaluation that rewards high agreement will therefore favor precisely the tools that add the least, which is a perverse thing to aim at. What we should want to know is whether the disagreements are the machine’s errors or its contributions, and that is a question about outcomes, answerable only by following patients forward and seeing whose judgment the world went on to vindicate.
Statistical interlude
(Note: this section is fully skippable if you are not interested in statistics.)
One of the reason I just had to use this paper as an exemplar is because it filled up my bingo card of issues in contemporary evaluation. Let’s talk about the statistics.
In fairness, one point in the paper’s favor should go first. I have repined that AI benchmarks tend to skip error estimates altogether and report only a top or mean score. This paper does better than that. It computes a standard error of measurement and a minimum detectable change (MDC) for every domain. The difficulty is that, having gone to the trouble of producing the error estimates, it then reads them the wrong way round.
The study reports mean AI-clinician differences of about 1.8 to 2.9 points and calls them statistically significant, with p-values bunched just under the conventional line, from .041 down to .004. It also reports the MDC, the values of which land between roughly 10.7 and 13.6 points. The differences the paper treats as its substantive findings, two to three points, are between a fifth and a quarter of the study’s own threshold for a difference large enough to be told apart from noise. The paper claims MDC “support[s] the meaningful interpretation of observed differences” when the differences sitting below MDC undercuts individual-level interpretation.
Note that this is not a mistake in frequentist terms. It is just what statistical significance does once the sample is large enough. It certifies that a difference is probably not exactly zero and stays silent about whether the difference is large enough to matter. Significance speaks to the reliability of a sign. It says nothing about magnitude, which is what MDC is.
A Bayesian approach, on the other hand, would model the sensor reading and the clinician reading as noisy observations of a latent functional ability, put priors on the two measurement processes and their errors, and give back a posterior over the quantity you actually care about: the probability that true function exceeds some standard, or the probability that the two instruments diverge by an amount large enough to change a decision. In place of “the difference is significant, p equals .004,” you would get “there is an 80% probability the AI overstates function by more than one clinically meaningful unit,” or, just as usefully, “the difference is almost certainly smaller than anything we would act on.” A clinician can actually use either of those.
Informative priors are not a luxury here. The expectations we bring to a system built for gait timing differ from the ones we bring to a system judging movement quality, because the two constructs differ in how cleanly they reduce to numbers. The study half-discovers this: agreement was highest for speed and lowest for symmetry and quality, which is the pattern you would predict once you accept that different domains warrant different priors about how well any instrument, human or machine, can pin them down. Remember: standards need not be standardized. A model for a stopwatch-like construct is not the model for a gestalt.
The measurement level of the study is just as instructive.
To compute a discrepancy between a sensor’s output and a clinician’s zero-to-ten rating, you first have to put the two on the same scale. The authors did this by min-max scaling each AI metric onto the clinician’s range, mapping its observed minimum and maximum onto the endpoints of the rating scale. That is defensible as an expedient but indefensible as a foundation, because it turns the two headline quantities of the paper, the size of the discrepancy and the absolute-agreement ICC, into downstream artifacts of an arbitrary transformation.
A sensor’s estimate of gait speed is in meters per second. It has a true zero and equal intervals. A clinician’s zero-to-ten rating of “adaptability” is an ordinal impression with no guarantee that the step from six to seven equals the step from eight to nine. Forcing these onto a shared ten-point metric and subtracting one from the other is just sleight of hand. It yields a number, and the number invites arithmetic, means and differences and effect sizes, that the underlying measurements do not license.6
Had we z-scored instead, or matched medians, or anchored to an external criterion, the discrepancies would take different values. The ICC for absolute agreement, unlike the ICC for consistency, is itself sensitive to exactly these shifts and rescalings. The paper’s headline findings all inherit their values from these arbitrary normalizations.
This is Equity?
If an AI assessment diverges from clinical judgment more for some patients than for others, deploying it distributes its errors unevenly across a population. The word “equity” is in the title but what the paper actually does is a subgroup analysis. It asks whether AI-clinician discrepancy varies with patient age, time since stroke, motor severity, and cognition, and finds that it does, with larger gaps for older patients and for those further out from their stroke. Should the machine prove least concordant for exactly the older, longer-recovering patients, an uncritical rollout would degrade assessment quality most for the people already least well served. That is the seed of an equity argument.
But it is only a seed. A study that took equity as its object rather than its keyword would have had to say whose equity, along which dimension, measured against what standard of fair treatment, and then disaggregate accordingly, ideally tying the differential error to a consequence a patient would feel. This paper does not disaggregate or socioeconomic position or any protected characteristic except sex (which it does not enter into the regression). It does not connect the discrepancies to differential decisions or outcomes. And because it has no external criterion, it cannot even say whether the larger discrepancy for older patients is the AI failing them or the clinicians failing them.
“Equity” here is a vibe more than a method. I point this out not because the authors are unusually guilty, since the term is used this loosely throughout health-AI research, but because our field, if it is going to be useful in this area, has to hold the line that equity is a claim about legitimate standards appropriately applied, and not a synonym for having looked at subgroups.
The finding that agreement is high for speed and low for symmetry and quality is, on a generous reading, the study’s most useful result, and it is an equity result in disguise. It says the tool is trustworthy for the parts of function that reduce cleanly to timing and shaky for the parts that require judgment, and the parts that require judgment are disproportionately the ones that matter for complex, atypical, or severe presentations. Different patients bring different mixes of those constructs. So the tool’s reliability is not uniform across patients, and the non-uniformity falls in a pattern with fairness implications. The paper holds all the pieces of that argument and assembles none of them, because it lacks the frame that would let it see the pieces as related.
Conclusions
The AI system in this study is probably fine. The sensors work, the pose estimation works, the metrics are plausible. The tool was never the problem. The problem was the question the tool was asked to answer, “do you agree with us,” which is the wrong question in a way that no amount of statistical care downstream can repair. Ask a machine whether it agrees with you and it will try to do that. Ask whether it is right, and how you would know, and for whom it is least right, and you have the beginnings of an evaluation.
My prediction is that we are going to see a great many studies shaped like this one over the next few years, as AI moves into every corner of clinical and social practice and the people closest to those practices reach for the tools they already have. Many of those tools measure agreement. Part of our job, more or less, is to keep insisting that the machine’s task was never to agree with us. It was to help us be right.
My point here is not to bother a group of rehabilitation scientists who did a competent job by the standards of their own literature. They did. By the local norms of health-professions agreement studies, this is an ordinary, publishable, unembarrassing paper.
But this is a failure mode I now wonder about as our field gets pulled by the tides of history into AI.
Last year, I wrote:
I think we will discover some reverse salients in evaluation as a discipline as it gets drawn into AI evaluation specifically. A reverse salient is a part of a sociotechnical system that is less developed than other parts, creating a gap in the advancement of the technology. Ideally, learning about the reverse salient causes innovators to swarm the bottleneck and innovate. Right now, the larger evaluation field is not fully ready to parent our fledgling subfield. In particular, our skill gaps and bad ideas will become society’s problem if they are applied to AI evaluation. On the other hand, formally-trained evaluators have much to contribute if we can rise to the occasion.
I think this paper marks several reverse salients. The absence of legitimate, externally anchored standards is one. The reflex of treating consensus as a proxy for truth is another. The gap between statistical significance and evaluative magnitude is a third. Metrological malpractice is a fourth. Each is a place where a formally trained evaluator, walking into a room full of capable clinicians and engineers, could say something the people in that room did not already know.
This is also pretty much my position about what AI should do in medical and other high-stakes settings - tell us something we don’t already know. If the plan is replace medical staff, then getting AI to agree with us is a plausible goal – I can imagine an emergency medical system like Robert Picardo’s hologram in Voyager or the MedPods in Prometheus, but obviously nobody sane wants to be treated by one of these on Earth when human medical staff are available.7 However, even in the case of replacing staff, we can agree that we want automated medical systems to perform as well as possible, not just at the level of the average clinician, who, as we discovered above, is about elusive as Quetelet’s homme moyen. If the plan, instead, is to augment decision-making, human-AI agreement literally doesn’t matter at all. In fact, there’s every chance we will accidentally smooth out the wrinkles in the AI’s digital brain trying to get it to agree with us. After all, this is why seeking a human second opinion is common practice. When the second opinion comes from a machine, an evaluation cycle optimized around meaningful outcomes actually stands a chance of correcting a human mistake or improving results.
In Bringing AI Evaluation into the Fold we walked around the cafeteria of evaluation subfields, product, educational assessment, program, and personnel, and asked which table AI evaluation should sit at. By that map, the study in front of us arguably belongs at the educational-assessment table rather than the program table because it is a measurement-validation exercise, asking whether the readings of a new instrument track those of an accepted one. What it carries with it, though, are the habits of clinical agreement research, and the distance between those habits and the ones an assessment specialist would bring is the source of most of what follows.
The 186 is the sum of two separate groups, 90 patients and 96 clinicians, rather than a count of paired observations. That would be harmless bookkeeping if the multivariable regression did not appear to run on all 186 of them. The reported test, F(6, 179), implies a model with six predictors fit to 186 cases. Four of those predictors belong to patients (age, time since stroke, motor severity, cognitive status) and two belong to clinicians (discipline and years of experience), while the outcome being predicted, the discrepancy score, is a property of a patient’s task performance. One row of that regression would therefore have to hold a discrepancy score alongside values on both the patient-level and the clinician-level predictors at the same time, and nowhere does the paper describe the matching of patients to clinicians that would let such a row exist. I think two reconstructions are available and neither is stated. If a row is a single patient-clinician assessment, the unit is a dyad, and a count of 186 is hard to reconcile with 90 patients each rated by some number of clinicians. If instead the 90 patient records and the 96 clinician records were simply stacked, then every row lacks values on the predictors that belong to the other group, and an ordinary complete-case model would drop all of them. I can’t figure out the actual unit of analysis here, so I can’t interpret the regression coefficients. I don’t think a reader should have to recover this from the degrees of freedom.
MDC refers to the smallest change in score that exceeds what measurement error alone would produce at that confidence level - a very handy idea if you aren’t already familiar with it.
Users of the Symmetric method will recognize the shape of this issue. It is the same challenge the Prior Interpreter tool was built to handle: you have a spread of expert and non-expert statements about a standard, and the task is usually not to average them but to weight them, with an explicit multiplier for expertise. Sylolypavan’s learnability filter is one data-driven answer to the weighting question: find the internally-consistent raters and upweight their scores.
More technically, this is an identification problem: with exactly two error-laden indicators and no external referent, the shared variance between them is not identified as true-construct variance. If the kinematic pipeline and the clinician's eye are both pulled by how textbook-normal a movement looks rather than by how functional it is, their correlated error inflates the agreement, and whatever latent thing you would extract from the pair is contaminated by that common method factor. You cannot pull shared signal apart from shared bias from inside the pair. You need a third real thing that both instruments must answer to.
A smaller observation compounds the worry. The methods describe the clinician scale as running zero to ten and the AI metrics as normalized to match it, yet every number in the results tables (means near 60, standard deviations near 11, standard errors near 4, MDC values above 10) lives unambiguously on a zero-to-hundred scale. An MDC of 13.6 is not even possible on a zero-to-ten instrument, since it exceeds the range. Something in the reporting is internally inconsistent, and the text gives no way to tell which description is the mistake. Likewise, the motor-severity variable is defined so that higher scores mean better function, and is then described in the results as “greater motor impairment severity,” so the direction of one of the paper’s strongest effects reverses depending on which sentence you trust.
After a brief Google, I realized I needed to add the word “sane” into that sentence.

