Program Evaluation

Program Evaluation

Causal Reasoning without a Control or Comparison Group

Don't give up yet.

Anthony Clairmont's avatar
Anthony Clairmont
Jun 04, 2026
∙ Paid

Note: In this paid post, I'm going to introduce a useful program evaluation method, explain when you would use it, work an example, and give you code you can copy to do the analysis yourself. I do paid posts for a couple of reasons: 1) because certain posts are highly technical and take me a long time to write, and 2) I want to support Substack in its current form as an ad-free platform.

Evaluators love a clean control group but I never seem to get one. Ideally, if you wanted to know what an intervention did, you would compare people who got it to people who didn’t, but it seems most real-world interventions don’t work that way.

At some point, you may have thought that we ought to be able to say something meaningful about situation when what varies is not whether participants received it, but how much.

A recent working paper by de Chaisemartin, Ciccia, D’Haultfœuille, and Knau (2026) addresses this directly. Their paper is technically demanding but I’ve tried to break down the core problem and its solution because they are very useful.

A Common Evaluation Scenario

Suppose we are evaluating a community mental health organization that rolls out a new therapy program to all clients simultaneously. Every client receives some amount of treatment, but the number of sessions completed varies substantially. Some clients attend one session, others attend 20. Nobody receives zero sessions, because the organization has no waitlist control group and the program administrators are so confident in the effectiveness of the intervention that they say it would be ethically unacceptable to withhold care. At the end of the program period, we want to estimate whether more sessions led to better outcomes, for instance lower depression scores on a standardized scale.

In design terms, the treatment dose is sessions attended. The outcome is depression score change. Every client is treated to some degree, and the evaluator must extract a causal estimate from variation in dose alone.

The authors call this situation a heterogeneous adoption design, or HAD. “Heterogeneous” means that the amount of treatment varies across units. “Adoption” means all units receive the treatment at the same point in time. The defining features are: all units start untreated, then all units receive a positive dose of treatment simultaneously, but the dose varies. Nobody gets a dose of size zero.

The authors conducted a literature review of top economics journals and found ten published papers, including work in the American Economic Review and the Quarterly Journal of Economics, where either no units are untreated or fewer than 1% are. HADs are not an edge case even in high-quality published research.

User's avatar

Continue reading this post for free, courtesy of Anthony Clairmont.

Or purchase a paid subscription.
© 2026 Albert Anthony Clairmont · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture