Thomas Aston (2026) has written a really interesting discussion of debates about what counts as realist evaluation and whether it is a unique kind of evaluation.
As I’ve written before (and Aston cites), I think theorising contexts, mechanisms, and outcomes (CMOs) is important, as is recognising a gap between theory, phenomena, and evidence. However, as Pawson (2024, p. 42) notes: “All [!] scientific investigation utilises explanations relating mechanisms and contexts to empirical patterns.” It’s science-as-usual. Pawson and Tilley know this – they acknowledge and cite earlier work. For example Pawson and Tilley (1997, p. 10) cite Palmer (1975, p. 150):
“Rather than ask, ‘What works for offenders as a whole?’ we must increasingly ask ‘Which methods work best for which types of offenders, and under what conditions or in what types of setting?'”
The link with science-as-usual is clear in Pawson and Tilley (2001, p. 324):
“We have argued that good evaluation is good social science. For us, this embraces the gallant aims of precision in articulation of theory, rigor in empirical testing, confederation in lines of inquiry, and cumulation in the body of findings. The ‘realist movement,’ of which we are a part, is often considered the brash upstart of the evaluation schools. In fact, it depends on these rather venerable ideas. The future, for us, thus lies in keeping faith with some of the grand old principles of social science and in not forgetting the hard-won lessons of the old studies.”
Evaluation method(ology) debates remind me of debates about psychological therapy brands, like CBT, psychodynamic, or humanistic approaches. Mick Power (2010) is an example therapist-academic who had a go at dismantling the brands, and pulled out graded exposure, transference, and challenging dysfunctional assumptions as example techniques that are used across a range of approaches. Others have done similar, e.g., the behavioural change technique taxonomy (Michie et al., 2013) extracts a long menu of approaches that can be combined to develop a programme.
I think evaluation would make faster progress if we routinely dismantled the big evaluation brands and asked of an apparently new and unique approach:
- Is it really unique?
- Who else has done similar?
- What is the approach an instance of?
- How could the approach be used in another type of evaluation?
- How could a technique be combined with another?
Westhorp and Feeny (2024) used this style of reasoning for realist evaluation, and showed how regression models with interaction terms can be used to test CMOs. Baumgartner and Falk (2023) explore regression-based alternatives to the Quine-McCluskey logical minimisation algorithm used in qualitative comparative analysis (QCA), closing the gap between QCA and statistics.
When it comes to questions of metaphysics, ask: what other approaches make the same or similar assumptions, e.g., concerning ontology and epistemology? For example, there’s a trace of the gap between mechanism and measure in questionnaire design: “operationalisation”, the process of developing a measure of, e.g., a psychological construct. This process begins with the premise that the phenomena are not directly observable. It’s psychometrics 101.
And this critical dismantling strategy applies to quantitative impact evaluation approaches too. Consider “synthetic controls” developed using the synthetic control method. They sounds unique and special; however, peek behind the curtain and they’re weighted averages, with weights estimated by balancing pre-intervention measurements of the outcome and covariates. So perhaps a comparison group developed using inverse probability of treatment weighting is also a “synthetic control”…? Maybe a t-test uses synthetic treatment and control groups, with unit weights…? Or consider entropy balancing, which, when targeting the ATT, is equivalent to inverse probability tilting (IPT; Słoczyński et al., 2025). The only difference is that entropy‑balancing weights sum to the comparison‑group sample size, whereas IPT weights sum to the intervention‑group sample size.
Build bridges between the methods and methodologies and they take up less space in your brain, making it easier to design robust evaluations and communicate clearly what you’re really setting out to do.
References
Aston, T. (2026). Evaluation blogs, podcasts, and webinars in the second half of 2025: Roundup review V. Evaluation.
Baumgartner, M., & Falk, C. (2023). Configurational Causal Modeling and Logic Regression. Multivariate Behavioral Research, 58(2), 292–310.
Michie, S., Richardson, M., Johnston, M., Abraham, C., Francis, J., Hardeman, W., Eccles, M. P., Cane, J., & Wood, C. E. (2013). The Behavior Change Technique Taxonomy (v1) of 93 Hierarchically Clustered Techniques: Building an International Consensus for the Reporting of Behavior Change Interventions. Annals of Behavioral Medicine, 46, 81–95.
Palmer, T. (1975). Martinson Revisited. Journal of Research in Crime and Delinquency, 12(2), 133–152.
Power, M. J. (2010). Emotion focussed cognitive therapy. John Wiley & Sons.
Słoczyński, T., Uysal, S. D., & Wooldridge, J. M. (2025). Covariate Balancing and the Equivalence of Weighting and Doubly Robust Estimators of Average Treatment Effects. IZA Institute of Labour Economics Discussion Paper, 18147.
Westhorp, G., & Feeny, S. (2025). Using surveys in realist evaluation. Evaluation Journal of Australasia, 25(1), 45–64.
Suggested citation: Fugard, A. (2026, June 1). Dismantling evaluation brands [blog post]. https://andifugard.info/dismantling-evaluation-brands/
This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀