Positivism

Joas and Knöbl (2009, p. 6) on how Popper demolished positivism in 1934:

As Popper lays out in his now very famous book The Logic of Scientific Discovery, which first appeared in 1934, in the case of most scientific problems we cannot be certain whether a generalization, that is, a theory or hypothesis, truly applies in all cases. In all probability, we will never be able to verify once and for all the astrophysical statement that ‘All planets move around their suns along an elliptical trajectory’, because we are unlikely ever to get to know all the solar systems in the universe and therefore we will presumably never be able to confirm with absolute certainty that every single planet does in fact follow an elliptical trajectory around its sun, as opposed to some other route. Much the same applies to the statement ‘All swans are white’. Even if you have seen thousands of swans and all of them were in fact white, you can ultimately never be certain that a black, green, blue, etc. swan will not show up at some point. As a rule, universal statements cannot therefore be confirmed or verified. To put it another way: inductive arguments (that is, inference from individual instances to a totality) are neither logically valid nor truly compelling arguments; induction cannot be justified purely in terms of logic, because we are unable to rule out the possibility that one observation may eventually be made that refutes the general statement thought to be corroborated. Positivists’ attempts to trace laws back to elementary observations or to derive them from elementary observations and verify them are thus doomed to failure. This was precisely Popper’s criticism.

Joas, H., & Knöbl, W. (2009). What is theory? In Social theory: Twenty introductory lectures (pp. 1–19). Cambridge University Press.

Folk psychological platitudes

“[…] if I wish to be understood by others, and if I wish to enlist others for my purposes, then I must present myself in a way that conforms to folk-psychological platitudes, or laws. To explain my actions – even my most outrageous ones – to others is to find an appropriate set of folk-psychological platitudes in the light of which I am understandable and predictable. […] Given the fundamental role of folk psychology, it should not come as a surprise that it is also the intuition most strongly protected by sanctions. Inability to come up with acceptable folk-psychological accounts of ourselves and others is sanctioned by disapproval, lack of acceptance, or even referral to psychiatric services. No wonder, therefore, that social psychologists find their subjects ‘telling more than they know’: rather than admit that they have no introspective access to many of their higher cognitive processes, subjects will tell folk-psychological stories of how their mind allegedly works.”
– Martin Kusch (1999, p. 239), Psychological knowledge: A social history and philosophy. Routledge.

Reckless and treacherous theorists

“[…] the most reckless and treacherous of all theorists is he who professes to let facts and figures speak for themselves, who keeps in the background the part he has played, perhaps unconsciously, in selecting and grouping them […]”

– Alfred Marshall (1885, February 24). The present position of economics. Inaugural lecture, University of Cambridge. Reprinted in Hodgson, G. M. (2005). Journal of Institutional Economics, 1(1), 121–137.

Seven persistent myths about evaluation

Social policy evaluation inherits concepts from multiple disciplines, and somewhere along the way many of them have become muddled. The list below sketches what I think are seven of the most harmful myths.

1. “Counterfactual” is a synonym for control/comparison group

There is a long history of research on counterfactual reasoning in the absence of a control group. For example, analyses of how people should and can determine the truth of a counterfactual conditional such as “If Oswald had not killed Kennedy, no one else would have” or “If I’d left the house earlier, I’d have caught that bus” (e.g., Adams, 1970; Halpern, 2015). We ponder counterfactuals like these all the time without running RCTs.

The evaluation literature has also considered how counterfactuals can be used without a comparison group (e.g., White, 2010) – including with qualitative evidence (e.g., Reichardt, 2022). In evaluation, the counterfactuals are conditionals like, “If intervention group students hadn’t been offered the mentoring programme, x% fewer would have passed their GCSEs.”

Some evidence for counterfactual outcomes is more robust than others. One example analysed by Reichardt (2022) involves asking programme participants what they think would have happened if they had not taken part in the programme. Their ability to answer this depends on their capacity for counterfactual reasoning, which may be limited for some programmes. For instance, many people believe that homeopathy is effective, even though its remedies are so highly diluted that they contain no active ingredients and cannot have any biological effect (Ernst, 2005). Counterfactual reasoning based on such beliefs could therefore falsely conclude that a programme caused improvements that would have occurred anyway.

Quantitative designs can be used to estimate counterfactual outcomes without a comparison group. One example is interrupted time series, which uses trends in pre-intervention outcomes to estimate post-intervention counterfactual outcomes. This is illustrated below:

The standard analysis estimates whether introducing the intervention shifted the outcome and whether it changed the slope, both relative to the inferred counterfactual outcomes.

When a comparison group is used to estimate counterfactual outcomes, it is a factual group – it consists of real data from real people. If it were a counterfactual group, it couldn’t be used to estimate anything. This is why I avoid the term “counterfactual group”.

2. RCTs only estimate overall average effects

If you think causal effects may vary by, e.g., participant characteristics or the context they are in, then you can test this by using moderator analysis and obtain estimates for each group. For example, trials funded by the Education Endowment Foundation usually investigate whether the magnitude of a programme’s effects differ for students who are eligible for free school meals – a proxy for socioeconomic disadvantage (e.g., Takala, et al., 2025). Westhorp and Feeny (2024) suggest using moderator analysis to test the effect of context in realist evaluations (see also Myth 4 below).

One challenge is that when testing for moderator effects for variables with a large number of levels – for example, in intersectional analyses of gender and ethnicity – models can become unwieldy very quickly. The UK census, for example, includes 19 ethnicities and 6 genders (including cis and trans identities). That results in 114 possible combinations. A promising solution is multilevel analysis of individual heterogeneity and discriminatory accuracy – MAIHDA, for short (Merlo, 2018).

3. (Trialists believe that) RCTs don’t need a programme theory

Individual causal effects are defined as the difference between what a person’s outcome would be following intervention and what it would be following control. We can’t observe this within‑person difference because each participant only experiences one condition. With random assignment, however, the average outcome of those in the intervention group minus the average outcome of those in the control group (a between‑person difference) provides an unbiased estimate of the average of individual causal effects (averages of the unmeasureable within-person differences) – even without any covariates. That’s the great thing about RCTs.

RCTs still need theory. How do you choose a control group? Outcome variables? How do you interpret estimated causal effects, i.e., what the difference between intervention and control outcomes means? Chen and Rossi (1980) explain how theory-based (what they call theory-driven) RCTs that include theory-informed covariates yield more precise estimates of effects (reducing the probability of Type II error), even though those covariates are not needed to control Type I error. That’s why sample size calculators ask how much variance in outcomes is thought to be explained by covariates. RCTs can also test mechanisms of change, e.g., using mediation analysis or through factorial designs.

A related myth is that trialists believe reality consists of variables. Paley and Lilford (2011, pp. 956-7) dispute this:

The view attributed to positivism, that reality is fragmented into variables, is a straw man. No one believes it, and classification into types is something that all researchers, both quantitative and qualitative, do. Variables are a product of measurement procedures; they are not part of the structure of reality.

Good theories of change go beyond the variables to discuss people and other entities and what they do (see Myth 6). Variables operationalise aspects of these.

Incidentally, trialists don’t need to be positivists either. Bogen and Woodward (1988) provide an accessible discussion of how to think about the relationship between theory and observed phenomena without falling into positivism.

4. Realist evaluation is a distinct type of evaluation

What people do in practice when they say they are conducting a realist evaluation or “using realist principles” often makes sense, and Pawson and Tilley’s work has helped bring theory to the fore in evaluation (I put at least one reading of theirs on my social theorising course when I was an academic). But the idea that you need a realist evaluation to find out “what works best for whom in what context” or to describe change in terms of context, mechanisms, and outcomes, is just absurd. Many evaluation approaches do this.

Paul (1967, p. 111) offers an earlier example of the what works for whom logic:

[…] the question towards which all outcome research should ultimately be directed is the following: What treatment, by whom, is most effective for this individual with that specific problem, and under which set of circumstances?

Pawson and Tilley (1997, p. 10) cite Palmer (1975, p. 150), who uses similar logic:

Rather than ask, “What works for offenders as a whole?” we must increasingly ask “Which methods work best for which types of offenders, and under what conditions or in what types of setting ?”

School climate research offers one example of contextual factors: the physical characteristics of schools; the formal and informal rules that operate within them; and the norms, beliefs, and values at play (Anderson, 1982). RCTs have examined the extent to which school climate moderates programme effects – i.e., how outcomes depend on the settings in which programmes are delivered (e.g., Low & Van Ryzin, 2014) – as well as the mediating role of climate, where a programme changes aspects of school climate and these changes help to explain its effects (e.g., Singla et al., 2021).

My reading of Pawson and Tilley (1997) and Pawson (2024) is that they’re introducing the scientific method and philosophy of science to evaluators who don’t have a formal training in a science. The ideas apply to any form of evaluation. For example, Pawson (2024, p. 42) argues that

All scientific investigation utilises explanations relating mechanisms and contexts to empirical patterns.

Science is not simply about observing regularities or estimating average causal effects. It is about explaining why observed patterns occur, how they are produced, and under what conditions they hold. This logic applies whether you are running an RCT, conducting a qualitative impact evaluation, or analysing administrative data using a quasi-experiment – and whether or not you call what you’re doing “realist evaluation”.

The “realist” in the term has a meaning from philosophy: ontological realism. Many people are realists without carrying a realist evaluator card.

5. “Theory-based evaluation” excludes trials and quasi-experiments

The original conceptualisations of theory-based evaluation, or its many synonyms (e.g., theory-driven evaluation, program theory evaluation, program theory–driven evaluation science) includes any approach that begins with a theory of change, uses it to design an evaluation that tests the theory, and revises the theory in light of findings (see, e.g., Chen & Rossi, 1983; Cook, 2000; Fitz-Gibbon & Morris, 1975; Weiss, 1972).

Quasi‑experiments are particularly dependent on programme theory. For example, studies that construct a comparison group using matching or weighting rely on identifying all important covariates. Once those covariates are specified, we can check whether matching/weighting has made the intervention and comparison groups equivalent at baseline; however, balance checks cannot tell us whether any important covariates were left out – theory is needed for that.

Additionally, a large number of causal models will be consistent with the data, so you need a theory to choose between them. For example, given only data, the following six simple causal models are statistically equivalent. Each model has three variables: outcome (out); condition, e.g., the programme being evaluated or usual practice (prog); and mediator (med).

Model 1 is the intended causal interpretation: the effect of the programme is partially explained by the mediator. Model 2, for example, says that programme and outcome are statistically associated with each other because the supposed mediator is actually a common cause.

The Magenta Book recommends using theory-based evaluation if you can’t find a comparison group (HM Treasury, 2020, p. 47). Annex A (analytical methods for use within an evaluation) locates trials and quasi-experiments outside the theory-based list. I think the Magenta Book is partly responsible for perpetuating Myth 5 in the UK – particularly through invitations to tender that follow the book, which leave evaluators having to play a semantic game to win work. I hope the upcoming revision corrects this (update: it doesn’t), and that it becomes routine practice to commission theory-based trials, quasi-experiments, and mixed methods evaluations, alongside theory-based qualitative impact evaluations and small-n case studies.

6. Theories of change are a kind of diagram

Theories of change are theories – of change. They may be illustrated using a diagram, but there needs to be an explanation, usually using prose, describing how the resources, entities (e.g., people, organisations, laws, etc.), and activities carried out lead to outputs and outcomes. Currently entrenched practices of developing a Theory of Change™ seem to lead people to forget what they learned about theory and theorising at university.

Pawson and Tilley (1997, p. 134) provide neat examples of how to organise theories of change explanations in tables. Tables like these are part of the “realist evaluation” tradition; however, the logic is clearly much more general (see Myth 4):

ContextMechanismOutcome
High numbers of prepayment meters, with a high proportion of burglaries involving cash from metersRemoval of cash meters reduces incentive to burgle by decreasing actual or perceived rewardsReduction in percentage of burglaries involving meter breakage; reduced risk of burglary at dwellings where meters are removed; reduced burglary rate overall

If a client forces you to include only a diagram in a report, plead with them to let you add the actual theory to an Appendix. We’ve been in this situation many times and used this solution once to date – let’s see if the report is published (update: it wasn’t).

7. Evaluations test “treatments”

Well, in a formally defined sense they do; however, not as people usually use the term. We’re not administering drugs – we’re evaluating, e.g., mentoring, approaches to teaching, emotional support. The term “treatment” carries with it medical baggage, which is very different to social programmes. My preference would be to use names that are as close as possible to the conditions. For instance, if we were evaluating ACME Therapy, the conditions would be something like ACME Therapy and usual practice or ACME Therapy and CBT.

The same goes for terms like contamination and dosage. Contamination is usually invoked when making the case for cluster randomisation, referring to the risk that individuals in the control group receive elements of the programme being evaluated too, e.g., students telling each other what they learned working with a mentor. Calling it “contamination” implies something toxic or infectious. A better term may be spillover, diffusion, or cross-group knowledge sharing.

Dosage is typically used to describe how much of a programme someone engages with, e.g., the number of sessions attended or hours of support. People aren’t absorbing a dose, they’re participating, doing activities that are suggested, sharing how they feel with a practitioner. Terms like engagement level, participation intensity, or simply sessions attended might better reflect what’s actually going on.

Using medical language in the context of trials in schools or criminal justice sounds especially creepy – it reminds me of A Clockwork Orange. I have been guilty of using these terms too. Sometimes it’s just too difficult to challenge medicalised tradition when there are pressing deadlines, and causal estimands like average treatment effect on the treated are pervasive in the methods literature (e.g., Cunningham, 2021, Section 4.1.2).

Revised 2 January 2026

References

Adams, E. W. (1970). Subjunctive and Indicative Conditionals. Foundations of Language, 6, 89–94.

Anderson, C. S. (1982). The Search for School Climate: A Review of the Research. Review of Educational Research, 52, 368–420.

Bogen, J., & Woodward, J. (1988). Saving the phenomena. The Philosophical Review, XCVII(3), 303–352.

Chen, H.-T., & Rossi, P. H. (1980). The Multi-Goal, Theory-Driven Approach to Evaluation: A Model Linking Basic and Applied Social Science. Social Forces, 59, 106–122.

Chen, H.-T., & Rossi, P. H. (1983). Evaluating With Sense: The Theory-Driven Approach. Evaluation Review, 7(3), 283–302.

Cook, T. D. (2000). The false choice between theory-based evaluation and experimentation. In A. Petrosino, P. J. Rogers, T. A. Huebner, & T. A. Hacsi (Eds.), New directions in evaluation: Program Theory in Evaluation: Challenges and Opportunities (pp. 27–34). Jossey-Bass.

Cunningham, S. (2021). Causal Inference: The Mixtape. Yale University Press.

Ernst, E. (2005). Is homeopathy a clinically valuable approach? Trends in Pharmacological Sciences, 26(11), 547–548.

Fitz-Gibbon, C. T., & Morris, L. L. (1975). Theory-based evaluation. Evaluation Comment, 5(1), 1–4. Reprinted in Fitz-Gibbon, C. T., & Morris, L. L. (1996). Theory-based evaluation. Evaluation Practice, 17(2), 177–184.

Halpern, J. Y. (2015). A Modification of the Halpern-Pearl Definition of Causality. Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence (IJCAI 2015), 3022–3033.

HM Treasury. (2020). Magenta Book.

Low, S., & Van Ryzin, M. (2014). The moderating effects of school climate on bullying prevention efforts. School Psychology Quarterly, 29(3), 306–319.

Merlo, J. (2018). Multilevel analysis of individual heterogeneity and discriminatory accuracy (MAIHDA) within an intersectional framework. Social Science & Medicine, 203, 74–80.

Paley, J., & Lilford, R. (2011). Qualitative methods: An alternative view. BMJ, 342, 956–958.

Palmer, T. (1975). Martinson Revisited. Journal of Research in Crime and Delinquency, 12(2), 133–152.

Paul, G. L. (1967). Strategy of outcome research in psychotherapy. Journal of Consulting Psychology, 31(2), 109–118.

Pawson, R., & Tilley, N. (1997). Realistic Evaluation. SAGE Publications Ltd.

Pawson, R. (2024). How to Think Like a realist: A methodology for social science. Edward Elgar Publishing Limited.

Reichardt, C. S. (2022). The Counterfactual Definition of a Program Effect. American Journal of Evaluation43(2), 158–174.

Singla, D. R., Shinde, S., Patton, G., & Patel, V. (2021). The Mediating Effect of School Climate on Adolescent Mental Health: Findings From a Randomized Controlled Trial of a School-Wide Intervention. Journal of Adolescent Health, 69(1), 90–99.

Takala, H., Kuo, T.-L., Duysak, E., Bhatti, S., Stoilova, E., Fletcher, A., McGuinness, N., McKaskill, M., & Fugard, A. (2025). Stop and Think: Learning Counterintuitive Concepts Evaluation Report. Education Endowment Foundation.

Weiss, C. H. (1972). Evaluation research: Methods of assessing program effectiveness. Prentice-Hall, Inc.

Westhorp, G., & Feeny, S. (2024). Using surveys in realist evaluation. Evaluation Journal of Australasia, 1035719X241292083.

White, H. (2010). A contribution to current debates in impact evaluation. Evaluation, 16(2), 153–164.

AI provenance problem

Earp et al. (2025) in a picture. You write some rough notes, magic it into a fully-formed idea using an LLM, but end up plagiarising a decades-old paper that was buried in the training set. Add that to the growing stack of concerns: AI slop, “hallucinations” (also known as falsehoods, misinformation, or BS), and looming climate catastrophe accelerated by the data and compute centres powering AI.

Earp, B. D., Yuan, H., Koplin, J., & Porsdam Mann, S. (2025). LLM use in scholarly writing poses a provenance problem. Nature Machine Intelligence.