Theory‑based (or driven) evaluations begin with a theory of change and then design the evaluation to test the causal mechanisms it proposes (Fitz‑Gibbon & Morris, 1975). A well‑constructed theory of change sets out how a programme’s resources are deployed to support delivery activities, and how these activities activate the causal mechanisms that generate outcomes as they evolve over time. The early – and, to my mind, more coherent – formulations of theory‑based evaluation were pluralistic, encompassing all methodological approaches (see, e.g., Chen, 2015): RCTs and quasi‑experiments alongside, for example, process tracing and qualitative comparative analysis.
Mechanisms can be defined in terms of entities and what they do to bring about change (Illari & Williamson, 2011). We can identify those entities and activities with the aid of substantive theories relevant to the policy area under investigation, and the mechanisms are described using the concepts provided by those theories (Ioannidis & Psillos, 2018). A substantive theory is particularly helpful when there is evidence for the mechanisms it proposes, rather than when it rests only on armchair speculation. However, all theories are necessarily incomplete and evolve as new evidence accumulates. Some theories are more detailed than others. Some have been more thoroughly tested than others.
Substantive theories describe mechanism at different levels of explanation. Sun et al. (2005) describe the levels as follows:
| Object of analysis | Type of analysis | Elements in model |
| Inter-agent | Social/cultural | Collections of agents |
| Agents | Psychological | Individual agents |
| Intra-agent | Componential [I think also psychological] | Modular construction of agents |
| Substrates | Physiological | Biological realisation of modules |
For an inter‑agent analysis, the focus is on how collections of agents (often people) interact with one another. Examples include models of crowd behaviour when a fire alarm sounds, or systemic processes that shape people’s experiences, such as racism, sexism, transphobia, and their intersections.
For an agent‑level analysis, the focus shifts to individual people. This is roughly the way we talk about individuals in everyday life, including their desires, beliefs, opportunities, and actions, for instance, spending money, talking, listening, going for a run, attending a mentoring session, or doing homework.
An intra‑agent analysis draws on concepts from psychology. This may include theories of cognitive control; for instance, why people sometimes respond automatically and at other times engage in deliberate thought, and how conflicts between competing cognitive systems are resolved. It may include theories of emotion, such as how people’s goals and their (often automatic) inferences about progress toward those goals give rise to feelings of happiness, sadness, fear, or anger. Theories of memory systems also sit here, including how information is temporarily represented in visuospatial and phonological working memory, and the capacity and processing limits of these systems.
Finally, analysis at the substrate level considers the biological processes that implement psychological‑level mechanisms or reflect the consequences of behavioural change. This is most visible in dietary interventions, where targets might include blood glucose or cholesterol levels. It can also include theories from cognitive neuroscience, such as which neural systems underpin cognitive control or memory.
I revisited the theories of change in two evaluations I worked on to see which elements from these four levels of analysis were present.
Case study 1: Basic Maths Premium
Basic Maths Premium was a Department for Education pilot that provided additional funding to post‑16 providers in disadvantaged areas to improve GCSE maths resit outcomes for students with prior attainment at grade 3 or below. It tested three funding models using an RCT: two with guaranteed elements and one based entirely on payment by results (PbR):
| Model | Unconditional funding | Conditional funding |
| A | £500 × number of eligible students | |
| B | £250 × number of eligible students | £250 × number of successful students |
| C | £500 × number successful students |
The theory of change was at the agent level. For unconditional funding, the theory was straightforward: funding would enable activities expected to improve outcomes, such as more teaching hours, smaller class sizes, and greater use of technology. By contrast, the theory underpinning conditional funding was less unconvincing. Financial incentives were assumed to boost staff motivation, which would in turn enhance teaching quality, student motivation, and ultimately learning outcomes. However, for Model C in particular, it was uncertain how providers were expected to finance any additional activities, given that all funding under this model depended on student pass rates after any teaching had ended. Two quotations from heads of maths, drawn from the implementation and process evaluation (IPE), captured the issue succinctly (Scott et al., 2024):
- “The fact that we’ve got this potential funding in the future that will reward us for that, that’s great, but we’ve still got to find the funds now to do what we do” (Model C – conditional £500 per passing student).
- “I think [payment by results is] really unfair because you would not really know how much money you were going to get. I don’t think we would have been able to spend any additional money on that basis, so for us it wouldn’t have really worked. […] I would never have been able to employ two staff on the basis that I might get a certain amount of students through a GCSE. That would have been too much of a financial risk” (Model A – unconditional £500 per student). Also a good example of counterfactual reasoning without a comparison group.
The IPE showed that actual spend was primarily driven by the unconditional funding:

Case study 2: Stop and Think
Stop and Think is a computer‑based intervention designed to help primary pupils overcome common misconceptions in maths and science. It focuses on strengthening children’s inhibitory control: the ability to pause, reflect, and override intuitive but incorrect responses, by guiding them through short, game‑like activities that present counter-intuitive problems.
The theory underlying the programme included intra-agent and substrate levels of analysis. Briefly, the idea is that when students learn new concepts, they need to overcome intuitively obvious prior beliefs. Mareschal (2016) summarises evidence of the cognitive mechanisms involved when intuitions and new learning clash with each other, e.g., the inhibition of pre-existing beliefs involves processes implemented in the dorsal lateral prefrontal cortex (DLPFC) and the anterior cingulate cortex (ACC). Picture below (Mareschal, 2016, p. 115):

One of the tests of the theory in the RCT evaluating Stop and Think (Takala, et al., 2025) involved the construction of a measure of misconceptions driven by prior beliefs. The impact of the programme on misconceptions was then tested quantitatively using mediation analysis.
Importantly, this variable-based analysis is not the mechanism. Instead, the misconceptions measure was an indirect operationalisation of a trace of the underlying mechanism illustrated in the picture above. This distinction is summarised by Paley and Lilford (2011, pp. 956-7):
“The view […] that reality is fragmented into variables, is a straw man. No one believes it, and classification into types is something that all researchers, both quantitative and qualitative, do. Variables are a product of measurement procedures; they are not part of the structure of reality.”
Conclusions
Being explicit about the entities involved, the activities they carry out, the substantive theories that link these activities to change, and the levels of analysis being used has potential to improve the quality of theories of change. In the lead‑up to developing a new theory, it can be valuable to revisit earlier theories and identify their weaknesses, with the aim of learning from past mistakes and doing better.
A persistent challenge, however, is that evaluators are often brought into the process too late, when the theory is already fixed or the programme is underway. Feasibility and pilot studies, with an emphasis on implementation and process evaluation, offer a crucial opportunity to deepen theories of change, fail fast, and adjust course before public funding is committed to programmes with unconvincing rationales.
References
Chen, H. T. (2015). Practical program evaluation: Theory-driven evaluation and the integrated evaluation perspective (2nd edition). Sage Publications.
Fitz-Gibbon, C. T., & Morris, L. L. (1975). Theory-based evaluation. Evaluation Comment, 5(1), 1–4. Reprinted in Fitz-Gibbon, C. T., & Morris, L. L. (1996). Theory-based evaluation. Evaluation Practice, 17(2), 177–184.
Illari, P. M., & Williamson, J. (2011). What is a mechanism? Thinking about mechanisms across the sciences. European Journal for Philosophy of Science, 2(1), 119–135.
Ioannidis, S., & Psillos, S. (2018). Mechanisms in practice: A methodological approach. Journal of Evaluation in Clinical Practice, 24(5), 1177–1183.
Mareschal, D. (2016). The neuroscience of conceptual learning in science and mathematics. Current Opinion in Behavioral Sciences, 10, 114–118.
Scott, M., Scandone, B., Griggs, J., Roberts, E., Bristow, T., Woolfe, E., Dey, M., & Fugard, A. (2024). Basic Maths Premium evaluation report. Education Endowment Foundation.
Sun, R., Coward, L. A., & Zenzen, M. J. (2005). On levels of cognitive modeling. Philosophical Psychology, 18, 613–637.
Takala, H., Kuo, T.-L., Duysak, E., Bhatti, S., Stoilova, E., Fletcher, A., McGuinness, N., McKaskill, M., & Fugard A. (2025). Stop and Think: Learning Counterintuitive Concepts Evaluation Report. Education Endowment Foundation.