What should the is-ought thesis state?

Hume’s is-ought thesis states that we cannot infer a normative statement about what we should do from a descriptive statement about what is the case. This post spells out some of the detail about what the thesis means, which should be of interest to policy evaluators since we’re in the business of producing evidence and offering advice on policy that’s consistent with that evidence.

Let’s start with Prior’s (1960) puzzle (mildly edited): what kind of statement is “Either it’s raining or cricket should be banned”? It combines a descriptive statement (“it’s raining”) with a normative one (“cricket should be banned”), but is the disjunction of the two descriptive or normative?

Suppose Prior’s puzzle is a descriptive statement. Then by adding it to the premise “it’s not raining”, we can derive the purely normative conclusion that “cricket should be banned”. Since the disjunction is true, but one of its disjuncts is false, then the other disjunct must be true. We have just drawn an is-ought inference.

In symbols,

\(\displaystyle \neg R, R \lor \mathsf{O} B\ \models\ \mathsf{O} B\),

where \(R\) denotes “it’s raining”, \(\mathsf{O} B\) denotes that we ought to ban cricket (the \(\mathsf{O}\) is the “ought” from deontic logic), \(\lor\) is disjunction (or), \(\neg\) is negation, and \(\models\) is logical consequence.

This is maybe easier to see by rewriting the disjunction as a conditional, “If it’s not raining, then cricket should be banned” (since \(\neg \phi \lor \psi = \phi \rightarrow \psi\)):

\(\displaystyle \neg R, \neg R \rightarrow \mathsf{O} B\ \models\ \mathsf{O} B\),

where \(\rightarrow\) is the material conditional. The conclusion is then drawn by modus ponens. This pattern is interesting for evaluation, since it mirrors the logic:

  1. The programme works*.
  2. If the programme works*, we should roll it out nationally.
  3. Therefore, we should roll it out nationally.

Where works* includes a range of statements about impact and process evaluation evidence, whether that evidence can be generalised to a broader population, whether the programme avoids causing harm, its cost‑effectiveness, and other considerations.

In symbols,

\(\displaystyle W^*, W^* \rightarrow \mathsf{O} R\ \models\ \mathsf{O} R\),

where \(W^*\) denotes the programme works* and \(R\) denotes that we should roll it out.

Suppose instead that Prior’s puzzle is a normative statement. Starting from the purely descriptive premise “it’s raining”, we can infer “Either it’s raining or cricket should be banned” by disjunction introduction. Again, we have drawn an is-ought inference.

In symbols,

\(\displaystyle R\ \models\ R \lor \mathsf{O} B\).

Rewriting using implication,

\(\displaystyle R\ \models\ \neg R \rightarrow \mathsf{O} B\),

the conclusion is one of the “paradoxes” of the material conditional: true under classical logic because the antecedent is false, but people tend to judge the conditional to be neither true nor false but irrelevant (Johnson-Laird & Tagart, 1969).

The solution is, “Either it’s raining or cricket should be banned” is neither descriptive nor normative: it’s mixed. The same applies to “If the programme works*, we should roll it out nationally”. Three refinements of the is-ought thesis are provided in a logic-heavy book by Schurz (1997), which has been on my reading stack for years. A more digestible summary is provided by Schurz (2014, pp. 2–3):

(H1) No non-logically true purely normative conclusion can be derived from a consistent set of purely descriptive premises.

(H2) Every mixed conclusion [i.e., combining descriptive and normative statements] which follows logically from a set of purely descriptive premises is normatively irrelevant in the sense that all of its normative subformulas are replaceable by other arbitrary subformulas, while preserving the validity of the inference [“salva validitate of the inference” in Schurz’s original].

(H3) No non-tautologous descriptive statement can be inferred from a consistent set of purely normative premises.

H1 is the thesis that applies most to evaluation: to make an evaluative judgement, you need mixed premises that blend normative and descriptive statements. H2 deals with weird uses of logic. I’ve used the principle of irrelevance it contains to help understand how people reason about sentences like “If Alex posted the letter, then he posted the letter or set fire to the letter”, which are true in classical logic but which people often judge to be false (Fugard et al., 2011). There’s a blog post about it yonder. H3 says that knowing or believing what should be true doesn’t tell you what is factually true.

There is, however, a well‑known problem with using the material conditional in combination with oughts (Chisholm, 1963), e.g., \(W^* \rightarrow \mathsf{O} R\) used above. Consider the following sentences and formalisations:

  1. It ought to be that Jones goes to assist his neighbours: \(\mathsf{O} g\).
  2. It ought to be that if Jones goes, then he tells them he is coming: \(\mathsf{O} (g \rightarrow t)\).
  3. If Jones doesn’t go, then he ought not tell them he is coming. \(\neg g \rightarrow \mathsf{O} \neg t\).
  4. Jones doesn’t go: \(\neg g\).

The English‑language statements feel consistent with each other. They are also independent in the sense that no one sentence follows from any of the others. A formalisation should preserve both features.

There are three paths through, using what has come to be known as standard deontic logic (SDL). McNamara and Van De Putte (2025, Section 2.1) provides an introduction to SDL. Section 4.1 provides the illustration of Chisholm’s (1963) problem, which I’ve just spelt out a little. Here’s a summary of the paths:

Path APath BPath C
(1′) \(\mathsf{O}g\)
(2′) \(\mathsf{O}(g \rightarrow t)\)
(3′) \(\neg g \rightarrow \mathsf{O}\neg t\)
(4′) \(\neg g\)
(1′) \(\mathsf{O}g\)
(2′) \(\mathsf{O}(g \rightarrow t)\)
(3″) \(\mathsf{O}(\neg g \rightarrow \neg t)\)
(4′) \(\neg g\)
(1′) \(\mathsf{O}g\)
(2″) \(g \rightarrow \mathsf{O}t\)
(3′) \(\neg g \rightarrow \mathsf{O}\neg t\)
(4′) \(\neg g\)
From (1′), (2′), \(\mathsf{O}t\).
From (3′), (4′), \(\mathsf{O}\neg t\).
Consistency is lost.
(1′) implies (3″).
Independence is lost.
(4′) implies (2″).
Independence is lost.

Following Path A, we deduce that Jones ought to tell them he is coming and ought not tell them he is coming, which is suspect. This uses the following SDL rule:

\(\mathsf{O}(\varphi \rightarrow \psi) \rightarrow (\mathsf{O}\varphi \rightarrow \mathsf{O}\psi)\text{,} \tag{OB-K}\)

which gives us \(\mathsf{O}g \rightarrow \mathsf{O} t\) (from 2′). This (alongside 1′) gives us \(\mathsf{O}t\) by modus ponens. We can also get \(\mathsf{O}\neg t\) using modus ponens (3′ and 4′). It also contradicts one of the axioms of SDL: if you should do something, then you shouldn’t also do the negation of that something:

\(\mathsf{O} \phi \rightarrow \neg \mathsf{O}\neg \phi \tag{NC}\)

Paths B and C attempt to use the same expression for the conditional oughts in sentences 2 and 3. Path B uses

\(\mathsf{O}(\phi \rightarrow \psi)\).

Path C uses

\(\phi \rightarrow \mathsf{O}\psi\),

which we encountered above when exploring Prior’s (1960) puzzle.

The problem with both attempts to save SDL is that the sentences become dependent, whereas in the original informal English they aren’t. For Path B, \(\mathsf{O}(\neg g \rightarrow \neg t)\) is vacuously true (when rewritten using OB-K) because \(\mathsf{O}g\) is true – a paradox of the material conditional again. Similarly for Path C, \(g \rightarrow \mathsf{O}t\) is vacuously true because \(\neg g\) is.

So however we choose to formalise the sentence “if you should do A, then you should do B” or “if A, then you should do B”, it is not something that can be captured within SDL – a conclusion reached in the late 1960s (Parent & Torre, 2018, p. 20). This conclusion is familiar from work in the psychology of reasoning, where the material conditional is replaced with systems that behave more like everyday inference.

One approach uses defeasible logics, which allow us to retract a conclusion when new information arrives and to treat some premises as having greater priority than others (e.g., Neves et al., 2002). Another approach uses probability logics, which model reasoning under uncertainty (e.g., Pfeifer & Kleiter, 2009). Probability logics typically rely on an underlying three‑valued semantics, in which a statement can be true, false, or neither. The third value is usually interpreted as something like “irrelevant” or “undetermined”.

A recent review of deontic logic (McNamara & Van De Putte, 2025) concludes that “there are a number of outstanding problems for deontic logic. Some see this as a serious defect; others see it merely as a serious challenge, even an attractive one.” I’ve been reading attempts to solve some of these problems, e.g., Horty (2012), which applies defeasible logic to deontic reasoning. See the follow-up.

References

Chisholm, R. M. (1963). Contrary-to-duty imperatives and deontic logic. Analysis, 24, 33–36.

Fugard, A., Pfeifer, N., & Mayerhofer, B. (2011). Probabilistic theories of reasoning need pragmatics too: modulating relevance in uncertain conditionals. Journal of Pragmatics, 43, 2034–2042.

Horty, J. F. (2012). Reasons as defaults. Oxford University Press.

Johnson-Laird, P., & Tagart, J. (1969). How implication is understood. The American Journal of Psychology, 82, 367–373.

McNamara, P., & Van De Putte, F. (2025). Deontic logic. In E. N. Zalta & U. Nodelman (Eds), The Stanford encyclopedia of philosophy (Winter 2025). Metaphysics Research Lab, Stanford University.

Neves, R. D. S., Bonnefon, J.-F., & Raufaste, E. (2002). An Empirical Test of Patterns for Nonmonotonic Inference. Annals of Mathematics and Artificial Intelligence, 34, 107–130.

Parent, X., & Torre, L. van der. (2018). Introduction to Deontic Logic  and Normative Systems. College Publications.

Pfeifer, N., & Kleiter, G. D. (2009). Framing human inference by coherence based probability logic. Journal of Applied Logic, 7, 206–217.

Prior, A. N. (1960). The autonomy of ethics. Australasian Journal of Philosophy, 38(3), 199–206.

Schurz, G. (1997). The Is-Ought Problem: An Investigation in Philosophical Logic. Springer.

Schurz, G. (2014). Cognitive success: Instrumental justifications of normative systems of reasoning. Frontiers in Psychology, 5(625).

ggauto

ggauto is designed to choose the best chart type, based on the type of data that you have. That means that you need to pre-process your data into the correct type before plotting it. This is a requirement to make ggauto possible, but is more generally a good idea because it forces you to understand what your data is before you plot it.”

Looks cool. Blog post here.

Install from CRAN: install.packages(“ggauto”)

The value of deferring strong evaluative judgement

This post continues my ruminations on values and evaluative judgements, and considers whether the scarcity of explicit evaluative judgements in reports is really a bad thing. It’s thinking in progress.

The story so far… Evaluation is frequently defined as “the process of determining the merit or worth” of things (Scriven, 1994, p. 152). However, a review of 13 broad evaluation approaches identified only three that provided any guidance on how to make these judgements (Schröter et al., 2026), and this is reflected in practice. A gallon (95% CI 0.8 to 1.2) of ink has been spilled arguing that we need more explicit evaluative judgements.

The argument, roughly and vastly oversimplifying, goes like this: in light of the is-ought gap, evaluative judgements require values. We should make those values explicit and blend them with the (theory‑laden) facts that the evaluation yields to obtain an explicit (value‑laden) judgement. Rubrics are one way to record agreed values and, its proponents argue, help us to reach an evaluative conclusion (King et al., 2013).

Three (value‑laden) facts trouble me about calls for evaluators to provide more explicitly argued evaluative judgements.

Firstly, arguments are often enthymemes: they rely on premises that are left implicit. This is well studied in philosophy, linguistics, and the psychology of reasoning (my PhD was on the latter). It would be extremely difficult to communicate at all if we had to spell out every premise, and it turns out that we are often very good at filling in the gaps. Grice, among others, described a set of conversational conventions that we seem to follow and that help us do this. For instance, if someone says that “Jane ate some of the ice cream”, we typically conclude that she did not eat all of it, even though that conclusion does not follow from classical logic. If she had eaten all of it, then to comply with one of Grice’s principles we would say so.

People draw similar enthymematic inferences when moving from an is to an ought. To take an easy example, from premises like “If you pull the dog’s tail again, then he’ll bite you”, people conclude “You should not pull the dog’s tail”, apparently implicitly inferring a bridging premise that fills in the is-ought gap (Elqayam et al., 2015). A key driver of these inferences is that descriptive sentences are often value-laden, e.g., being bitten by a dog is judged to be bad. Findings from evaluations are usually value-laden too, concerning outcomes such as improved health or educational attainment. However, logicians and philosophers strive to make any is-ought gaps explicit in their analyses.

Secondly, systematic reviews are frequently required to assess evidence, and what you tend to find is that it takes time for patterns in findings to be identified. There is variation in study quality, particularly in the earliest evaluations of a programme. Over time, moderators of change are identified – if you’re lucky, and enough studies have been conducted in sufficiently many contexts. There is also the issue of publication bias, which can take a while to detect, and early optimism about effect sizes is often attenuated. All of this means it can be unwise to rely too heavily on a single evaluation. Yet, in practice, evaluators are not funded to conduct a systematic review once they have finished evaluating a programme in a single context.

Thirdly, there is obviously vast diversity in the values people hold. For example, a YouGov poll a few years back found that around half of Conservative and Leave voters thought the British Empire was something to be proud of, and that former colonies were better off for having been colonised. Roughly 40 percent of them said they would like Britain still to have an empire! Among Labour and Remain voters, only about 20 percent expressed pro‑empire views. Listen to callers on LBC Radio and you will hear equally varied views on, e.g., racism and immigration, and consequent policy suggestions.

Diversity of values applies across a wide range of policy areas that have a huge impact on people’s lives, and that many of us do or will work on. Consider, for example, contemporary debates on holding people in immigration removal centres; the activities of big tech companies; welfare benefit policy and conditionality; how transgender people are treated; or the experiences of people in mental health inpatient units.

Mabry (2010, p. 84) summarises the problem faced by evaluators co-creating rubrics:

“Those who promote attention to the values of stakeholders beyond those of program managers or funders […] often refer optimistically to the importance of building consensus. But the diversity of stakeholder interests may be irreconcilable, and evaluation’s capacity to clarify differences may cement dissensus. Moreover, […] every taxpayer, every citizen, every resident is a remote stakeholder, introducing a diversity of social values that could overwhelm an evaluation.”

Of course, the fact that it is often challenging to reach consensus on values does not mean that we shouldn’t try to do so. Alternatively, we may have to draw more than one evaluative judgement depending on whose values are added to the premises of the argument, or prioritise the values of service users. However, as Mabry’s argument suggests, we may often be stepping into territory that an evaluator’s judgement alone cannot settle.

Regardless of how explicit we are about value judgements, evaluation – like the rest of science – is value‑laden (Ward, 2026). Values shape which evaluations are commissioned, how they are designed, what outcomes get measured and ignored, and how findings are used. But I wonder whether the scarcity of explicit evaluative judgements in evaluation reports can be explained by taking seriously the reality of policymaking, and who uses or could potentially use the findings from an evaluation.

One way to understand this reality is to look at how policy actually gets made. For example, Sabatier (1988) argues that, rather than sitting neatly within a single department, most policymaking depends on shifting coalitions of actors. Adapting his argument for the UK context, this would include ministers and civil servants, opposition parties, local authorities (both elected members and officers), regulators, professional bodies, charities, think tanks, campaigners, academics, journalists, and service providers. Each brings their own priorities and values, and each has differing levels of power – and therefore differing influence – over how evidence is used.

Evaluations could offer more value by explicitly considering this diverse range of potential users of the findings, beyond the funder. When perusing Hansard, I’m always delighted to see an opposition MP asking when an evaluation is due to be published. Delighted, not always because I share their likely conclusions, but because it shows that evaluations are being used. I am unsure where, within the wider mix of actors described above, evaluation as a profession should position itself. Perhaps we shouldn’t be ashamed of delegating some evaluative judgements to others, provided we ensure that we supply the evidence needed to support those judgements.

Sometimes the values at stake, though left implicit in conclusions, are obvious, e.g., if an evaluation concludes that a programme reduces people’s risk of suicide. In other cases they are less so, and the policy landscape is marked by deep value clashes. Consider, for example, policies concerning asylum seekers or transgender people. Some policy options will cross a threshold that evaluators cannot ignore, leaving us ethically compelled to spell out the values – particularly the value clashes between policymakers and those most directly affected by policies.

Revised 30 March 2026

References

Elqayam, S., Thompson, V. A., Wilkinson, M. R., Evans, J. St. B. T., & Over, D. E. (2015). Deontic introduction: A theory of inference from is to ought. Journal of Experimental Psychology: Learning, Memory, and Cognition, 41, 1516–1532.

King, J., McKegg, K., Oakden, J., & Wehipeihana, N. (2013). Evaluative rubrics: a method for surfacing values and improving the credibility of evaluation. Journal of MultiDisciplinary Evaluation, 9, 11–20.

Mabry, L. (2010). Critical social theory evaluation: Slaying the dragon. New Directions for Evaluation, 127, 83–98.

Sabatier, P. A. (1988). An advocacy coalition framework of policy change and the role of policy-oriented learning therein. Policy Sciences, 21, 129–168.

Schröter, D., Becho, L. W., & Montrosse-Moorhead, B. (2026). The garden of evaluation approaches: Supporting explicit, theory-informed evaluation practice. Evaluation.

Scriven, M. (1994). Evaluation as a discipline. Studies in Educational Evaluation, 20(1), 147–166.

Ward, Z. B. (2026). What does it mean to say that science is value-laden? In K. C. Elliott & T. Richards, The Routledge Handbook of Values and Science (pp. 74–83). Routledge.

Whose values of merit or worth…?

Some helpful thoughts from Linda Mabry (2010) on value clashes in evaluative judgements – perhaps what we could call the fundamental problem of evaluative judgement, the value-based sibling of the fundamental problem of causal inference:

“Those who promote attention to the values of stakeholders beyond those of program managers or funders […] often refer optimistically to the importance of building consensus. But the diversity of stakeholder interests may be irreconcilable, and evaluation’s capacity to clarify differences may cement dissensus. Moreover, for federally funded programs, […] every taxpayer, every citizen, every resident is a remote stakeholder, introducing a diversity of social values that could overwhelm an evaluation.” (p. 84)

“[…] should decision-makers favor the evaluator’s values over those of program personnel and other stakeholders? For a critical social theory evaluator, this issue demands introspection, a willingness to interrogate one’s own conception of appropriate use of the evaluation. Theoretically at least, it is as possible for evaluators to misunderstand appropriate use as it is for clients to do so. While the client can count on lived experience of the program to guide ideas about appropriate use, he or she is invested in the personal values reflected in the program; while an external evaluator can count on fresh eyes and systematically collected data, he or she is invested in the personal values reflected in the evaluation. There being no disinterested view, notions of appropriate use always involve someone’s subjective values.” (pp. 90-91)

“[…] the evaluator cannot simply presume it appropriate for his or her conclusions or values to overrule those of decision-makers and other stakeholders. Errors of two types are possible. On one hand, clients might be well advised to exercise healthy skepticism in considering the work of a short-term outsider, one whose findings might point them toward unproductive territories. On the other hand, stakeholders pinched by evaluation results have been known to engage in blatant self-protection […].” (p. 91)

Mabry, L. (2010). Critical social theory evaluation: Slaying the dragon. New Directions for Evaluation, 127, 83–98.

The value-ladenness of science, and what it means for evaluation

There’s a long history of work arguing that science is value‑laden and involves evaluative judgements. But what sorts of values and evaluative thinking are involved? I read an analysis by Zina Ward (2026) to get a sense of the latest thinking, and pondered what it might mean for our discipline of evaluation.

Ward reminds us that value‑ladenness is obvious in science. For instance, more research funding is devoted to understanding and curing diseases in humans than in koalas. A recent example is the UK’s £2 billion investment in quantum research. This investment is motivated by anticipated applications such as secure communication, faster algorithms for scientific problems, and advanced sensing. Each of these reflects underlying values: that communication should be secure; that certain scientific problems are worth prioritising; and that military capabilities, such as detecting submarines that evade current technologies, should be strengthened. Ethical values also play a role in shaping the sorts of research that is conducted.

There are at least four different ways that a choice is value-laden, according to Ward. Value-ladenness can be:

  • rational,
  • motivational,
  • causal, or
  • objectual.

Choices that are value-laden in the rational sense provide the justification for a choice. For example, we may want to promote certain kinds of research given priorities in society and some overall ideology. What actually motivates a scientist to conduct research (motivational value-ladenness) may also align with these rational values; however, scientists often conduct research for personal reasons. That has been the case for scientists focusing on Covid-related research (e.g., they lost a loved one to Covid or know someone with long Covid) or working on trans-inclusive theories of gender (e.g., they are trans or have a loved one who is trans). Outside the realm of science, Ward gives the example of someone who cites reducing their carbon footprint and reducing animal suffering as justifications for becoming vegetarian (rational), whereas in reality they did so to fit in with their vegetarian friends (motivational, but also rational if used as a justification).

A choice is causally value-laden if values influence the choices someone makes. For example, the code of ethics for a profession such as psychology constrains the sorts of research that can be conducted. This may also be an example of motivational value-ladenness – psychologists want to conduct research that aligns with these ethical frameworks, and may have been involved in developing them. However, some scientists would be motivated to conduct research that is unethical without the causal constraints – there are plenty of examples of this in history.

Finally, a choice, or an evaluation of a potential choice, can be objectually value-laden. This concerns the impact that the choice has on the world, in relation to values. Choosing to fund more human than koala health studies is an example. Another would be reforming private family law, e.g., through the recently announced national rollout of the Child Focused Model. These choices cannot be made based on facts alone, as the is-ought problem reminds us.

When evaluation is defined as “the process of determining the merit or worth” of things (Scriven, 1994, p. 152), it reflects this latter objectual value-ladenness. Given this definition, it is striking that a recent review of 13 broad evaluation approaches identified only seven that considered judgements of merit or worth “essential”, and only three provided any guidance on how to make these judgements (Schröter et al., 2026). This gap between definition and practice requires some consideration; however, I am still formulating my views on this. Where I’ve got to can be summarised in the following two paragraphs.

Firstly, objectual evaluations are pervasive across science and policy making – evaluation, as the field operates in practice, does not have a monopoly on evaluative thinking. Now it could be that there is a need for a transdisciplinary genre of evaluation (see, e.g., Scriven, 2008) – something similar to logic, psychometrics, or statistics. Conferences for such a discipline would invite anyone who reasons about objectual values, whether they be scientists, educators, vegetarians, or anyone else.

I’d conjecture instead that the current field of evaluation that we know and love is really focused on policy evaluation, and policy evaluation concerns a range of activities other than evaluative thinking. It’s about conducting social research on policies and programmes. One important aspect of policy evaluation is reasoning about objectual values – just as it is across the sciences. But there is more to policy evaluation than this, for example how to develop theories of change, apply participatory approaches in evaluation design, understand methods that can be used to test theories of change, and a huge number of practical considerations involved when conducting research on policy at scale. Finally, I’d venture the conjecture that evaluation in the transdisciplinary sense already exists. It just has a different name and lives somewhere in departments of philosophy and/or politics.

References

Schröter, D., Becho, L. W., & Montrosse-Moorhead, B. (2026). The garden of evaluation approaches: Supporting explicit, theory-informed evaluation practice. Evaluation.

Scriven, M. (1994). Evaluation as a discipline. Studies in Educational Evaluation, 20(1), 147–166.

Scriven, M. (2008). The concept of a transdiscipline: And of evaluation as a transdiscipline. Journal of MultiDisciplinary Evaluation, 5(10), 65–66.

Ward, Z. B. (2026). What does it mean to say that science is value-laden? In K. C. Elliott & T. Richards, The Routledge Handbook of Values and Science (pp. 74–83). Routledge.

Deepening the theories of mechanisms used in policy evaluation

Theory‑based (or driven) evaluations begin with a theory of change and then design the evaluation to test the causal mechanisms it proposes (Fitz‑Gibbon & Morris, 1975). A well‑constructed theory of change sets out how a programme’s resources are deployed to support delivery activities, and how these activities activate the causal mechanisms that generate outcomes as they evolve over time. The early – and, to my mind, more coherent – formulations of theory‑based evaluation were pluralistic, encompassing all methodological approaches (see, e.g., Chen, 2015): RCTs and quasi‑experiments alongside, for example, process tracing and qualitative comparative analysis.

Mechanisms can be defined in terms of entities and what they do to bring about change (Illari & Williamson, 2011). We can identify those entities and activities with the aid of substantive theories relevant to the policy area under investigation, and the mechanisms are described using the concepts provided by those theories (Ioannidis & Psillos, 2018). A substantive theory is particularly helpful when there is evidence for the mechanisms it proposes, rather than when it rests only on armchair speculation. However, all theories are necessarily incomplete and evolve as new evidence accumulates. Some theories are more detailed than others. Some have been more thoroughly tested than others.

Substantive theories describe mechanism at different levels of explanation. Sun et al. (2005) describe the levels as follows:

Object of analysisType of analysisElements in model
Inter-agentSocial/culturalCollections of agents
AgentsPsychologicalIndividual agents
Intra-agentComponential [I think also psychological]Modular construction of agents
SubstratesPhysiologicalBiological realisation of modules

For an inter‑agent analysis, the focus is on how collections of agents (often people) interact with one another. Examples include models of crowd behaviour when a fire alarm sounds, or systemic processes that shape people’s experiences, such as racism, sexism, transphobia, and their intersections.

For an agent‑level analysis, the focus shifts to individual people. This is roughly the way we talk about individuals in everyday life, including their desires, beliefs, opportunities, and actions, for instance, spending money, talking, listening, going for a run, attending a mentoring session, or doing homework.

An intra‑agent analysis draws on concepts from psychology. This may include theories of cognitive control; for instance, why people sometimes respond automatically and at other times engage in deliberate thought, and how conflicts between competing cognitive systems are resolved. It may include theories of emotion, such as how people’s goals and their (often automatic) inferences about progress toward those goals give rise to feelings of happiness, sadness, fear, or anger. Theories of memory systems also sit here, including how information is temporarily represented in visuospatial and phonological working memory, and the capacity and processing limits of these systems.

Finally, analysis at the substrate level considers the biological processes that implement psychological‑level mechanisms or reflect the consequences of behavioural change. This is most visible in dietary interventions, where targets might include blood glucose or cholesterol levels. It can also include theories from cognitive neuroscience, such as which neural systems underpin cognitive control or memory.

I revisited the theories of change in two evaluations I worked on to see which elements from these four levels of analysis were present.

Case study 1: Basic Maths Premium

Basic Maths Premium was a Department for Education pilot that provided additional funding to post‑16 providers in disadvantaged areas to improve GCSE maths resit outcomes for students with prior attainment at grade 3 or below. It tested three funding models using an RCT: two with guaranteed elements and one based entirely on payment by results (PbR):

ModelUnconditional fundingConditional
funding
A£500 × number of eligible students
B£250 × number of eligible students£250 × number of successful students
C£500 × number successful students

The theory of change was at the agent level. For unconditional funding, the theory was straightforward: funding would enable activities expected to improve outcomes, such as more teaching hours, smaller class sizes, and greater use of technology. By contrast, the theory underpinning conditional funding was less unconvincing. Financial incentives were assumed to boost staff motivation, which would in turn enhance teaching quality, student motivation, and ultimately learning outcomes. However, for Model C in particular, it was uncertain how providers were expected to finance any additional activities, given that all funding under this model depended on student pass rates after any teaching had ended. Two quotations from heads of maths, drawn from the implementation and process evaluation (IPE), captured the issue succinctly (Scott et al., 2024):

  • “The fact that we’ve got this potential funding in the future that will reward us for that, that’s great, but we’ve still got to find the funds now to do what we do” (Model C – conditional £500 per passing student).
  • “I think [payment by results is] really unfair because you would not really know how much money you were going to get. I don’t think we would have been able to spend any additional money on that basis, so for us it wouldn’t have really worked. […] I would never have been able to employ two staff on the basis that I might get a certain amount of students through a GCSE. That would have been too much of a financial risk” (Model A – unconditional £500 per student). Also a good example of counterfactual reasoning without a comparison group.

The IPE showed that actual spend was primarily driven by the unconditional funding:

Case study 2: Stop and Think

Stop and Think is a computer‑based intervention designed to help primary pupils overcome common misconceptions in maths and science. It focuses on strengthening children’s inhibitory control: the ability to pause, reflect, and override intuitive but incorrect responses, by guiding them through short, game‑like activities that present counter-intuitive problems.

The theory underlying the programme included intra-agent and substrate levels of analysis. Briefly, the idea is that when students learn new concepts, they need to overcome intuitively obvious prior beliefs. Mareschal (2016) summarises evidence of the cognitive mechanisms involved when intuitions and new learning clash with each other, e.g., the inhibition of pre-existing beliefs involves processes implemented in the dorsal lateral prefrontal cortex (DLPFC) and the anterior cingulate cortex (ACC). Picture below (Mareschal, 2016, p. 115):

One of the tests of the theory in the RCT evaluating Stop and Think (Takala, et al., 2025) involved the construction of a measure of misconceptions driven by prior beliefs. The impact of the programme on misconceptions was then tested quantitatively using mediation analysis.

Importantly, this variable-based analysis is not the mechanism. Instead, the misconceptions measure was an indirect operationalisation of a trace of the underlying mechanism illustrated in the picture above. This distinction is summarised by Paley and Lilford (2011, pp. 956-7):

“The view […] that reality is fragmented into variables, is a straw man. No one believes it, and classification into types is something that all researchers, both quantitative and qualitative, do. Variables are a product of measurement procedures; they are not part of the structure of reality.”

Conclusions

Being explicit about the entities involved, the activities they carry out, the substantive theories that link these activities to change, and the levels of analysis being used has potential to improve the quality of theories of change. In the lead‑up to developing a new theory, it can be valuable to revisit earlier theories and identify their weaknesses, with the aim of learning from past mistakes and doing better.

A persistent challenge, however, is that evaluators are often brought into the process too late, when the theory is already fixed or the programme is underway. Feasibility and pilot studies, with an emphasis on implementation and process evaluation, offer a crucial opportunity to deepen theories of change, fail fast, and adjust course before public funding is committed to programmes with unconvincing rationales.

References

Chen, H. T. (2015). Practical program evaluation: Theory-driven evaluation and the integrated evaluation perspective (2nd edition). Sage Publications.

Fitz-Gibbon, C. T., & Morris, L. L. (1975). Theory-based evaluation. Evaluation Comment, 5(1), 1–4. Reprinted in Fitz-Gibbon, C. T., & Morris, L. L. (1996). Theory-based evaluation. Evaluation Practice, 17(2), 177–184.

Illari, P. M., & Williamson, J. (2011). What is a mechanism? Thinking about mechanisms across the sciences. European Journal for Philosophy of Science, 2(1), 119–135.

Ioannidis, S., & Psillos, S. (2018). Mechanisms in practice: A methodological approach. Journal of Evaluation in Clinical Practice, 24(5), 1177–1183.

Mareschal, D. (2016). The neuroscience of conceptual learning in science and mathematics. Current Opinion in Behavioral Sciences, 10, 114–118.

Scott, M., Scandone, B., Griggs, J., Roberts, E., Bristow, T., Woolfe, E., Dey, M., & Fugard, A. (2024). Basic Maths Premium evaluation report. Education Endowment Foundation.

Sun, R., Coward, L. A., & Zenzen, M. J. (2005). On levels of cognitive modeling. Philosophical Psychology, 18, 613–637.

Takala, H., Kuo, T.-L., Duysak, E., Bhatti, S., Stoilova, E., Fletcher, A., McGuinness, N., McKaskill, M., & Fugard A. (2025). Stop and Think: Learning Counterintuitive Concepts Evaluation Report. Education Endowment Foundation.

Evaluation as a social science

“We have argued that good evaluation is good social science. For us, this embraces the gallant aims of precision in articulation of theory, rigor in empirical testing, confederation in lines of inquiry, and cumulation in the body of findings. The ‘realist movement,’ of which we are a part, is often considered the brash upstart of the evaluation schools. In fact, it depends on these rather venerable ideas. The future, for us, thus lies in keeping faith with some of the grand old principles of social science and in not forgetting the hard-won lessons of the old studies.” (Pawson & Tilley, 2001, p. 324)

“All scientific investigation utilises explanations relating mechanisms and contexts to empirical patterns.” (Pawson, 2024, p. 42)

“Despite the habitual use of the appellation ‘social science’, many of my colleagues would reject any claim to follow science, dismiss any interest in causality, deny any need for objectivity, and scorn the possibility of generalisation. They are beyond hope. I don’t seek to convert them. But in following their chosen paths these various tribes – constructivists, post-modernists, emancipators, critics, essayists, relativists, and so on – have found time to say why causality, objectivity, and generality are false idols. So, in defending the science in social science, their criticisms also need to be overturned.” (Pawson, 2024, p. xviii)

“What I’ve come up with here might well be entitled The Old Rules of Sociological Method. I have attempted to extract and justify some realist principles for conducting social research on the back of a generous portfolio of existing examples. Those illustrations reach across many research domains and a broad portfolio of practical methods. But they remain a pinprick; I could have called upon a thousand others. Accordingly, there is another way of perceiving my efforts. The book is no more and no less than an attempt to codify and formalise existing practices. I have tried to capture a tradition.” (Pawson, 2024, 251)

References

Pawson, R., & Tilley, N. (2001). Realistic Evaluation Bloodlines. American Journal of Evaluation, 22(3), 317–324.

Pawson, R. (2024). How to Think Like a realist: A methodology for social science. Edward Elgar Publishing Limited.

The bargraphs of evaluation approaches

Scroll down for a redrawing of Schröter, Becho, and Montrosse-Moorhead’s (2026) characterisation of broad evaluation approaches. R code here.

Their assessments of each approach on 8 dimensions are linked here:

References

Schröter, D., Becho, L. W., & Montrosse-Moorhead, B. (2026). The garden of evaluation approaches: Supporting explicit, theory-informed evaluation practice. Evaluation.

Master’s in Public Policy (MPP) talk at Cambridge

This week I gave a talk to the MPP class. Slides here. I covered the usual things, e.g., how all policy evaluation is theory-based and counterfactuals don’t need a comparison group. A few new additions:

  • I claimed that Realistic Evaluation, i.e., the (Pawson & Tilley, 1997) book, is an intro to scientific methods and philosophy of science, rather than introducing a distinct form of evaluation. I think this is clearest to see in Pawson (2024), and it’s one of the most muddled sections of the Magenta Book.
  • The importance of theory for interpreting the meaning of covariates as well as merely selecting them. This is particularly important for intersectional understandings of selection into programmes, e.g., you need a theory to work out why, e.g., some combinations of gender and ethnicity, are more likely to be referred to a targetted programme. Is it driven by gender and/or racial stereotypes at time of selection, or broader societal discrimination before selection, for example?
  • An example from the Green Book (para 4.17) of a type of counterfactual reasoning, used when appraising different ways to realise a policy. I think the reasoning behind “deadweight” and “additionality” outcomes may use future less vivid conditionals. Still reading.