Resignations

Interesting dataset from the Institute for Government. (I last updated this after Streeting resigned.)

Here’s a cumulative count of resignations that were coded by IfG as being due to “disagreement”, for a selection of PMs. The dashed horizontal line shows Starmer’s count:

Expanding out to all PMs in the data (from Thatcher to Sunak), how long did they have left as the number of resignations grew? (Note: different colours – also remember the small number of PMs!)

Another attempt to visualise (again, remember small n – the lines give a range not uncertainty!). Purple marks where Starmer is currently at, though the data excludes him since he’s still PM.

Counterfactual without a comparison group

“Following the approval of Bill 56, Building a More Competitive Economy Act, 2025, all automated speed enforcement cameras were deactivated on November 13, 2025. The city has since been collecting and analyzing speed data at the original eight pilot locations to evaluate the impact of the camera removals. Monthly updates are provided for each site, including average speed, 85th percentile speed, speed-limit compliance rates, and the proportion of high-end speeders. This data is presented alongside baseline conditions prior to ASE camera installation and results observed during the enforcement period.”

Data here (hat tip @sjamieit.bsky.social). Here’s a pic, drawn using ggplot in R.

Value free governments

“[… I]t is accepted that the social scientists’ contribution cannot be value free but, surprisingly, much less attention has been given to the much more obvious fact that governments are not value free either. The fact that they are not means that the problem facing the social scientist in government is not so much that his own value system colours his research and recommendations, but that these values may be out of tune with those of the government he is advising.” (Sharpe, 1976, pp. 75–76)

References

Sharpe, L. J. (1976). Government as Clients for Social Science Research. Zeitschrift Für Soziologie, 5(1), 70–79.

Useful science (and evaluation?) for policy

“A policy problem is not usually the same as a scientific problem, and may have several scientific problems incorporated within it.”

“Scientists can be advocates, or they can provide the best possible balanced assessment of the evidence but they cannot do both simultaneously. It has to be clear to policymakers which horse they are riding. Papers seen as advocacy are likely to be discounted.”

“Since the policy process tends to be very fast, papers must be timely. An 80% right paper before a policy decision is made it is worth ten 95% right papers afterwards, provided the methodological limitations imposed by doing it fast are made clear.”

“Sensible policymakers prefer a paper they understand, including its flaws, to one they do not, however sophisticated and apparently precise it looks. It is possible to be simple whilst being rigorous.”

“Policymaking is a professional skill; most scientists have no experience of it and it shows. […] Worse, trying to work up to a policy position can unconsciously bias scientists towards trying to get a neat policy narrative from a complex picture, or downplay inconvenient facts.”

“Some scientists seem to assume a two-stage process where the individual research is conducted, and then policies made on the basis of that. The reality should be a three-stage process: that original research is conducted, then research from multiple instances and disciplines is synthesised, and then policies made on the basis of all the available synthesized evidence.”

References

Whitty, C. J. M. (2015). What makes an academic paper useful for health policy? BMC Medicine, 13(1), 301, s12916-015-0544–0548. https://doi.org/10.1186/s12916-015-0544-8

The kind of researcher needed for policy research and evaluation (Donnison, 1972)

David Donnison’s (1972, p. 532) view on the qualities required for someone to be an effective policy researcher or evaluator:

“The research workers to look for should of course be the most creative and intelligent we can find. But their motivation is equally important and harder to assess. They must have a policy-oriented cast of mind, or be likely to acquire that in time. They must enjoy thinking in a sustained, rigorous and independent way about the problems of government – which means they take seriously the version of a problem perceived by the government of the moment, but do not confine their own perception of the issues to that version alone. They must want to make the world a better place for people already living in it: the needs of future generations should not be wholly neglected, but neither should they be the main concern. And they must enjoy studying how that could be achieved with the resources potentially available for the task – which calls for a capacity to envisage the alternative uses to which such resources might be put. […]”

“Responsible scholars recognise that if they are to present conclusions which might ultimately affect many of their fellow-citizens they must be more, not less, careful about the quality of their evidence and the logic of their argument than they would be if they intended only to communicate with their academic colleagues. […] They should not be tempted to become politicians – a reputable but entirely different role. They should remain firmly rooted in the academic world – whether they work in universities or not – where they can keep in touch with their colleagues and students and draw on their help.”

Donnison, D. (1972). Research for policy. Minerva, 10(4), 519–536.

P.S. I’m not convinced by the comment on future generations – perinatal programmes and net zero policies are example areas where evaluators should and do consider impacts on future generations.

“It is not unheard of for permanent secretaries, in a sense, to try and work in a way which allows there to be a decision”

Edward Morello: Well, we should put on record that you have served under four different Prime Ministers. At any point during your tenure as perm sec, were you asked by a No. 10 organisation to withhold information from the Foreign Secretary? How usual is that request?

Sir Philip Barton: I am worried that everyone is going to think that the centre of Government spends its whole time sort of conniving behind the backs of everyone else in Government. That is not the reality. The reality is people trying to get on with delivering the business of the Government the vast majority of the time. My best answer to that question is that it is unusual. However, in the end, it is particularly around things where you might have a policy disagreement or difference of view between the Prime Minister and the Secretary of State.

It is not unheard of for permanent secretaries, in a sense, to try and work in a way which allows there to be a decision, and a consensus view in the Government to then move ahead and take it forward. In that sort of situation, it is not unheard of for a permanent secretary to be privy to something that they are asked not to pass on to their Secretary of State. I would describe it as not unheard of, but I do not want to give the impression that this is standard operating procedure.

Q344 Chair (Emily Thornberry): But this is absolutely extraordinary. The permanent secretary at the Foreign Office can be told, “Don’t tell your boss about this” when they are the Foreign Secretary? Has that ever happened to you before this?

Sir Philip Barton: Yes, it has.

Q345 Chair: Really?

Sir Philip Barton: Yes.

Chair: Well, you learn something new every day.

(Oral evidence: Work of the Foreign, Commonwealth and Development Office, HC 385, Tuesday 28 April 2026)

Starmer vs Putin approval ratings

Two interesting opinion polls, in the genre of value-laden evidence. Wearing your evaluative thinking spectacles, what do you make of them?

What normative values informed your judgements? For example, you probably wouldn’t want to be just comparing the metrics!

What should the ratings be?

Starmer

Starmer’s favourability tracker (YouGov)

Putin

Putin’s approval rating (Levada)

Applying a deontic logic to policy evaluation

One of the aims of policy evaluation is to support policymakers in making decisions. In practice, policymaking is shaped by coalitions of actors (Sabatier, 1988), including ministers and civil servants, opposition parties, professional bodies, charities, think tanks, campaigners, academics and journalists. Evaluation should aim to inform the deliberations of all these actors, not only those of whoever commissioned the work – regardless of whether it is remotely feasible to include them in a participatory process as part of the evaluation. Since each actor may hold normative values that clash with others’ values (understatement of the century), they may draw different conclusions about what should be done even when presented with the same evidence.

If the evidence produced by evaluators is to be relevant to a range of actors, we potentially need to reason about conflicting values. Deontic logics, which make it possible to reason about oughts, provide one way of structuring this reasoning. Given the complexities of applying deontic logics formally, as we will see shortly, and the large number of normative values involved, it is unlikely that evaluators will use them in a fully formal way. However, just as process tracing is informed by Bayesian logic (Bennett, 2009), I am curious whether informal reasoning about values can be informed by deontic logics. This post is a first go to find out.

There are many systems of deontic logic. A previous post ruled out standard deontic logic (SDL), which, confusingly, ceased to be standard in the late 1960s (Parent & Torre, 2018, p. 20). An alternative I’d like to explore is Horty’s (2012) approach, a prioritised default logic, which has the following key characteristics:

  1. Classical logic is monotonic, in the sense that adding more premises to an argument can never lead to the retraction of a conclusion: either the set of conclusions stays the same or grows (note the parallel with monotonic functions). Horty’s logic is nonmonotonic, meaning that it does allow conclusions to be withdrawn. The classic example concerns a bird named Tweety. All birds fly, so we conclude that Tweety flies. However, if we subsequently learn that Tweety is a penguin, we revise that conclusion and conclude that Tweety doesn’t fly.
  2. Horty’s nonmonotonic logic is implemented using default rules (defaults for short), written as \(\phi \rightarrow \psi\), which are read as: if we have established \(\phi\), then we should conclude \(\psi\) by default (for example, that birds fly). This inference can be overruled if another default supports a contradictory conclusion (for example, that penguins do not fly).
  3. Defaults represent reasons for believing things (e.g., Tweety can’t fly because Tweety is a Penguin) or reasons for doing things (e.g., meeting a friend for lunch because we promised to).
  4. Some default rules have higher priority than others, e.g., if \(\delta_1\) and \(\delta_2\) are two defaults, then \(\delta_1 > \delta_2\) means that \(\delta_1\) has higher priority than \(\delta_2\), so can overrule it. This might be due to the rule being more specific, e.g., if we’re reasoning about penguins we should prioritise defaults about penguins rather than about birds more generally. When defaults refer to normative values, the priorities determine which values are more important than others.
  5. The system therefore makes it possible to reason about moral conflicts. One of Horty’s examples involves a promise to meet a friend for lunch and a moral obligation to save a drowning child encountered en route to lunch. A reasonable ordering on defaults representing these norms is that saving a drowning child takes priority over fulfilling a lunch promise, so the logic concludes that the child should be rescued.

To illustrate how this works, let’s explore the following simplified example:

  • An evaluation has shown that a programme leads to an outcome valued by a policymaker: \(P\).
  • The same evaluation has shown an unanticipated harmful consequence of the programme that was not considered by policymakers or evaluators at the outset, but was highlighted by service users: \(H\).
  • If the valued outcome is found, then we should roll out the programme nationally: \(d_1 = P \rightarrow R\).
  • If harmful outcomes are found, then it should not be rolled out nationally: \(d_2 = H \rightarrow \neg R\).

We could complicate this further by introducing differences in the quality of evidence. For example, \(P\) might have been established through a rigorous impact evaluation, whereas \(H\) might have been identified through qualitative interviews with a small sample – a common way in which unintended consequences are discovered. However, to keep the discussion simple, let’s assume that there is no difference in the quality of evidence.

How do we evaluate evidence using Horty’s prioritised default logic? Since it is a logic, the rules are all formally defined, and applying them in practice can be challenging. Horty’s (2012) explanation spans several chapters. Here follows a concise summary; refer to the original text for fuller explanation and worked examples.

A default theory consists of three elements, which for the example above would be:

  1. \(\mathcal{W} = \{ P, H \}\) is the starting point for our inferences, representing the evidence found by the evaluation and any background assumptions. This would include relationships between the variables, which we don’t have in our simplified policy example. For the Tweety example, it would include that penguins are birds.
  2. \(\mathcal{D} = \{ d_1, d_2 \}\) is the set of defaults described above. For a default \(\phi \rightarrow \psi\), \(\phi\) is the premise and \(\psi\) the conclusion. We can refer to the premises and conclusions of a default or set of defaults, \(\Delta\), using \(\text{premise}(\Delta)\) and \(\text{conclusion}(\Delta)\).
  3. \(<\) is the ordering on the defaults, which specifies which take precedence over others. This is a partial ordering, in the sense that some or all defaults may have the same precedence. Let’s begin with no ordering, so the defaults concerning desired (\(d_1\)) and harmful (\(d_2\)) outcomes have equal importance.

To draw inferences, we need to consider one or more scenarios: these are a subset of the defaults, \(S \subseteq \mathcal{D}\). A scenario could include some, all, or none of the defaults. Binding defaults are defaults in the full theory such that the following conditions hold for a scenario:

  1. The default \(\delta \in \mathcal{D}\) is triggered, in the sense that its premise logically follows from the background theory and conclusions of the defaults in the scenario: \(\mathcal{W} \cup \text{conclusion}(S) \vdash \text{premise}(\delta)\).
  2. The default \(\delta \in \mathcal{D}\) is not conflicted, i.e., its conclusion does not contradict the background theory and scenario: it is not the case that \(\mathcal{W} \cup \text{conclusion}(S) \vdash \neg \text{conclusion}(\delta)\).
  3. The default \(\delta \in \mathcal{D}\) is not defeated by another default in the theory, i.e., there is no triggered default \(\delta^\prime > \delta\) such that \(\delta^\prime\) and \(\delta\) arrive at contradictory conclusions: it is not the case that \(\mathcal{W} \cup \text{conclusion}(\delta^\prime) \vdash \neg \text{conclusion}(\delta)\).

A scenario, \(S \subseteq \mathcal{D}\), is a proper scenario if it doesn’t include any extra defaults in \(\mathcal{D}\) that aren’t binding and doesn’t miss out any defaults that are binding. A proper scenario includes the good reasons for drawing an inference, and is what we (or an algorithm) are trying to find. More than one scenario may be proper.

We want to know what logically follows from each proper scenario, known as the extension, \(\mathcal{E}\), since it extends beyond the background theory and default rules to what they logically imply. This is a set of conclusions, defined

\(\mathcal{E} = \text{Th}\bigl(\mathcal{W} \cup \text{conclusion(S)}\bigr)\),

where \(\text{Th}(\Gamma)\) is the set of propositions, \(\phi\), such that \(\Gamma \vdash \phi\), i.e., every proposition that logically follows from \(\Gamma\).

Finally, how do we conclude that some proposition \(\phi\) ought to be the case, \(\mathsf{O} \phi\)? There are two ways to do this, depending on how we deal with conflicts:

  1. Conflict account: \(\mathsf{O} \phi\) follows if and only if \(\phi\) is included in the extension of at least one proper scenario.
  2. Disjunctive account: \(\mathsf{O} \phi\) follows if and only if \(\phi\) is included in extensions of all proper scenarios.

It is straightforward to find the proper scenarios for our example. We know both defaults are triggered, since the evidence \(\mathcal{W} = \{ P, H \}\), and the defaults are:

\(\displaystyle \begin{aligned}
d_1 &= P \rightarrow R \\
d_2 &= H \rightarrow \neg R
\end{aligned}\)

The conclusions of the two defaults contradict each other (\(R\) and \(\neg R\)), so if we included them both in a scenario they would be conflicted. Neither of the defaults could be defeated by the other since we have not put an ordering on them, and defeat requires an ordering. So there are two proper scenarios: \(\{ d_1 \}\) and \(\{ d_2 \}\), and two sets of extensions, call them \(\mathcal{E}_1\) and \(\mathcal{E}_2\).

  • \(\mathcal{E}_1\) includes \(\{ P, H, R \}\), so we can conclude \(\mathsf{O} R\).
  • \(\mathcal{E}_2\) includes \(\{ P, H, \neg R \}\), so we can conclude \(\mathsf{O} \neg R\).

Under the conflict account, this means we have two oughts: \(\mathsf{O} R\) and \(\mathsf{O} \neg R\). Note the extensions \(\mathcal{E}_1\) and \(\mathcal{E}_2\) both also include \(R \lor \neg R\) (using disjunction introduction), so by the disjunctive account we can conclude \(\mathsf{O}(R \lor \neg R)\). Therefore, with no ordering on the defaults, we can only say that we either ought to roll out the programme or not roll it out. Not particularly helpful!

To arrive at a decision, we need to put an ordering on the defaults. We could reason that an important normative value is first do no harm, which could be formalised as \(d_2 > d_1\). In this case, the only proper scenario is \(\{ d_2 \}\). This is because a proper scenario cannot have any defeated defaults, and in scenarios \(\{ d_1 \}\) and \(\{ d_1, d_2 \}\), \(d_1\) would be defeated by \(d_2\). We can’t use the empty scenario since it excludes binding defaults, so it is not a proper scenario. The extension of the proper scenario leads to the conclusion \(\mathsf{O} \neg R\): we ought not roll out the programme.

I noted earlier that SDL is ruled out as an appropriate logic for deontic reasoning. It does not, for example, handle conflicts in obligations. However, there’s something slightly puzzling about how Horty (2012) derives oughts: if \(\phi\) is in the extension of a scenario, we can conclude \(\mathsf{O} \phi\) for that scenario. But defaults can include non-deontic conclusions, such as the inference that penguins do not fly. It seems odd to move from this belief to the conclusion that penguins ought not fly. Fuhrmann (2017) suggests instead ditching Horty’s move from extensions to oughts and supplementing the system with SDL. Our original set of defaults would then be:

\(\displaystyle \begin{aligned}
d_1 &= P \rightarrow \mathsf{O} R \\
d_2 &= H \rightarrow \mathsf{O} \neg R
\end{aligned}\)

The idea is that instead of using classical logical consequence for building extensions, we use SDL. An extension of a scenario can then include an ought “for free”, since default conclusions can include oughts. The conflict and disjunctive accounts still work for examples Fuhrmann tries, yielding reasonable inferences, and we also still have the very helpful machinery of default logic. I don’t know whether adding SDL introduces other problems, though.

Does Horty’s prioritised default logic, perhaps supplemented with SDL, help evaluation? Firstly, it offers a deontic logic that formalises some elements of evaluative thinking. Even if we do not use it in a fully formal way, it provides clues about what evaluative thinking needs to include and highlights some of the issues that may arise. For example, we need to anticipate what the potential findings might be and how different normative values may lead to different evaluative judgements.

Secondly, it suggests that one way to do this is through default rules – formally or informally specified – with orderings where possible to reduce the number of conflicting conclusions. This allows some defaults to take priority over others. The justifications for that ordering lie outside the logic itself and form part of what coalitions of actors will debate.

Finally, as alluded to above, we also need an ordering on the evidence, and we need to ensure evaluations are designed so that, for example, if harms are identified, evidence of those harms cannot simply be dismissed on the grounds that the evidence is of insufficient quality. For example, suppose we have evidence from an impact evaluation (e.g., a quasi-experiment) of no harms (\(\neg H_I\)) and evidence from a process evaluation (e.g., qualitative interviews) of harms (\(H_P\)). We could setup defaults to interpret this evidence as:

\(\displaystyle \begin{aligned}
i_1 &= H_I \rightarrow H \\
i_2 &= \neg H_I \rightarrow \neg H \\
p_1 &= H_P \rightarrow H \\
p_2 &= \neg H_P \rightarrow \neg H
\end{aligned}\)

where \(H\) means we have inferred the programme has caused harm. Order the defaults so that \(i_x > p_x\) for \(x = 1,2\). This would mean that with conflicting evidence of harms, the impact evaluation evidence would always trump the process evaluation evidence. So, given the evidence \(\{ \neg H_I, H_P \}\), the conclusion would be no harm, \(\neg H\). This mirrors how process evaluation evidence of perceived impact is often, in practice, defeated by impact evaluation evidence of no impact. These interpretations can be anticipated before any data is collected and feed into the design of an evaluation.

I think it is unrealistic to expect an evaluation to settle debates about normative values, for the reasons explored in a previous post. However, deontic logic could help us think through what evidence needs to be produced to support others when they have those debates, making explicit both the assumed priorities among different forms of evidence and the ways in which this evidence is taken to justify particular policy decisions.

References

Bennett, A. (2009). Process Tracing: A Bayesian Perspective. In J. M. Box-Steffensmeier, H. E. Brady, & D. Collier (Eds), The Oxford Handbook of Political Methodology (pp. 702–721). Oxford University Press.

Fuhrmann, A. (2017). Deontic Modals: Why Abandon the Default Approach. Erkenntnis, 82(6), 1351–1365.

Horty, J. F. (2012). Reasons as defaults. Oxford University Press. (Final version preprint available here.)

Parent, X., & Torre, L. van der. (2018). Introduction to Deontic Logic  and Normative Systems. College Publications.

Sabatier, P. A. (1988). An advocacy coalition framework of policy change and the role of policy-oriented learning therein. Policy Sciences, 21, 129–168.

Bixonimania

Research involves immersing yourself in the literature; engaging in a struggle to try to keep up with key findings in your field; critically engaging with studies, learning from their strengths and unintentional errors and mishaps to improve your work; attending conferences and debating in and between sessions to reach an informed view on what’s going on; maintaining collections of papers that may be relevant in future, e.g., using reference managers like Zotero to keep track and aid searching, annotation, and citation.

LLMs can’t engage and think for you, can’t magically place the ideas in your brain, fails to find papers that a simple literature search uncovers. But the literature is being flooded by complete nonsense, driven by some researchers’ credulous reliance on what LLMs produce. Bixonimania is a great illustration of what can go wrong, but the prevalent examples are more subtle – worth a read of the Nature writeup: Scientists invented a fake disease. AI told people it was real.