What should the is-ought thesis state?

Hume’s is-ought thesis states that we cannot infer a normative statement about what we should do from a descriptive statement about what is the case. This post spells out some of the detail about what the thesis means, which should be of interest to policy evaluators since we’re in the business of producing evidence and offering advice on policy that’s consistent with that evidence.

Let’s start with Prior’s (1960) puzzle (mildly edited): what kind of statement is “Either it’s raining or cricket should be banned”? It combines a descriptive statement (“it’s raining”) with a normative one (“cricket should be banned”), but is the disjunction of the two descriptive or normative?

Suppose Prior’s puzzle is a descriptive statement. Then by adding it to the premise “it’s not raining”, we can derive the purely normative conclusion that “cricket should be banned”. Since the disjunction is true, but one of its disjuncts is false, then the other disjunct must be true. We have just drawn an is-ought inference.

In symbols,

\(\displaystyle \neg R, R \lor \mathsf{O} B\ \models\ \mathsf{O} B\),

where \(R\) denotes “it’s raining”, \(\mathsf{O} B\) denotes that we ought to ban cricket (the \(\mathsf{O}\) is the “ought” from deontic logic), \(\lor\) is disjunction (or), \(\neg\) is negation, and \(\models\) is logical consequence.

This is maybe easier to see by rewriting the disjunction as a conditional, “If it’s not raining, then cricket should be banned” (since \(\neg \phi \lor \psi = \phi \rightarrow \psi\)):

\(\displaystyle \neg R, \neg R \rightarrow \mathsf{O} B\ \models\ \mathsf{O} B\),

where \(\rightarrow\) is the material conditional. The conclusion is then drawn by modus ponens. This pattern is interesting for evaluation, since it mirrors the logic:

  1. The programme works*.
  2. If the programme works*, we should roll it out nationally.
  3. Therefore, we should roll it out nationally.

Where works* includes a range of statements about impact and process evaluation evidence, whether that evidence can be generalised to a broader population, whether the programme avoids causing harm, its cost‑effectiveness, and other considerations.

In symbols,

\(\displaystyle W^*, W^* \rightarrow \mathsf{O} R\ \models\ \mathsf{O} R\),

where \(W^*\) denotes the programme works* and \(R\) denotes that we should roll it out.

Suppose instead that Prior’s puzzle is a normative statement. Starting from the purely descriptive premise “it’s raining”, we can infer “Either it’s raining or cricket should be banned” by disjunction introduction. Again, we have drawn an is-ought inference.

In symbols,

\(\displaystyle R\ \models\ R \lor \mathsf{O} B\).

Rewriting using implication,

\(\displaystyle R\ \models\ \neg R \rightarrow \mathsf{O} B\),

the conclusion is one of the “paradoxes” of the material conditional: true under classical logic because the antecedent is false, but people tend to judge the conditional to be neither true nor false but irrelevant (Johnson-Laird & Tagart, 1969).

The solution is, “Either it’s raining or cricket should be banned” is neither descriptive nor normative: it’s mixed. The same applies to “If the programme works*, we should roll it out nationally”. Three refinements of the is-ought thesis are provided in a logic-heavy book by Schurz (1997), which has been on my reading stack for years. A more digestible summary is provided by Schurz (2014, pp. 2–3):

(H1) No non-logically true purely normative conclusion can be derived from a consistent set of purely descriptive premises.

(H2) Every mixed conclusion [i.e., combining descriptive and normative statements] which follows logically from a set of purely descriptive premises is normatively irrelevant in the sense that all of its normative subformulas are replaceable by other arbitrary subformulas, while preserving the validity of the inference [“salva validitate of the inference” in Schurz’s original].

(H3) No non-tautologous descriptive statement can be inferred from a consistent set of purely normative premises.

H1 is the thesis that applies most to evaluation: to make an evaluative judgement, you need mixed premises that blend normative and descriptive statements. H2 deals with weird uses of logic. I’ve used the principle of irrelevance it contains to help understand how people reason about sentences like “If Alex posted the letter, then he posted the letter or set fire to the letter”, which are true in classical logic but which people often judge to be false (Fugard et al., 2011). There’s a blog post about it yonder. H3 says that knowing or believing what should be true doesn’t tell you what is factually true.

There is, however, a well‑known problem with using the material conditional in combination with oughts (Chisholm, 1963), e.g., \(W^* \rightarrow \mathsf{O} R\) used above. Consider the following sentences and formalisations:

  1. It ought to be that Jones goes to assist his neighbours: \(\mathsf{O} g\).
  2. It ought to be that if Jones goes, then he tells them he is coming: \(\mathsf{O} (g \rightarrow t)\).
  3. If Jones doesn’t go, then he ought not tell them he is coming. \(\neg g \rightarrow \mathsf{O} \neg t\).
  4. Jones doesn’t go: \(\neg g\).

The English‑language statements feel consistent with each other. They are also independent in the sense that no one sentence follows from any of the others. A formalisation should preserve both features.

There are three paths through, using what has come to be known as standard deontic logic (SDL). McNamara and Van De Putte (2025, Section 2.1) provides an introduction to SDL. Section 4.1 provides the illustration of Chisholm’s (1963) problem, which I’ve just spelt out a little. Here’s a summary of the paths:

Path APath BPath C
(1′) \(\mathsf{O}g\)
(2′) \(\mathsf{O}(g \rightarrow t)\)
(3′) \(\neg g \rightarrow \mathsf{O}\neg t\)
(4′) \(\neg g\)
(1′) \(\mathsf{O}g\)
(2′) \(\mathsf{O}(g \rightarrow t)\)
(3″) \(\mathsf{O}(\neg g \rightarrow \neg t)\)
(4′) \(\neg g\)
(1′) \(\mathsf{O}g\)
(2″) \(g \rightarrow \mathsf{O}t\)
(3′) \(\neg g \rightarrow \mathsf{O}\neg t\)
(4′) \(\neg g\)
From (1′), (2′), \(\mathsf{O}t\).
From (3′), (4′), \(\mathsf{O}\neg t\).
Consistency is lost.
(1′) implies (3″).
Independence is lost.
(4′) implies (2″).
Independence is lost.

Following Path A, we deduce that Jones ought to tell them he is coming and ought not tell them he is coming, which is suspect. This uses the following SDL rule:

\(\mathsf{O}(\varphi \rightarrow \psi) \rightarrow (\mathsf{O}\varphi \rightarrow \mathsf{O}\psi)\text{,} \tag{OB-K}\)

which gives us \(\mathsf{O}g \rightarrow \mathsf{O} t\) (from 2′). This (alongside 1′) gives us \(\mathsf{O}t\) by modus ponens. We can also get \(\mathsf{O}\neg t\) using modus ponens (3′ and 4′). It also contradicts one of the axioms of SDL: if you should do something, then you shouldn’t also do the negation of that something:

\(\mathsf{O} \phi \rightarrow \neg \mathsf{O}\neg \phi \tag{NC}\)

Paths B and C attempt to use the same expression for the conditional oughts in sentences 2 and 3. Path B uses

\(\mathsf{O}(\phi \rightarrow \psi)\).

Path C uses

\(\phi \rightarrow \mathsf{O}\psi\),

which we encountered above when exploring Prior’s (1960) puzzle.

The problem with both attempts to save SDL is that the sentences become dependent, whereas in the original informal English they aren’t. For Path B, \(\mathsf{O}(\neg g \rightarrow \neg t)\) is vacuously true (when rewritten using OB-K) because \(\mathsf{O}g\) is true – a paradox of the material conditional again. Similarly for Path C, \(g \rightarrow \mathsf{O}t\) is vacuously true because \(\neg g\) is.

So however we choose to formalise the sentence “if you should do A, then you should do B” or “if A, then you should do B”, it is not something that can be captured within SDL – a conclusion reached in the late 1960s (Parent & Torre, 2018, p. 20). This conclusion is familiar from work in the psychology of reasoning, where the material conditional is replaced with systems that behave more like everyday inference.

One approach uses defeasible logics, which allow us to retract a conclusion when new information arrives and to treat some premises as having greater priority than others (e.g., Neves et al., 2002). Another approach uses probability logics, which model reasoning under uncertainty (e.g., Pfeifer & Kleiter, 2009). Probability logics typically rely on an underlying three‑valued semantics, in which a statement can be true, false, or neither. The third value is usually interpreted as something like “irrelevant” or “undetermined”.

A recent review of deontic logic (McNamara & Van De Putte, 2025) concludes that “there are a number of outstanding problems for deontic logic. Some see this as a serious defect; others see it merely as a serious challenge, even an attractive one.” I’ve been reading attempts to solve some of these problems, e.g., Horty (2012), which applies defeasible logic to deontic reasoning. See the follow-up.

References

Chisholm, R. M. (1963). Contrary-to-duty imperatives and deontic logic. Analysis, 24, 33–36.

Fugard, A., Pfeifer, N., & Mayerhofer, B. (2011). Probabilistic theories of reasoning need pragmatics too: modulating relevance in uncertain conditionals. Journal of Pragmatics, 43, 2034–2042.

Horty, J. F. (2012). Reasons as defaults. Oxford University Press.

Johnson-Laird, P., & Tagart, J. (1969). How implication is understood. The American Journal of Psychology, 82, 367–373.

McNamara, P., & Van De Putte, F. (2025). Deontic logic. In E. N. Zalta & U. Nodelman (Eds), The Stanford encyclopedia of philosophy (Winter 2025). Metaphysics Research Lab, Stanford University.

Neves, R. D. S., Bonnefon, J.-F., & Raufaste, E. (2002). An Empirical Test of Patterns for Nonmonotonic Inference. Annals of Mathematics and Artificial Intelligence, 34, 107–130.

Parent, X., & Torre, L. van der. (2018). Introduction to Deontic Logic  and Normative Systems. College Publications.

Pfeifer, N., & Kleiter, G. D. (2009). Framing human inference by coherence based probability logic. Journal of Applied Logic, 7, 206–217.

Prior, A. N. (1960). The autonomy of ethics. Australasian Journal of Philosophy, 38(3), 199–206.

Schurz, G. (1997). The Is-Ought Problem: An Investigation in Philosophical Logic. Springer.

Schurz, G. (2014). Cognitive success: Instrumental justifications of normative systems of reasoning. Frontiers in Psychology, 5(625).

Drawing an is-ought

Hume’s Treatise famously argued that we cannot infer an “ought” from an “is”. This has presented an enduring problem for science: how should we produce a set of recommendations for what should be done following the results of a study? If a new cancer treatment dramatically improves remission rates, should study authors simply shrug, present the results, and leave the recommendations to politicians? What if a treatment causes significant harms – can we recommend that the treatment be banned? Or suppose we have ideas for future studies that should be carried out and want to summarise them in the conclusions…?

The solution, if it is one, is that any recommendations require a set of premises stating our values. These values necessarily assert something beyond the evidence, for instance that if a treatment is effective then it should be provided by the health service. In practice, such values are often left implicit and assumed to be shared with readers. But there are interesting examples where it is apparently possible to draw an is-ought inference without assuming values.

One example, due to Mavrodes (1964), begins with the premise

If we ought to do A, then it is possible to do A.

This seems reasonable enough. It would, for instance, be horribly dystopian to require that people behave a particular way if it were impossible for them to do so. Games like chess and tennis have rules that are possible – if they were impossible then it would make playing the games challenging. Let’s see what happens if we apply a little logic to this premise.

Sentences of the form

If A, then B

are equivalent to those of the contrapositive form

If not-B, then not-A

This can be seen in the truth table below, where 1 denotes true and 0 denotes false. The values of the last two columns are equivalent:

A B not-A not-B If A, then B If not-B, then not-A
1 1 0 0 1 1
1 0 0 1 0 0
0 1 1 0 1 1
0 0 1 1 1 1

Together, this means that if we accept the premise

If we ought to do A, then it is possible to do A,

and the rules of classical logic, we must also accept

If it is not possible to do A, then it is not the case that we ought to do A.

But here we have an antecedent that is an “is” and a consequent that is an “ought”: logic has licenced an is-ought!

Worry not: there has been debate in the literature… See Gillian Russell (2021) for a recent analysis.

References

Mavrodes, G. I. (1964). “Is” and “Ought.” Analysis, 25(2), 42–44.

Russell, G. (2022). How to Prove Hume’s Law. Journal of Philosophical Logic, 51(3), 603–632.

A psychoanalyst walks into a bar(red subject)

A psychoanalyst walks into a bar with a book on logic and set theory. He orders a whisky. And another. Twelve hours and a lock-in later, all he has to show for the evening is a throbbing headache and some indecipherable rubbish scrawled on a napkin.

That’s the only conceivable explanation for these diagrams from The Subversion of the Subject and the Dialectic of Desire in the Freudian Unconscious, by Jacques Lacan (published in the Écrits collection):

But, surely this notation means something? After all, Lacan is famous and academics across the world dedicate their lives to understanding his genius.

Also the notation f(x) is a function, f, applied to argument x – that’s recognisable from maths. So the I(A) and s(A) must mean something…?

To illustrate how function notation is usually used, consider the Fibonacci sequence, which pops up in all kinds of interesting places in nature. It is defined as follows:

f(0) = 0,
f(1) = 1,
f(n) = f(n-1) + f(n-2), for n > 1.

In English, this says that the first two numbers in the sequence are 0 and 1 and the numbers following are obtained by summing the previous two. So the sequence goes: 0, 1, 1, 2, 3, 5, 8, 13, 21, 34, …

The function notation “does something”. It provides a way of defining and referring to (here, mathematical) concepts. I claim that the brief explanation above would make some kind of sense to most people who can add two numbers together.

Less well-known, but appearing in university philosophy courses, is the lozenge symbol, ◊, which means “possible” in a particular kind of logic called modal logic. So if R stands for “it’s raining” then ◊R stands for “it’s possible that it’s raining”. It seems plausible that there is something meaningful here in Lacan’s use of the symbol too.

Here is Lacan, “explaining” his notation to his almost entirely non-mathematical readership:

Huh?

Lacan doesn’t try to explain what the notation means; he doesn’t seem to want readers to understand. Maybe he is just too clever and if only we persevered we would get what he means. Elsewhere in the same text, Lacan uses arithmetic to argue that “the erectile organ can be equated with \(\sqrt{-1}\)”. I’m told this is a joke because \(\sqrt{-1}\) is an imaginary number. Maybe trainee psychoanalysts learn about complex numbers so get the joke. I doubt it though. Maybe all Lacanian discourse is dadaist performance – that at least would make some sense.

Alan Sokal and Jean Bricmont have written a book-length critique of Lacan’s maths and others’ similar misuse of natural science concepts. Having read lots of mathematical texts and seen how authors make an effort to introduce their notation, I think it’s entirely possible Lacan is a fraud, ◊(Lacan is a fraud). That might sound harsh, but forget how famous he is and just look at the pretentious rubbish he writes.

A Connectionist Computational Model for Epistemic and Temporal Reasoning

Many researchers argue that logics and connectionist systems complement each other nicely. Logics are an expressive formalism for describing knowledge, they expose the common form across a class of content, they often come with pleasant meta-properties (e.g. soundness and completeness), and logic-based learning makes excellent use of knowledge. Connectionist systems are good for data driven learning and they’re fault tolerant, also some would argue that they’re a good candidate for tip-toe-towards-the-brain cognitive models. I thought I’d give d’Avila Garcez and Lamb (2006) a go [A Connectionist Computational Model for Epistemic and Temporal Reasoning, Neural Computation 18:7, 1711-1738].

I’m assuming you know a bit of propositional logic and set theory.

The modal logic bit

There are many modal logics which have properties in common, for instance provability logics, logics of tense, deontic logics. I’ll follow the exposition in the paper. The gist is: take all the usual propositional logic connectives and add the operators □ and ◊. As a first approximation, □P (“box P”) means “it’s necessary that P” and ◊P (“diamond P”) means “it’s possible that P”. Kripke models are used to characterise when a model logic sentence is true. A model, M, is a triple (Ω, R, v), where:

  • Ω is a set of possible worlds.
  • R is a binary relation on Ω, which can be thought of as describing connectivity between possible worlds, so if R(ω,ω’) then world ω’ is reachable from ω. Viewed temporally, the interpretation could be that ω’ comes after ω.
  • v is a lookup table, so v(p), for an atom p, returns the set of worlds where p is true.

Let’s start with an easy rule:

(M, ω) ⊨ p iff ω ∈ v(p), for a propositional atom p

This says that to check whether p is true in ω, you just look it up. Now a recursive rule:

(M, ω) ⊨ A & B iff (M, ω) ⊨ A and (M, ω) ⊨ B

This lifts “&” up to our natural language (classical logic interpretation thereof) notion of “and”, and recurses on A and B. There are similar rules for disjunction and implication. The more interesting rules:

(M, ω) ⊨ □A iff for all ω’ ∈ Ω such that R(ω,ω’), (M, ω’) ⊨ A

(M, ω) ⊨ ◊A iff there is an ω’ ∈ Ω such that R(ω,ω’) and (M, ω’) ⊨ A

The first says that A is necessarily true in world ω if it’s true for all connected worlds. The second says that A is possibly true if there is at least one connected world for which it is true.

A sketch of logic programs and a connectionist implementation

Logic programs are sets of Horn clauses, A1 & A2 & … & An → B, where Ai is a propositional atom or the negation of an atom. Below is a picture of the network that represents the program {B & C & ~D → A, E & F → A, B}.

A network representing a program

The thresholds are configured so that the units in the hidden layer, Ni, are only active when the antecedents are all true, e.g. N1 is only active when B, C, and ~D have the truth value true. The thresholds of the output layer’s units are only active when at least one of the hidden layer connections to them is active. Additionally, the output feeds back to the inputs. The networks do valuation calculations through the magic of backpropagation, but can’t infer new sentences as such, as far as I can tell. To do so would involve growing new nets and some mechanism outside the net interpreting what the new bits mean.

Aside on biological plausibility

Biological plausibility raises its head here. Do the units in this network model – in any way at all – individual neurons in the brain? My gut instinct says, “Absolutely no way”, but perhaps it would be better not even to think this as (a) the units in the model aren’t intended to characterise biological neurons and (b) we can’t test this particular hypothesis. Mike Page has written in favour of localists nets, of which this is an instance [Behavioral and Brain Sciences (2000), 23: 443-467]. Maybe more on that in another post.

Moving to modal logic programs and nets

Modal logic programs are like the vanilla kind, but the literals may have one of the modal operators. There is also a set of connections between the possible worlds, i.e. a specification of the relation, R. The central idea of the translation is to use one network to represent each possible world and then apply an algorithm to wire up the different networks correctly, giving one unified network. Take the following program: {ω1 : r → □q, ω1 : ◊s → r, ω2 : s, ω3 : q → ◊p, R(ω1,ω2), R(ω1,ω3)}. This wires up to:

A network representing a modal logic program

Each input and output neuron can now represent □A, ◊A, A, □~A, ◊~A, or ~A. The individual networks are connected to maintain the properties of the modality operators, for instance □q in ω1 connects to q in ω2 and ω3 since R(ω1, ω2), R(ω1, ω3), so q must be true in these worlds.

The Connectionist Temporal Logic of Knowledge

Much the same as before, except we now have a set of agents, A = {1, …, n}, and a timeline, T, which is the set of naturals, each of which is a possible world but with a temporal interpretation. Take a model M = (T, R1, …, Rn, π). Ri specifies what bits of the timeline agent i has access to, and π(t) gives a set of propositions that are true at time t.

Recall the following definition from before

(M, ω) ⊨ p iff ω ∈ v(p), for a propositional letter p

Its analogue in the temporal logic is

(M, t) ⊨ p iff t ∈ π(p), for a propositional letter p

There are two extra model operators: O, which intuitively means “at the next time step” and K which is the same as □, except for agents. More formally:

(M, t) ⊨ OA iff (M, t+1) ⊨ A

(M, t) ⊨ KA iff for all u ∈ T such that Ri(t,u), (M, u) ⊨ A

Now in the translation we have network for each agent, and a collection of agent networks for each time step, all wired up appropriately.

Pages 1724-1727 give the algorithms for net construction. The proof of soundness of translation relies on d’Aliva Garcez, Broda, and Gabbay (2002), Neural-symbolic learning systems: Foundations and applications.

Some questions I haven’t got around to working out the answers to

  • How can these nets be embedded in a static population coded network. Is there any advantage to doing so?
  • Where is the learning? In a sense it’s the bit that does the computation, but it doesn’t correspond to the usual notion of “learning”.
  • How can the construction of a network be related to what’s going on in the brain? Really I want a more concrete answer to how this could model the psychology. The authors don’t appear to care, in this paper anyway.
  • How can networks shrink again?
  • How can we infer new sentences from the networks?

Comments

I received the following helpful comments from one of the authors, Artur d’Avila Garcez (9 Aug 2006):

I am interested in the localist v distributed discussion and in the issue of biological plausibility; it’s not that we don’t care, but I guess you’re right to say that we don’t “in this paper anyway”. In this paper – and in our previous work – what we do is to say: take standard ANNs (typically the ones you can apply Backpropagation to). What logics can you represent in such ANNs? In this way, learning is a bonus as representation should precede learning.

The above answers you question re. learning. Learning is not the computation, that’s the reasoning part! Learning is the process of changing the connections (initially set by the logic) progressively, according to some set of examples (cases). For this you can apply Backprop to each network in the ensemble. The result is a different set of weights and therefore a different set of rules – after learning if you go back to the computation you should get different results.