The green whistle test

I have some lived-experience to report and a brief reflection on an RCT. Content warning: it involves a brief description of an injury, but I lived happily ever after!

I recently dislocated my shoulder after slipping on a traverse climbing wall – ten minutes before the end of an introductory bouldering session. Thankfully, I wasn’t far off the ground. As far as these things go, it was a “good” dislocation: I feel fine and am starting physio imminently.

It wasn’t my first. Over a decade ago, I had a dislocation and was given nitrous oxide (gas and air) at A&E while they reattached my arm. This time, I was handed Penthrox – also known as the green whistle – a vape-sized handheld device that delivers methoxyflurane.

Subjectively, Penthrox felt much better than gas and air. During the procedure, I’d describe my mental state as “relaxed, pain-free, and a bit messed up but in a good way.” Once my shoulder was back in place, I would say I was “elated and very chatty” (emphasis on the “very”: apologies to everyone else waiting for an x-ray with me).

In contrast, gas and air didn’t do as much for the pain, though it did have the “messed up” component, which was helpfully distracting – at least as I remember it. There were plenty of confounding variables: injury severity (previously I needed an ambulance, this time not) and memory effects among them. So I wondered what the evidence says about the average differences between the two.

I came across a systematic review (Porter et al., 2017). Both gas and air and Penthrox outperformed placebo; however, no statistically significant difference was found between them in direct comparisons. But hang on – someone ran a study comparing Penthrox with placebo? I’m not sure I’d have met the inclusion criteria, but imagining being invited to take part in a placebo-controlled trial while hanging onto my arm, I suspect my response would have involved expressive language and a firm decline.

The justification for the placebo-control was: “Use of an active comparator, although preferred, would have posed considerable challenges to keep the study blind because of the unique mode of delivery and smell of methoxyflurane” (Coffey et al., 2014, p. 614).

An interesting case example to discuss.

References

Coffey, F., Wright, J., Hartshorn, S., Hunt, P., Locker, T., Mirza, K., & Dissmann, P. (2014). STOP!: A randomised, double-blind, placebo-controlled study of the efficacy and safety of methoxyflurane for the treatment of acute pain. Emergency Medicine Journal, 31(8), 613–618.

Porter, K. M., Siddiqui, M. K., Sharma, I., Dickerson, S., & Eberhardt, A. (2017). Management of trauma pain in the emergency setting: Low-dose methoxyflurane or nitrous oxide? A systematic review and indirect treatment comparison. Journal of Pain Research, 11, 11–21.

RCTs can and do use sensible control groups

If there are effective treatments for a condition, then a placebo-controlled randomised trial would generally be unethical (see below for exceptions discussed in the literature). But RCTs do not need to use placebo control.

There is a common misconception that an RCT should have a “pure” control (whatever that means) and if an active control of some description is used, e.g., another treatment, then it becomes A/B testing. This is false. A/B testing is just a type of trial and the term is more commonly used in marketing, e.g., in studies of variations of a website design.

RCTs can and do have active controls. Education trials and psychological therapy trials often use usual practice controls. Once effective Covid vaccines had been found, trials investigated the comparative efficacy and safety of different vaccines.

Example discussions

“Often, trials are conducted to compare a new intervention against current practice. The new intervention might be a small change, or a set of small changes to current practice; or it could be a whole new approach which is proving to be successful in a different country or context, or that has sound theoretical backing.” (Haynes, Goldacre, & Torgerson, 2012, p. 20)

“An RCT is not necessarily a test between doing something and doing nothing. Many interventions might be expected to do better than nothing at all. Instead, trials can be used to establish which of a number of policy intervention options is best.” (Haynes, Goldacre, & Torgerson, 2012, p. 21)

“Placebo controls are clearly inappropriate for conditions in which delay or omission of available treatments would increase mortality or irreversible morbidity in the population to be studied. For conditions in which forgoing therapy imposes no important risk, however, the participation of patients in placebo-controlled trials seems appropriate and ethical, as long as patients are fully informed.” (Temple & Ellenberg, 2000, p. 460)

“In cases where an available treatment is known to prevent serious harm, such as death or irreversible morbidity in the study population, it is generally inappropriate to use a placebo control. There are occasional exceptions, however, such as cases in which standard therapy has toxicity so severe that many patients have refused to receive it.” (European Medicines Agency, 2001, p. 16)

“Many placebo-controlled trials are conducted as add-on trials, where all patients receive a specified standard therapy or therapy left to the choice of the treating physician or institution.” (European Medicines Agency, 2001, p. 16)

References

European Medicines Agency (2001). ICH Topic E 10: Choice of Control Group in Clinical Trials.

Haynes, L., Goldacre, B., & Torgerson, D. (2012). Test, learn, adapt: developing public policy with randomised controlled trials. Cabinet Office and Behavioural Insights Team.

Temple, R., & Ellenberg, S. S. (2000). Placebo-controlled trials and active-control trials in the evaluation of new treatments. Part 1: ethical and scientific issues. Annals of Internal Medicine, 133(6), 455-463.

Approaches to consent in public health research in secondary schools

“Seeking active parent consent can undermine secondary school students’ autonomy, and limit participation, particularly among disadvantaged students, so biasing research. Our analysis suggests that active student consent and passive parent/carer consent be standard practice for most research procedures in secondary schools. More intrusive data collection, such as blood and saliva samples, would require parent/carer active consent since such procedures would be defined as diagnostic procedures so being classed as an investigational product. However, we would argue that for questionnaire completion, observation or routine data, student consent and autonomy should have primacy with parents having the right and means to receive full information, ask questions and withdraw their children from research should they wish. This approach gives proper primacy to student autonomy while also respecting parent/carer autonomy.”

Also helpful thoughts on issues arising with consent for whole-class interventions.

Bonell, C., Humphrey, N., Singh, I., Viner, R. M., & Ford, T. (2023). Approaches to consent in public health research in secondary schools. BMJ Open, 13, e070277.

Harvey at PwC

LLM-driven text analysis is becoming a norm, allowing people to process huge volumes of text they wouldn’t otherwise have the capacity to do. Although outputs can be checked, the large volume of inputs processed means there are fundamental limits on how comprehensively analyses can be checked.

PwC announced yesterday that it is trialling the use of Harvey, built on Chat GPT, to “help generate insights and recommendations based on large volumes of data, delivering richer information that will enable PwC professionals to identify solutions faster.”

They say that “All outputs will be overseen and reviewed by PwC professionals.” But what about how the data was processed in the first place…?

“Randomista mania”, by Thomas Aston

Thomas Aston provides a helpful summary of RCT critiques, particularly in international evaluations.

Waddington, Villar, and Valentine (2022), cited therein, provide a handy review of comparisons between RCT and quasi-experimental estimates of programme effect.

Aston also cites examples of unethical RCTs. One vivid example is an RCT in Nairobi with an arm that involved threatening to disconnect water and sanitation services if landlords didn’t settle debts.

Migration and the Value of Social Networks

I haven’t read this working paper yet – just struck by this dataset:

“We leverage a rich new source of ‘digital trace’ data to provide a detailed empirical perspective on how social networks influence the decision to migrate. These data capture the entire universe of mobile phone activity in Rwanda over a five-year period. Each of roughly one million individuals is uniquely identified throughout the dataset, and every time they make or receive a phone call, we observe their approximate location, as well as the identity of the person they are talking to. From these data, we can reconstruct each subscriber’s 5-year migration trajectory, as well as a detailed picture of their social network before and after migration

An infinite trolley problem

Remember the trolley problem? There are now thousands of variants of this ethical conundrum. There’s a curious infinite variant, where there are as many people as integers on the top track and as many people as real numbers on the bottom:

My first thought on seeing this was, aha, finally a chance to apply Cantor’s diagonal argument to something useful: to solve the problem of whether or not to pull the lever.

Let’s start with an easier version. Suppose the segment of the lower rail where people are tied is bounded in length to a whisker short of 100 meters. Each person is tied at a position on that segment somewhere greater than or equal to 0 metres and less than 100 metres from the beginning of the segment. Another way to write that range of positions is [0,100).

All the infinitely many positions in [0, 100) have to be used. So there’s somebody at exactly 0 metres, someone at 43.54377239879432 metres, someone at 3.5 metres, and so on for all the real number positions in [0, 100).

Number the people 0, 1, 2, 3, … starting at the end of the track from which the train is approaching. Now let’s work along that infinite philosophically imagined mound of people and construct a position on the track as follows from where they are lying along the rail (in metres). You have a very precise measuring tape.

From person 0, take the number to the left of the decimal point on their measurement and compute 99 minus that number. From person 1, take 9 minus the 1st number to the right of the decimal point. From person 2 take 9 minus the 2nd digit, and so on. So from person i, take 9 minus the ith decimal digit (pad out the digits with zeros where necessary).

Here are some examples of positions:

Person 0: 0.1455487...
Person 1: 0.5534524...
Person 2: 1.2364765...
Person 3: 2.4500000...
Person 4: 3.6273692...
...

From these, we would calculate the following:

  • \(99 – 0 = 99\)
  • \(9 – 5 = 4\)
  • \(9 – 3 = 6\)
  • \(9 – 0 = 9\)
  • \(9 – 3 = 6\)

So now we have a position on the track, 99.4696… m along.

Think about what we have done here. We have worked along all the people and calculated a new position in [0,100) where nobody is tied to the track. This is because the new position differs from each existing position on at least one digit, by construction. But we are supposed to have one person at all infinitely many real number positions on the track. Here we have found a gap, a real number that isn’t being used.

We could add someone else at 99.4696… m. But if we did that, we could just follow the procedure above again to find another gap.

It follows that our original assumption is false: it is not possible to stack infinitely many people along infinitely many real-numbered positions along an almost 100 metre long segment of track. Assuming we could has led to a contradiction.

We can’t do it for [0,100). That means there’s no hope for doing it for all real numbers since [0,100) is a subset of the reals.

So it is not possible to tie heaps of people to a track so that there are as many people as there are real numbers. People are rather discrete, countable, beings, and I could have stopped at “Number the people 0, 1, 2, 3, …”.

It turns out that the two tracks must have the same countably infinite number of people, so it doesn’t matter whether you pull the lever.