QSL: Oostenderadio

2237 UTC, 10 June 2026, 2761 kHz USB, using an XHDATA D-808 with XHDATA AN-80 antenna.

“I hereby wish to give you confirmation that you actually heard a broadcast of the Coaststation OOSTENDERADIO on frequency 2761 kHz. at 102233 [sic] UTC JUN 2026.

Our frequency 2761 kHz is a shore to ship working frequency, also used for broadcasting our Maritime Safety Information at fixed times, starting at 02.33 UTC and repeating these safety messages every 4 hours.

Oostenderadio also keeps watch 24 hours a day on 2182 kHz. This frequency is also used to announce all our broadcasts (Maritime Safety Information), as well as to reply on a call from any vessel for a radio-check.

We also keep watch 24 hours a day on the coupled frequency 3178 kHz, and reply on 2484 kHz, as well as on 4095 kHz and reply on 4387 kHz. The range is a large part of the North Sea.

The transmitters we use, are of the type Marconi H1141 with a power of 10 kW, located at Wingene, 51°11N 002°49E.”

Website

History:

1918-1930 rebuilding of the coast station – callsign OST – located at Oostende.
1929 r/telephony operational – callsign OSU.
1932-1940 modernisation of Oostenderadio – located at Stene.
1947 transmitters located at Middelkerke – receivers located at Stene.
April 1st 1964 opening of a new station located at Oudenburg.
full remote control of all transmitters located at Middelkerke and at Ruiselede.
Dec.15th 1977 opening of a ultra-modern station in the center of Oostende.
r/telex fully modernised and operational.
radio traffic on VHF, MF and HF – r/telephony and r/telegraphy.
1987 full automatic Navtex.
Nov. 1995 closing down of HF r/telegraphy.
Feb. 1999 DSC fully operational.
2001 Oostenderadio becomes part of the Belgian Navy.
March 2004 closing down of r/telex.
Oct. 30th 2006 station located in a new Comm Center at Zeebrugge 51°20N 003°12E.
March 2016 Oostenderadio relocates in Oostende, with MRCC.
TECHNICAL DATA
MF/HF transmitters Marconi H1140 – 1/10 kW and NAVTEX transmitters SAIT-Devlonics CST 3001 –
1 kW – located at Ruiselede/Wingene 51°11 N 002°49 E.

Association of vehicle emission charging zones with adult emergency hospital admissions: an interrupted time series analysis of the toxicity charge and Ultra Low Emission Zone in London, UK

Interesting comparative interrupted time series doing the rounds in the press at the moment.

Chamberlain, R. C., Hodsoll, J., Cai, C., Griffiths, C., Kelly, F., Beevers, S., Stewart, G., Blangiardo, M., Elliott, P., Davies, B., & Fecht, D. (2026). Association of vehicle emission charging zones with adult emergency hospital admissions: An interrupted time series analysis of the toxicity charge and Ultra Low Emission Zone in London, UK. Environment International, 213, 110324. https://doi.org/10.1016/j.envint.2026.110324

Whom and context again

I reread Mayne’s (2019) on contribution analysis today and see he also allows a “what works for whom”:

“… a more interesting and important contribution claim is around this evaluation question: How and why has the intervention (or component) made a difference, or not, and for whom?” (p. 174-5).

He also allows ToCs to describe contexts: “Using nested ToCs to unpack a complex intervention and its context has worked well in numerous situations…” (p. 178).

Mayne (2012) put context in the ToCs too: “External factors – context and rival explanations – influencing the intervention are assessed and are either shown not to have made a significant contribution or, if they did, their relative contribution is recognized.” (p. 273)

So there’s another theory-based brand with a “what works for whom in what context”, plus it uses generative causation 🙃

Of course I jest. Any approach can use generative causation, whether RCT or qualitative impact evaluation. Any approach can explore moderators of impact.

References

Mayne, J. (2012). Contribution analysis: Coming of age? Evaluation, 18(3), 270–280.

Mayne, J. (2019). Revisiting contribution analysis. Canadian Journal of Program Evaluation, 34(2), 171–191.

PM vs. Reform leader (3 June 2026)

PM: “The grieving family have asked us not to respond in the way that the leader of Reform has responded. They have lost their son in the most appalling circumstances, and they make a simple plea of us as human beings to please not exploit that. We all need to reflect on the words of Henry’s father.

“My response—and the response of others, to be fair—has been focused on the lessons to be learned so that we can deliver justice. The hon. Gentleman’s response has been to appeal for rage. That is his response to a father who has lost his son and asked for that not to happen. Exploiting this tragedy to create grievance and division would be wrong in any circumstances, but to do it when the family are expressly saying, “Please don’t,” is unforgivable. It shows exactly who he is.”

(PMQs, Wednesday 3 June 2026)

Extended two-way fixed effects

selectTWFE looks interesting: “Estimates both a vanilla two-way fixed effects (TWFE) model and an extended TWFE (ETWFE) model, then selects between them using Cochran’s Q test for heterogeneity. When ETWFE wins, reports the heterogeneity fraction (I-squared) and cohort-time estimates with empirical Bayes shrinkage and Bonferroni multiplicity correction.”

Dismantling evaluation brands

Thomas Aston (2026) has written a really interesting discussion of debates about what counts as realist evaluation and whether it is a unique kind of evaluation.

As I’ve written before (and Aston cites), I think theorising contexts, mechanisms, and outcomes (CMOs) is important, as is recognising a gap between theory, phenomena, and evidence. However, as Pawson (2024, p. 42) notes: “All [!] scientific investigation utilises explanations relating mechanisms and contexts to empirical patterns.” It’s science-as-usual. Pawson and Tilley know this – they acknowledge and cite earlier work. For example Pawson and Tilley (1997, p. 10) cite Palmer (1975, p. 150):

“Rather than ask, ‘What works for offenders as a whole?’ we must increasingly ask ‘Which methods work best for which types of offenders, and under what conditions or in what types of setting?'”

The link with science-as-usual is clear in Pawson and Tilley (2001, p. 324):

“We have argued that good evaluation is good social science. For us, this embraces the gallant aims of precision in articulation of theory, rigor in empirical testing, confederation in lines of inquiry, and cumulation in the body of findings. The ‘realist movement,’ of which we are a part, is often considered the brash upstart of the evaluation schools. In fact, it depends on these rather venerable ideas. The future, for us, thus lies in keeping faith with some of the grand old principles of social science and in not forgetting the hard-won lessons of the old studies.”

Evaluation method(ology) debates remind me of debates about psychological therapy brands, like CBT, psychodynamic, or humanistic approaches. Mick Power (2010) is an example therapist-academic who had a go at dismantling the brands, and pulled out graded exposure, transference, and challenging dysfunctional assumptions as example techniques that are used across a range of approaches. Others have done similar, e.g., the behavioural change technique taxonomy (Michie et al., 2013) extracts a long menu of approaches that can be combined to develop a programme.

I think evaluation would make faster progress if we routinely dismantled the big evaluation brands and asked of an apparently new and unique approach:

  1. Is it really unique?
  2. Who else has done similar?
  3. What is the approach an instance of?
  4. How could the approach be used in another type of evaluation?
  5. How could a technique be combined with another?

Westhorp and Feeny (2024) used this style of reasoning for realist evaluation, and showed how regression models with interaction terms can be used to test CMOs. Baumgartner and Falk (2023) explore regression-based alternatives to the Quine-McCluskey logical minimisation algorithm used in qualitative comparative analysis (QCA), closing the gap between QCA and statistics.

When it comes to questions of metaphysics, ask: what other approaches make the same or similar assumptions, e.g., concerning ontology and epistemology? For example, there’s a trace of the gap between mechanism and measure in questionnaire design: “operationalisation”, the process of developing a measure of, e.g., a psychological construct. This process begins with the premise that the phenomena are not directly observable. It’s psychometrics 101.

And this critical dismantling strategy applies to quantitative impact evaluation approaches too. Consider “synthetic controls” developed using the synthetic control method. They sounds unique and special; however, peek behind the curtain and they’re weighted averages, with weights estimated by balancing pre-intervention measurements of the outcome and covariates. So perhaps a comparison group developed using inverse probability of treatment weighting is also a “synthetic control”…? Maybe a t-test uses synthetic treatment and control groups, with unit weights…? Or consider entropy balancing, which, when targeting the ATT, is equivalent to inverse probability tilting (IPT; Słoczyński et al., 2025). The only difference is that entropy‑balancing weights sum to the comparison‑group sample size, whereas IPT weights sum to the intervention‑group sample size.

Build bridges between the methods and methodologies and they take up less space in your brain, making it easier to design robust evaluations and communicate clearly what you’re really setting out to do.

References

Aston, T. (2026). Evaluation blogs, podcasts, and webinars in the second half of 2025: Roundup review V. Evaluation.

Baumgartner, M., & Falk, C. (2023). Configurational Causal Modeling and Logic Regression. Multivariate Behavioral Research, 58(2), 292–310.

Michie, S., Richardson, M., Johnston, M., Abraham, C., Francis, J., Hardeman, W., Eccles, M. P., Cane, J., & Wood, C. E. (2013). The Behavior Change Technique Taxonomy (v1) of 93 Hierarchically Clustered Techniques: Building an International Consensus for the Reporting of Behavior Change Interventions. Annals of Behavioral Medicine, 46, 81–95.

Palmer, T. (1975). Martinson Revisited. Journal of Research in Crime and Delinquency, 12(2), 133–152.

Power, M. J. (2010). Emotion focussed cognitive therapy. John Wiley & Sons.

Słoczyński, T., Uysal, S. D., & Wooldridge, J. M. (2025). Covariate Balancing and the Equivalence of Weighting and Doubly Robust Estimators of Average Treatment Effects. IZA Institute of Labour Economics Discussion Paper, 18147.

Westhorp, G., & Feeny, S. (2025). Using surveys in realist evaluation. Evaluation Journal of Australasia, 25(1), 45–64.

Quasi-experiment sample size

Two papers caught my attention on sample‑size calculation: one for difference‑in‑differences designs (Hedberg & Hedges, 2026) and another for propensity score weighted designs (Liu et al., 2026).

Difference-in-differences (DiD)

Hedberg and Hedges (2026) have developed a method for estimating minimum detectable effect size (MDES) for DiD, including R code in an appendix. The approach takes account of variance explained by covariates (which could include fixed effects for units) and covariate imbalance. It also allows for unequal sample sizes. To run in reverse and estimate sample sizes, there’s a multitude of ways to get R to search across a range of sample sizes and minimise the difference between the achieved and target MDES.

Some things to bear in mind:

It handles both repeated cross-sectional and mixed between-within designs; for the latter, use fixed effects (so \(n-1\) covariates for \(n\) units) alongside an estimate of the variance explained. The parameters include the number of observations, not participants. So for a 2×2 design with pre and post measures from each participant, you will need to multiply the number of participants by 2.

There’s an error in the published code:

if (rho1 > 1) {
  V <- V*(1/(1-rho1^2))
}
if (rho2 > 1) {
  V <- V*(1-rho2^2)
}

Those rhos should never be over 1, and earlier checks won’t allow them to be. The code should check if they are over zero:

if (rho1 > 0) {
  V <- V*(1/(1-rho1^2))
}
if (rho2 > 0) {
  V <- V*(1-rho2^2)
}

I spotted this during testing: increasing the variance explained by covariates did not affect the MDES, and the results differed from the simulation checks. The corrected version resolves this, and the authors very kindly and promptly confirmed the fix. I understand that updated code will be available on the journal’s website soon. In the meantime, the edits noted above should address the issue.

Propensity score weighting

I had been using Austin’s (2021) work to approximate VIFs from c-statistics of propensity score models, which represent lack of overlap in propensity score distributions between groups. The resulting VIFs can then be used to inflate sample sizes estimated from standard power calculators.

Liu, Yang, and Li (2026) have developed an alternative approach that does the sample size estimation in one go, depending on prevalence of the intervention, and an overlap parameter, \(\phi\), which is 1 with perfect overlap and decreases as the propensity score distributions separate. The approach has been implemented in the R package PSpower, which is on CRAN.

It’s speedy. Here’s an example with intervention prevalence at 50%, MDES 0.2, and varying the overlap from 80% to 99%, for the ATE, ATT, and ATO estimands. ATO is least affected by lack of overlap because it deliberately defines the estimand on the region where there is most overlap and down-weights observations that are far from there. ATO has a policy interpretation too: impact for people in the region of decisional equipoise.

References

Austin, P. C. (2021). Informing power and sample size calculations when using inverse probability of treatment weighting using the propensity score. Statistics in Medicine, 2, 1–14.

Liu, B., Yang, C., & Li, F. (2026). Sample size and power calculations for causal inference of observational studies (Version 5). arXiv, to appear in Annals of Statistics.

Hedberg, E. C., & Hedges, L. V. (2026). Computing Statistical Power for the Difference in Differences Design. Evaluation Review, 50(1), 149–180.

Invest in your staff

Relevant in the age of over-reliance on AI, particularly LLMs:

“I have always said you have to invest in technology but do not rely on it, but you should invest in your staff and you can always rely on them.”
– Tim O’Toole, Managing Director of London Underground (2003–2009), reflecting on the heroic staff response to the 7/7 attacks amid failed radio links [source]

Bending evaluation timelines to meet policy deadlines

There is a simple formula for delivering evaluation findings in time to inform policymaking: \(T = S + R + D + I + A\).

Suppose that, for a programme to be considered effective, it must demonstrate a practically significant causal impact at least \(I\) months after the end of programme activities.

The programme duration is \(D\) months from the point at which a participant is enrolled in the evaluation.

Given recruitment rates, participants are enrolled over a period of \(R\) months to ensure sufficient numbers to achieve the required statistical power. The final participant therefore completes all programme activities \(R + D\) months after recruitment begins.

Study setup takes \(S\) months. This includes developing a theory of change, study design, developing surveys and topic guides, finalising the evaluation protocol, ethical approval, and recruiting study sites.

Analysis and reporting takes \(A\) months. This includes time waiting for administrative data requests to be fulfilled or for fieldwork to conclude, data management, statistical disclosure control and output clearance, peer review, and revisions.

The total study time is therefore \(T = S + R + D + I + A\) months. (You could add more steps, but it wrecks the anagram.)

For example, suppose the timeline is:

Setup3 months
Recruitment period6 months
Duration of the programme3 months
Impact time point6 months
Analysis and reporting3 months

The total duration \(T = 21\) months. So start the study 1 June 2026, it will report by 1 March 2028. And this evaluation only estimates impact at 6 months.

Suppose you want findings earlier than this as you need to inform an urgent policy decision. What could you do? Some options:

1. Analyse existing data for a similar programme, e.g., using meta-analysis of studies in a systematic review or quasi-experimental analysis of admin data. You can also gather implementation and process evidence to evaluate whether the current programme is being implemented correctly, even if you don’t have time to estimate impacts (use the review or secondary analysis of earlier similar programmes for that).

2. Investigate shorter-term outcomes, relying on existing evidence from similar programmes to argue that effects will endure. For example, it could be that you are really interested in impact at 10 years; however, the literature suggests that an outcome you can measure at 1 year is strongly correlated with the longer-term impact and the theory of change suggests it is an important link in the causal chain.

3. Shorten the programme duration, risking smaller impact and an underpowered study, so you may also need to boost the sample size, e.g., by lengthening the recruitment period.

4. Increase recruitment rates, e.g., by including more study sites or contacting more potential participants (depending on the nature of the programme).

5. Shorten the recruitment period, risking an underpowered study so you will likely have a wide confidence interval that straddles zero. Light a candle to guide interpretation.

6. Combine 4 and 5 to achieve the required sample size faster.

7. Run a simpler study, reducing setup, analysis, and reporting time. For example, you might measure only one or two key outcomes and conduct highly focused implementation and process evaluation.

8. Combine analysis of existing data with running a new evaluation: use option 1 (analyse existing data) to inform your policy design now, and also run the \(T\)-month evaluation to help out your future policy colleagues when they review the literature in a few years’ time.

9. Use time and relative dimensions in space-adjusted analysis to deliver the originally planned evaluation in time to feed into the policy decisions.

A policy evaluator about to embark on a three year evaluation that needs to report its findings in a few weeks’ time

Option 9 seems to be what policymakers expect. My preference would be to accept our spacetime structure and go for option 8 where possible (with the help of 2) or 1 if not. In any case, ensuring that evaluations can complete in time to inform policy requires long-term planning that is resilient to changes of PM, Ministers, and governments.

Thanks Edisa for inspiring a mild but crucial tweak to the variable names!

“The body does not keep the score”

Interesting example of theorising in paper by Kotler et al. (2026), “The body does not keep the score”. I haven’t read van der Kolk’s (2014) classic The Body Keeps the Score, which they critique, but I don’t think that critique is central to their claims. I also don’t think their analysis needs Friston’s “free energy principle”.

Key quotes:

“… trauma over-weights the precision of danger priors: the brain assigns excessive confidence to threat predictions, constraining inference based on the prior premise of enduring danger. The result is hypervigilance, flashbacks, and avoidance—symptoms of a system caught in self-confirming predictions.”

“Internal threat expectations dominate the search for—and attention to—sensory evidence of danger; unattenuated interoceptive signals (racing heart, tight chest, etc.) are interpreted as confirmation of danger rather than imprecise noise. The ‘score’ the body appears to keep is thus an artifact of circular inference: the brain predicts pain, senses arousal, and takes that arousal as proof that pain persists. The body participates in trauma, but as messenger, not archive.”

“If the old story held that ‘the body keeps the score,’ the emerging narrative elides somatic chauvinism. It’s subtler, and more hopeful. The body does not keep the score; the brain keeps predicting it. When prediction becomes too rigid, experience repeats itself, not because it is stored, but because it cannot yet be reinterpreted. Flow—and other states that expand metastability offer the nervous system a chance to update its model of the world, to reassign precision where it belongs, and to rediscover safety in uncertainty.”

Flow is often attributed to Mihály Csíkszentmihályi, and is the name for the feeling of absorption in an activity that people tend to experience when they perceive that

“the environment contains high enough opportunities for action (or challenges), which are matched with the person’s own capacities to act (or skills). When both challenges and skills are high, the person is not only enjoying the moment, but is also stretching his or her capabilities with the likelihood of learning new skills and increasing self-esteem and personal complexity.” (Csíkszentmihályi & LeFevre, 1989, p. 816)

An example of when people frequently experience this flow state is when driving. There’s a lot to unpack in the “metastability” aspect of the theory. They describe this in terms of neural states, being able to “fluidly switch among semi-stable network states” (Kotler et al., 2026, p. 1), but their examples seem to signal that they mean being flexible in responses to situations, e.g., not only perceiving potential threats.

References

Csíkszentmihályi, M., & LeFevre, J. (1989). Optimal experience in work and leisure. Journal of Personality and Social Psychology, 56, 815–822.

Kotler, S., Mannino, M., Fox, G., & Friston, K. (2026). The body does not keep the score: Trauma, predictive coding, and the restoration of metastability. Frontiers in Systems Neuroscience, 20, 1812957. https://doi.org/10.3389/fnsys.2026.1812957

van der Kolk, B. A. (2014). The Body Keeps the Score: Brain, Mind, and Body in the Healing of Trauma. Viking Press.