Evaluation is often divided into two kinds: formative evaluation, which is conducted to support the improvement of a programme, and summative evaluation, which is conducted to make (often high-stakes) decisions about a programme, such as whether to expand its delivery so it reaches more people or discontinue it.
Huey Chen (1996) challenged this dichotomy. The issue is that formative evaluation can lead to summative conclusions, e.g., the decision not to continue with a fully-powered impact evaluation. Conversely, summative evaluation can lead to formative improvements that can be tested in the next summative trial.
A related distinction, better understood as a spectrum, concerns the stance evaluators adopt before setting off on an evaluation. The following two David Shrigley pieces may offer a way to distil the extremes at either end of this spectrum:

A illustrates puff-piece evaluation: it is compromised by conflict of interest and often conducted by those who devised or deliver the programme under evaluation. If an organisation exists to deliver a programme, then evaluators within that organisation are going to find it challenging to evaluate themselves out of a job. I have seen this kind in the wild, for example evaluations of apps conducted by the company providing them. Don’t mention AI.
B illustrates adversarial evaluation: it’s assumed from the outset that the programme has no impact or causes harm. These are usually independent evaluations. The examples that come to mind are evaluations of interventions that persist despite a long history of evidence that they cannot work and the aim of the evaluation is to finally make a compelling case that funding for them should cease.
Reality is usually more nuanced and evaluations land somewhere in between. Programmes sometimes work, but maybe not as well as we hoped or not for everyone we thought would benefit. Still, I think identifying your location on this spectrum before you begin an evaluation can help clarify what safeguards are needed to support its objectivity. What do you believe the findings will be and why?
The various replication crises expose how we mere mortals are individually biased. We achieve objectivity collectively, through social systems such as peer review and pre-registered protocols. If users of an evaluation, whether potential programme beneficiaries or policy teams, suspect unmanaged bias, then it will be difficult to persuade them to trust what you find.
References
Chen, H. (1996). Typology for Program Evaluation. Evaluation Practice, 17(2), 121–130.
Suggested citation: Fugard, A. (2025, August 3). Extremes along an evaluation spectrum [blog post]. https://andifugard.info/extremes-along-an-evaluation-spectrum/
This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀