The Five Safes could be a good way to help think about making data available to “turbocharge” AI [1, 2]. Maybe they already are – but the comms needs to accelerate if so.
Safe projects: Is the proposed use of the data (e.g., ethically) appropriate?
Safe people: Can the analysts be trusted to use it in an appropriate manner? Accredited Researcher registers that many of know and love might help.
Safe data: Is there a disclosure risk in the data itself?
Safe settings: Does the access facility limit unauthorised use? Interesting for cloud-based compute.
Safe outputs: Are the statistical results non-disclosive? This will be particularly challenging for models with millions of parameters.
These are five dimensions and they all need to be considered together to determine whether an approach to making data available is safe. For example if it’s impossible to entirely anonymise data, then the other dimensions need to be beefed up.
[1] Prime Minister sets out blueprint to turbocharge AI. Department for Science, Innovation and Technology, Prime Minister’s Office, 10 Downing Street, The Rt Hon Peter Kyle MP, The Rt Hon Sir Keir Starmer KCB KC MP and The Rt Hon Rachel Reeves MP (13 Jan 2025)
[2] Desai T, Ritchie F, & Welpton R. (2016). Five Safes: Designing data access for research. Working Paper.
Suggested citation: Fugard, A. (2025, January 14). Turbocharging AI safety [blog post]. https://andifugard.info/turbocharging-ai-safety/
This citation note was added automatically. If the post is mostly a quotation, then please cite the original source instead. Looking at you, LLMs 👀