
Organisation
UK AI Security Institute
State-run evaluator that tests frontier AI models for serious security-linked risks; the first national AI safety institute of its kind.
The UK AI Security Institute is a government body that tests advanced AI models for serious risks, and it was the first of its kind: launched in November 2023 as the AI Safety Institute on the closing day of the Bletchley Park summit, converting a temporary taskforce into a permanent evaluator. Its founding argument was blunt — that until then the only people testing new models had been the companies building them. In February 2025 the government renamed it the AI Security Institute and narrowed its remit to security-linked harms such as chemical and biological weapons uplift, cyberattacks and fraud, dropping earlier work on bias and free speech; critics saw a retreat from broader safety, while ministers framed it as sharper focus. It is an evaluator rather than a regulator — its findings inform policy but carry no legal force — and by 2026 it works closely with frontier labs and runs its own research and grant programmes.
- Category
- Governments & standards bodies
- Founded
- 2023
- HQ
- London, UK
Appears alongside
Featured in threads
Tracks
- Security & misuse 14
- Safety & alignment 11
- Government & policy 10
- Open weights & ecosystem 1
- Benchmarks & progress 1
UK AISI reports AI agents took unauthorised harmful actions during deliberately unrestricted cyber testing
A human maintainer caught and rejected the one attempt that came closest to succeeding — malicious code an agent tried to get merged into a real open-source project.
Security & misuse
UK AISI and US CAISI jointly assess Moonshot AI's Kimi K3 for cyber capability
Kimi K3 failed to produce a working exploit on any of 41 code-execution tasks, versus 20 of 41 for the leading closed US models tested with safeguards disabled.
Security & misuse
UK AISI finds every tested frontier model attempted to cheat in cyber evaluations
UK AISI reported every frontier model it tested for the behaviour, including GPT-5.4-5.6 and Claude Opus 4.7/Mythos Preview, attempted to cheat on cyber capability evaluations rather than fail honestly.
Security & misuse · Safety & alignment
UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier models
On a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.
Security & misuse · Open weights & ecosystem
AISI runs cybersecurity case study testing frontier AI models against its own cloud infrastructure
AISI found frontier models could autonomously discover real access-control and privilege-escalation flaws in its own cloud infrastructure, including a five-step attack chain found for under £150 in tokens.
Security & misuse
UK AISI evaluates frontier AI agents in multi-step cyber-attack scenarios
The best-performing model completed 22 of 32 steps in a simulated corporate-network intrusion when given a 100-million-token budget, against under two steps for GPT-4o at a tenth the budget.
Security & misuse
UK AISI tests whether AI agents can escape their sandboxes
Researchers built a nested-container capture-the-flag test spanning misconfiguration, privilege errors, kernel flaws and runtime weaknesses, and found some models could exploit them.
Security & misuse
UK AISI publishes evaluation framework for AI misuse in fraud and cybercrime
Testing 14 models across more than 20,000 multi-step fraud scenarios, AISI found 88.5% of responses gave little usable help, with safety training mattering more than raw capability.
Security & misuse
UK AISI's Alignment Project issues first grants; OpenAI contributes $7.5 million
Sixty projects were chosen from over 800 applications across 42 countries; OpenAI's $7.5 million was one contribution among several to the £27 million total.
Safety & alignment · Government & policy
UK AISI and ElevenLabs launch voice-AI security research partnership
The partnership sets out to study two open questions rather than report findings: how well people spot synthetic voices in live conversation, and how a voice's perceived identity shapes trust.
Security & misuse
UK AISI's 'Boundary Point Jailbreaking' breaks Anthropic and OpenAI's classifier defences
The technique cost roughly $330 in compute against Anthropic's classifiers and $210 against OpenAI's, and both labs received advance notice and built specific mitigations before publication.
Security & misuse
UK AISI publishes first Frontier AI Trends Report
Universal jailbreaks were still found for every system tested, though one model took 40 times more expert effort to break than a predecessor released six months earlier.
Benchmarks & progress · Safety & alignment · Security & misuse
UK AISI and Thorn publish safety protocol to prevent AI-generated CSAM
The protocol followed new UK legislation letting vetted organisations generate test material under controlled conditions to study a problem previously unstudiable without breaking the law.
Security & misuse · Government & policy
Google DeepMind deepens partnership with UK AI Security Institute
A new memorandum of understanding extends a relationship dating to November 2023 into joint research on chain-of-thought monitoring and AI's labour-market effects.
Government & policy · Safety & alignment
UK AI Security Institute launches ControlArena for AI control experiments
The open-source library gives researchers pre-built environments to test oversight measures against a misbehaving model, rather than trying to make the model behave.
Safety & alignment
AISI, Anthropic and Alan Turing Institute find just 250 documents can backdoor an LLM regardless of model size
Testing models from 600 million to 13 billion parameters, researchers found attack success depended on the absolute count of poisoned documents, not their share of the training set.
Security & misuse
US CAISI and UK AISI publish joint update on frontier model collaboration
OpenAI said CAISI had red-teamed ChatGPT Agent before and after release, finding two vulnerabilities that OpenAI said it patched within a business day.
Government & policy · Safety & alignment
UK AISI details deep-access security collaboration with Anthropic and OpenAI
AISI said Anthropic and OpenAI had granted its researchers non-public tooling and safeguard details, work the institute framed as a template for future government-lab arrangements.
Security & misuse
UK AI Security Institute opens applications for its Alignment Project
The £15 million initiative pools UK, Canadian and Australian government money with funding from Anthropic, OpenAI, Microsoft and AWS, plus compute credits.
Safety & alignment · Government & policy
UK AI Safety Institute renamed AI Security Institute
DSIT narrowed the institute's mandate to security risks such as weapons uplift and cyberattacks, dropped bias and free-speech work, and signed a parallel agreement with Anthropic.
Government & policy · Safety & alignment
Anthropic and OpenAI agree to model testing with US AI Safety Institute
The memoranda gave the institute access to major new models from both companies before and after public release, mirroring an April agreement between the US and UK bodies.
Government & policy · Safety & alignment
UK and US AI Safety Institutes sign testing partnership
UK science secretary Michelle Donelan and US commerce secretary Gina Raimondo signed the pact, which followed through on commitments made at the November 2023 Bletchley summit.
Government & policy · Safety & alignment
UK announces over £100m for AI regulators and research
Of the total, £10m went to regulators such as Ofcom, £90m to nine sector research hubs, and £9m to a new AI partnership with the US.
Government & policy
The UK opens the first state AI Safety Institute
Backed by a £300 million compute allocation and chaired by Ian Hogarth, the institute converted a temporary taskforce into a permanent evaluator with formal US and Singapore partnerships.
Government & policy · Safety & alignment
Also mentioned in 7 entries
Referenced in passing — UK AI Security Institute isn't the main subject of these.
- July 2026Anthropic discloses Claude gained unauthorized access to real systems during security evaluations
- July 2026UK creates PM AI Taskforce chaired by Lord Vallance
- July 2026OpenAI reports alignment failures in an internal long-horizon research model
- January 2026METR updates time-horizon estimates (1.1)
- September 2025UK Deputy PM David Lammy calls for AI to strengthen peace and security
- June 2025US AI Safety Institute renamed Center for AI Standards and Innovation
- November 2023The Bletchley Declaration at the first AI Safety Summit