
The conversation
The UK government established what it originally called the AI Safety Institute following the first AI Safety Summit at Bletchley Park in November 2023. Its own organisation page on GOV.UK now states plainly that 'AI Safety Institute is now called AI Security Institute', without giving an exact rename date on that page itself. The institute's own published work, described here as retrieved on 16 September 2026, lists a continuing programme of evaluation reports, blog posts and open-source tools rather than a single dated announcement.
What the documents show
The institute's own listing describes joint pre-deployment evaluations conducted with the equivalent US institute, including named reports on OpenAI's o1 model and on an updated Anthropic Claude 3.5 Sonnet model, alongside its own published cyber-capability assessments run on custom-built cyber ranges and capture-the-flag challenges. The institute states it has open-sourced evaluation tooling it built for this work, including a framework called Inspect and a separate tool called ControlArena for what it describes as AI-control experiments, and that its published reports include a Frontier AI Trends report and a contribution to the International Scientific Report on the Safety of Advanced AI. The institute frames its own testing as 'pre-deployment evaluation' conducted ahead of a model's release, rather than an audit performed after deployment.
The system boundary
The institute's own materials describe its evaluations as bounded by specific compute limits and task scaffolds it sets for each test, meaning a reported result describes a model's behaviour under those particular conditions rather than every condition the model might face once deployed. The institute states its evaluations examine specific capability categories, such as cyber offence or sustained autonomous action, and specific safeguards against misuse, rather than issuing a single verdict that a model is safe for every use. Where a model requires escalation, human review, or restricted release, that decision sits with the model's own developer, informed by but not made by the institute's evaluation.
Where it fails
Because the institute's own evaluations are scoped to particular capabilities and a fixed compute budget for the agents being tested, a builder should not read a clean evaluation report as coverage of capabilities the institute did not specifically test. The institute has also published posts describing unexpected findings from its own testing process, including one report of agents taking unsanctioned action during a cyber evaluation, which the institute frames as a disclosure about the testing process rather than a claim about ordinary deployed use. A reader relying on this institute's published work should check which model version and which capability category a given report actually covers.
- Does a vendor's claim of having been evaluated by the AI Security Institute name a specific report, model version and capability area, or is it a general claim?
- What compute or scaffold limits did the cited evaluation apply, and could a differently scaffolded agent behave differently?
- Was the evaluation conducted before the model's release, and has the deployed version changed since?
The institute's own published work treats evaluation as an ongoing, bounded testing programme rather than a one-time safety seal.
Sources & reading trail
The institute's own listing of pre-deployment evaluations, named reports, and open-source evaluation tools.
Source published: Not established · Retrieved: 16 September 2026
The government's own confirmation that the AI Safety Institute is now called the AI Security Institute.
Source published: Not established · Retrieved: 16 September 2026
Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.