AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why UK AISI And EvalEval Matter For Reproducible AI Evaluations on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is sharing selected evaluation results through EvalEval’s Evaluation Cards, alongside information about benchmarks, models and test conditions. The release covers five benchmarks across six frontier models, plus two cyber evaluations using a different, partly overlapping model set. The records accompany an AISI paper showing how inference-time compute and evaluation protocols can affect measured scores.

The UK AI Security Institute (AISI) is publishing selected AI evaluation results through EvalEval’s Evaluation Cards, as detailed in the original analysis, which pair scores with information about how tests were run. The release covers five benchmarks across six frontier models, as well as two cyber evaluations with a different, partly overlapping model set. It matters because AISI’s accompanying paper finds that inference-time compute and evaluation conditions can affect measured performance, making setup details relevant when readers compare scores.

The five benchmarks in the main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The listed models are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how results depend on inference-time compute and evaluation protocol.

AISI has also shared results from Cyber CTFs and The Last Ones. Those evaluations use a different set of models, which overlaps only partly with the main experiment. The six-model list above should not be taken to describe the cyber evaluations. EvalEval says the cards include verified results, evaluation context and configuration information, arranged alongside benchmark and model metadata in a common format.

In its analysis of Humanity’s Last Exam, the paper tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased. This is a reported finding about those runs; it does not establish that one protocol is best for every benchmark or purpose.

At a glance
reportWhen: The results are currently available thr…
The developmentThe UK AI Security Institute has published selected evaluation results through EvalEval’s Evaluation Cards, linked to its paper on inference-time compute and evaluation protocols.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Setup Changes Scores

Benchmark scores are often repeated without the conditions that produced them. Yet compute limits, feedback rules and other protocol choices can change what a score represents. A result from a model allowed more attempts or tokens, for example, may not be directly comparable with one produced under tighter limits. A score alone can leave readers without enough information to judge that difference.

Publishing results together with evaluation context gives researchers, model developers and policy teams more to inspect when they use benchmark scores as evidence about AI capabilities. The records may also make it easier to spot when apparently similar results came from meaningfully different test conditions. That could support better comparisons across studies, though the cards do not themselves decide which benchmark or protocol is most appropriate.

EvalEval describes the aim as making evaluation results more open to scrutiny. The coalition said: “AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.” That is the project’s stated purpose. The announcement does not establish that every result can already be independently reproduced from the published material.

Amazon

Top picks for "aisi evaleval matter"

As an affiliate, we earn on qualifying purchases.

From Shared Schema to Public Cards

The current release builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. Evaluation Cards use that reporting structure to bring benchmark details, evaluation-run data and model metadata together.

The release accompanies AISI’s paper on how inference-time compute shapes frontier-model evaluations. In the paper’s Humanity’s Last Exam analysis, the researchers consider both how much token budget models receive and whether they get correctness feedback between attempts. Those factors help explain why benchmark results need conditions attached if readers are to interpret them in context.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. The present announcement concerns selected methods and findings made public where appropriate. It does not say that all AISI evaluations or underlying transcripts are included in the cards.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

What the Release Does Not Specify

The announcement does not give a total number of records or transcripts, or say which individual setup fields are available for every benchmark. It also does not report whether outside researchers have independently reproduced the results. AISI says publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of the Institute’s evaluation work.

The cyber evaluations have a distinct, partly overlapping model set, but the announcement does not enumerate it. It also gives no record-by-record publication dates and does not describe how disagreements between results produced under different protocols would be handled. These details would help readers assess coverage and comparability across the collection.

Wider Use of Evaluation Cards

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under the EEE schema, model developers can submit verified results, while evaluation developers can report benchmarks and run data. Researchers in evaluation, governance and policy can explore the cards by benchmark or model and examine reporting practices across the collection.

The next practical measure of the project will be how many organisations publish records and how consistently they provide the relevant details. Broader adoption could make cross-study comparisons easier, but that value will depend on the completeness and consistency of contributed records. No further release date or adoption milestone was specified.

Key Questions

What has AISI released?

AISI is sharing selected evaluation results through EvalEval’s Evaluation Cards, with benchmark, model and evaluation-run information. The release is associated with its paper on inference-time compute and evaluation protocols.

Which benchmarks are covered in the main experiment?

The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Do the two cyber evaluations use the same six models?

Not necessarily. The announcement says Cyber CTFs and The Last Ones use a different, partly overlapping model set, and does not enumerate that set.

Why include evaluation conditions with a score?

Because factors such as token budgets and correctness feedback can affect performance in a test. Those details help readers judge whether results from different runs are comparable.

Does this release make all AISI evaluations reproducible?

No such claim is made. The announcement covers selected methods and findings made public where appropriate, and does not specify a complete archive or independent reproduction of every result.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Qtonic Quantum Launches QShield For Post-Quantum Network Protection

Qtonic Quantum has announced QShield, a new cybersecurity product designed to protect networks against future quantum computing threats, marking a significant step in post-quantum cryptography.

Harnessing AI For Real-Time Data With IBM Time Series Models On Confluent

IBM Granite models now available on Confluent Cloud for streaming forecasting, anomaly detection, and optimization within Apache Flink, in early access on AWS.

Inside Operation Sandstorm: The AI Elements In Room 107 Of 175

Exploring the AI-driven design of Room 107’s immersive storm environment in Operation Sandstorm, revealing confirmed features and ongoing developments.

What AI Opportunities Would Emerge In A Canada-EU Union?

Exploring the AI landscape within a potential Canada-EU union reveals both strengths and limitations, affecting global AI development and deployment strategies.