Oneirix

The thinking behind the testing

Research
& insights.

Original research. Practical questions. Evidence that informs our evaluations.

Our work includes research published at NeurIPS, proprietary benchmarks and investigation of model computation during training and inference.

Perspectives on AI evaluation

A benchmark score is only part of the picture.

A score summarises performance on a set of tasks. It tells us how a model performed under those conditions, while leaving questions about how it reached its answers and how that behaviour changes during computation.

Our proprietary benchmarks sit alongside measurements of internal computation during training and inference, where model access permits. These complementary views help us investigate failure patterns beyond the final output.

We connect each finding to the tasks, conditions and access used in the investigation. That scope gives the evidence its meaning and makes its limits clear.

Explore model evaluation