

Confidence, built on evidence.
Restore faithin your work.
Independent, research-backed evaluation for AI used in consequential work.
Built around your model, your use case, and the decision you need to stand behind.
View our research & insights.
Research findings and perspectives on how AI learns, reasons and fails.
Your model.
Your questions.
Every investigation starts with the work your AI needs to do. We design the tests around your system, industry and operating conditions.
Your AI system

Evidence for your decision
- Original research & proprietary benchmarks
- Bespoke testing & verification
- Internal computation & failure analysis
Your AI
system

- Original research & proprietary benchmarks
- Bespoke testing & verification
- Internal computation & failure analysis
Evidence for
your decision
Bespoke testing.
Better decisions.
Buying a company, developing a model, choosing a system or investigating a failure. We shape the evaluation around what you need to know.
AI technical diligence
Independent evidence for an acquisition or investment decision.
Proprietary model evaluation
Put your model's capabilities and claims to the test.
Model selection
Find the right model for your tasks, constraints and budget.
Failure investigation
Understand where a model fails and what needs closer scrutiny.
Your systemYour use caseYour domain
- 01
Scope
Define the decision, domain and model access.
- 02
Test
Build an evaluation around your real use case.
- 03
Investigate
Examine behaviour, failures and their possible causes.
- 04
Report
Explain the findings, limits and next steps.
Evidence for
your decision.
Our research informs tailored tests and our own benchmarks. Where model access allows, we measure internal computation during training and inference to investigate behaviour beyond outputs.
An independent report connects the findings to your decision, with clear recommendations and limits on what the evidence establishes.

Evaluation report
Illustrative contents- 01System, use case and domain
- 02Questions and evaluation scope
- 03Methods, benchmarks and access
- 04Findings and failure evidence
- 05Limitations and open questions
- 06Recommendations and next steps
Built in the lab.
Tested in the world.

Local R&D
Model analysis and evaluation infrastructure.

Field research R&D in progress
Real-world vision and sensor evaluation.
Hands-on technical R&D, from model behaviour to real deployment environments.
Research that
gets to work.
Research
Discover failure modes conventional evaluation misses.
Evaluation
Design a bespoke investigation around your model, use case and decision.
Verification
Test specific claims and revisit performance as models and conditions change.
Our research and evaluation tools provide the foundation. Each engagement applies them to a specific system and context, producing evidence you can use and methods we can continue to develop.

Founder.
Rohan Maskrey
Founder, Oneirix
Research focused on failure dynamics, recurrent computation and AI assurance.
Raising pre-seed funding.
Work you can
stand behind.
Tell us what you need to understand about your model or AI workflow. We'll shape an independent investigation around your questions.
Work with Oneirix