Overview
News
Technologies
Salaries
Products
People
Growth
Financials

Overview

Open-world evaluations for measuring frontier AI capabilities. The case for long, messy, real-world tasks to evaluate AI agents.

News

News CRUX 2 months ago
Can AI agents conduct open-ended AI research? Early evidence from two case studies
We gave frontier agents research questions from two non-public NeurIPS submissions and had the original authors grade the results. The papers produced by the agents were unambiguously rejected.
Read more
Report
Blog 2 months ago
Update on our long-horizon AI R&D evaluations
CRUX 2 is testing whether AI agents can answer novel, open-ended AI research questions.
Read more
Report