Actively hiring
Follow
Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
Open-world evaluations for measuring frontier AI capabilities. The case for long, messy, real-world tasks to evaluate AI agents.
News
Can AI agents conduct open-ended AI research? Early evidence from two case studies
We gave frontier agents research questions from two non-public NeurIPS submissions and had the original authors grade the results. The papers produced by the agents were unambiguously rejected.
Read more
Report
Update on our long-horizon AI R&D evaluations
CRUX 2 is testing whether AI agents can answer novel, open-ended AI research questions.
Read more
Report
