Overview
News
Salaries
Products
People
Growth
Offices
Financials

Overview

Humanity's Last Exam. Long Phan *1, Alice Gatti *1, Ziwen Han *2, Nathaniel Li *1.

News

News Humanity's Last Exam 11 days ago
Introducing HLE-Diamond
A refined subset of Humanity's Last Exam, following a year-long process of cleaning and refinement with input from research communities.
Read more
Report
News Nature 9 months ago
A benchmark of expert-level academic questions to assess AI capabilities
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve more than 90% accuracy on popular benchmarks such as Measuring Massive
Read more
Report
News developers.googleblog.com 1 year ago
Start building with Gemini 2.5 Flash- Google Developers Blog
Gemini 2.5 Flash, is now in preview, offering improved reasoning while prioritizing speed and cost efficiency for developers.
Read more
Report
News nytimes.com 1 year ago
The New York Times When A.I. Passes This Test, Look Out The creators of a new test called 'Humanity's Last Exam' argue we may soon lose the ability to create tests hard enough for A.I. models.
The creators of a new test called 'Humanity's Last Exam' argue we may soon lose the ability to create tests hard enough for A.I. models.
Read more
Report
Introducing Claude Sonnet 4.5
Claude Sonnet 4.5 is the best coding model in the world, strongest model for building complex agents, and best model at using computers.
Read more
Report
News api-docs.deepseek.com
DeepSeek-R1 Release DeepSeek API Docs
Performance on par with OpenAI-o1
Read more
Report