Actively hiring
Artificial Intelligence
As featured in
Follow
Overview
News
Salaries
Products
People
Growth
Offices
Financials
News
Introducing HLE-Diamond
A refined subset of Humanity's Last Exam, following a year-long process of cleaning and refinement with input from research communities.
Read more
Report
A benchmark of expert-level academic questions to assess AI capabilities
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve more than 90% accuracy on popular benchmarks such as Measuring Massive
Read more
Report
Start building with Gemini 2.5 Flash- Google Developers Blog
Gemini 2.5 Flash, is now in preview, offering improved reasoning while prioritizing speed and cost efficiency for developers.
Read more
Report
The New York Times When A.I. Passes This Test, Look Out The creators of a new test called 'Humanity's Last Exam' argue we may soon lose the ability to create tests hard enough for A.I. models.
The creators of a new test called 'Humanity's Last Exam' argue we may soon lose the ability to create tests hard enough for A.I. models.
Read more
Report
Reuters AI experts ready 'Humanity's Last Exam' to stump powerful tech A team of technology experts issued a global call on Monday seeking the toughest questions to pose to artificial intelligence systems, which increasingly have handled popular benchmark tests like child's play.
The creators of a new test called 'Humanity's Last Exam' argue we may soon lose the ability to create tests hard enough for A.I. models.
Read more
Report
Introducing Claude Sonnet 4.5
Claude Sonnet 4.5 is the best coding model in the world, strongest model for building complex agents, and best model at using computers.
Read more
Report
Report
