Actively hiring
Publishing
Digital Media
Follow
Overview
News
Salaries
Products
People
Growth
Offices
Financials
News
MiniMax M3 vs GLM 5.2
A workbench report comparing MiniMax M3 and GLM 5.2 as autonomous coding workers across scored and observed Python coding tasks.
Read more
Report
Defensible Deep Research with Open-Weight Models
Low-cost open-weight passes read and draft from fetched sources; a frontier coordinator checks claims and confidence tiers show where the evidence is strong or thin.
Read more
Report
State of the Agent - Claude Code Agent Census
Interactive census of 375 Claude Code agent primitives across 69 open-source plugins, measuring domain coverage, overlap, boundaries, and uncertainty guidance.
Read more
Report
agent-evals - LLM Coding Agent Evaluation
Evaluate LLM coding agents with overlap analysis, boundary testing, and metacognitive scoring for Claude Code, Cursor, Copilot, GPT, and Gemini.
Read more
Report
