434,399open jobs
14,946companies
65,229added this week
Browse all
Salary
$165k – $357k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
MultiOn builds artificial intelligence agents that browse and act on the web for users. Its API automates bookings, purchases and research tasks. The company focuses on reliable multi step web actions.

Think Different. Build the Future.

Our Mission

Build everyday AGI. Trustworthy, consumer-grade agents that redefine human-AI collaboration for millions. Software shouldn’t wait for commands; it should partner with you, amplifying what you can do every single day.

Why AGI, Inc.

We’re a stealth team of elite founders and AI researchers, with backgrounds spanning Stanford, OpenAI, and DeepMind. We’re industry leaders in mobile and computer-use agents, bringing these capabilities to consumer scale.

Grounded in years of agent research, our AI is designed with trustworthiness and reliability as core pillars, not afterthoughts.

We are supported by tier-1 investors who funded the first generation of AI giants; now they’re backing us to build the next: everyday AGI. (Watch the demo)

If you see possibility where others see limits, read on.

You decide what "better" means.

Models, agents, and product features all ship behind one question: did this actually get better? Without a strong evals function, the lab ships vibes. With one, every training run, every prompt change, every agent capability moves a number we trust - and the team makes decisions on real signal, not the loudest opinion in the room.

You'll build the eval harness for AGI - across model capability, agentic behavior, on-device performance, and end-user experience. You'll set the bar for what counts as "shipped" and protect it from the gravity of product deadlines.

Tasks you will own

  • The eval suites that gate every model and agent release - capability, behavior, regressions, and human-rated rubrics that catch what automated evals miss

  • The dashboards and tooling that make researcher experiment loops fast and leadership decisions easy

  • The bar - what counts as ready to ship, and how we know

Areas where you will assist

  • Research, by making sure what we measure is what we want

  • Product engineers, by instrumenting real-user behavior on real devices

  • Partnerships, by translating "did it get better" into language an OEM partner can hold us to

Skills you'll be expected to teach

  • How to measure non-deterministic systems - agent eval, tool use, long-horizon tasks, multilingual behavior

  • How to push back on a metric that's being gamed without breaking the team

Skills you'll be expected to learn

  • On-device perf trade-offs and how they show up in real-user evals

  • What QA-ing AI at OEM scale actually looks like

  • The realities of shipping consumer agents to production partners

Timeline of success

After 30 days - You've audited every eval we run today and produced a sharp doc on what's good, what's noise, and what's missing. You've fixed the most embarrassing gap.

After 60 days - You've stood up a new eval surface - agentic, on-device, or behavioral - and the team is making real decisions on its output. Researchers come to you before launching a run, not after.

After 90 days - Releases now ship against your eval bar, not a vibe-check. You've caught a regression that would have shipped, and cleared a launch the team was nervous about. You're shaping the research roadmap by surfacing where we're flat, where we're climbing, and where we're lying to ourselves.

Compensation

Competitive cash and meaningful equity. Top-tier relocation and immigration support. SF, in person.

How to apply

Send a link to an eval, benchmark, or measurement system you built - and one paragraph on what decision it changed. Plus your resume or LinkedIn. Every exceptional candidate hears back within 48 hours.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
434,399 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Remote/Hybrid • 4+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
Senior AI Engineer 1 day ago
Remote/Hybrid • 9+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
$126k – $265k per year (Estimated) • Remote • 8+ years exp
Python
SQL
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
Remote/Hybrid • 8+ years exp
Python
SQL
C#
C#
.NET
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
$66k – $137k per year (Estimated) • Equity • Remote/Hybrid • 8+ years exp • Tokyo
AI/ML
AI Agents
DevOps
SLI/SLO/SLA
Robotics
Digital Twin
Apply
iOS Engineer 4 months ago
$147k – $315k per year (Estimated) • In office • Full-Time • PhD • San Francisco
AI/ML
AI Agents
MLX ML
OpenAI
Computer Use
Mobile
Core ML
DevOps
GitHub
Apply
AI Researcher 1 year ago
$168k – $362k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
RLHF
Quantization
Knowledge Distillation
Function Calling
AI Agents
Mixture of Experts
OpenAI
DPO
SFT
GRPO
Post-training
Computer Use
Speculative Decoding
KV Cache
Tool Use
Model Distillation
DevOps
GitHub
Apply
$163k – $351k per year (Estimated) • In office • Full-Time • San Francisco
Databases
PostgreSQL
AI/ML
AI Agents
LLM
OpenAI
Computer Use
Tool Use
DevOps
GitHub
Apply
Product Designer 1 year ago
$138k – $259k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
AI Agents
OpenAI
Computer Use
Design
Figma
Sketch
Apply
AI Product FDE 1 year ago
$152k – $317k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
AI Agents
LLM
OpenAI
Edge AI
Computer Use
Apply
$60k – $100k per year • In office • Full-Time • 3+ years exp • San Francisco
Apply
$77k – $170k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Apply
$140k – $210k per year • Equity • In office • Full-Time • 4+ years exp • San Francisco
Python
Go
TypeScript
DevOps
GCP
CI/CD
Kubernetes
Management
Stripe
Apply
$200k – $250k per year • Equity 1–2% • In office • Full-Time • 3+ years exp • San Francisco
Apply
$160k – $250k per year • Equity 1–2% • In office • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
PyTorch
DevOps
AWS
Apply
See all jobs
This is one of many
434,399 more open roles from verified company boards, updated every day.