1,282,209open jobs
74,317companies
212,999added this week
Browse all
Salary
$120k – $200k per year
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Oct 3, 2026.

Overview
Company
Impact
Profile match
Surface silent failures, pull context across traces, improve your agent before users churn.

Lemma is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.

Why this role exists

Agents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.

Making agents reliable is the problem Lemma exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.

The hardest part of our product is deciding what counts as a failure.

There is no ground truth here, and no benchmark to climb. Every customer's agent is different, what "wrong" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.

This role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.

What you’ll do

  • Own detection quality. Find the failures that don't look like failures: compliant but wrong, omissions, and patterns that only show up across thousands of traces
  • Turn implicit signals into evidence. Rephrasing, abandonment, retries, and the other ways users tell you something broke without saying so
  • Build the evals for our own system. If we can't measure precision on a problem with no labels, we can't improve it
  • Make patch generation trustworthy. Reproduce the failure, verify the fix, and know when to stay quiet instead of opening a bad PR
  • Keep it affordable. LLM-as-judge on every event is easy. Doing it at a cost per event that doesn't eat the business is the actual job
  • Read real customer traces every week. The best ideas here come from staring at production, not papers

What we’re looking for

  • High slope over years of experience. New grads and dropouts welcome
  • A track record of shipping things people actually use
  • Hands-on with LLMs in production: evals, LLM-as-judge, embeddings, and knowing when a smaller model or no model is the right call
  • Real research taste. You can tell a real improvement from noise, even when there's no ground truth to check against
  • Bonus: you were the customer once. You ran agents in production and got burned

Who you’ll work with

You'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.

Onsite in San Francisco. We don’t sponsor visas.

We keep this short on purpose. Target is an offer within two weeks of first contact.

Intro call with a founder (30 minutes). What we're building, what you've built, whether the shape of the role actually fits what you want next.

Take-home project and deep dive. We give you a real problem from Lemma. You work on it on your own time, then we walk through it together. The conversation matters more to us than the artifact.

Paid work trial (onsite in-person). Real problem, real codebase, sitting with the team. You find out what working here actually feels like before you commit, which tells you more than anything we could say about it.

May ask for additional references.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,282,209 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$195k – $255k per year • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • New York • Menlo Park
Python
Databases
PostgreSQL
Google BigQuery
BigQuery
AI/ML
NLP
LLM
Machine Learning
DevOps
Terraform
GCP
Analytics
Metabase
Apply
$193k – $262k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Cupertino
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
DeepSeek
Stable Diffusion
JAX
Llama
TensorFlow
PyTorch
AWS Trainium
MLIR
DevOps
AWS
Amazon EC2
GitHub
Amazon S3
Apply
$143k – $193k per year • Equity • In office • Full-Time • 5+ years exp • Master's Degree • San Diego
Python
Java
SQL
C++
Perl
AI/ML
Spark
Machine Learning
Apply
$143k – $193k per year • Equity • In office • Full-Time • 5+ years exp • Master's Degree • Seattle
Python
Java
SQL
C++
Perl
AI/ML
Spark
Machine Learning
Apply
≈ $97k – $220k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Nashville
Python
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
Edge AI
Agentic Workflows
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Apply
$200k – $250k per year • In office • Full-Time • 8+ years exp • San Francisco
Python
TypeScript
Python
FastAPI
AI/ML
Model Context Protocol
AI Agents
LLM
Apply
Founding Engineer 3 hours ago
$156k – $182k per year • Remote (United States) • Full-Time • 2+ years exp • New York
Python
JavaScript
TypeScript
SQL
Python
FastAPI
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Analytics
ETL/ELT
Apply
$112k – $160k per year • Hybrid • Burlington
Python
JavaScript
TypeScript
SQL
AI/ML
Model Context Protocol
Function Calling
Speech Recognition
LLM
LiveKit
Vapi
Voice Agents
Tool Use
Mobile
Twilio
DevOps
Rest API
Cybersecurity
HIPAA
QA
Postman
Apply
≈ $133k – $245k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Huntsville
Python
Java
TypeScript
C++
AI/ML
AI Agents
Text-to-Speech
DevOps
Helm
GitLab CI
CI/CD
Kubernetes
Cybersecurity
SBOM
Apply
≈ $147k – $256k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Los Angeles
SQL
Databases
Google BigQuery
BigQuery
AI/ML
AI Agents
Apply
Product Engineer 4 days ago
$120k – $200k per year • Equity 0.5–2% • In office • Full-Time • San Francisco
TypeScript
AI/ML
AI Agents
DevOps
GitHub
Management
Slack
Apply
In office • Internship • San Francisco • Munich
Python
JavaScript
TypeScript
Python
FastAPI
Pydantic
Databases
PostgreSQL
pgvector
AI/ML
Langfuse
Pydantic AI
LLM
LLM Guardrails
Frontend
React.js
Vite
TanStack Router
Chakra UI
Apply
ML Research Scientist 10 hours ago
$300k – $500k per year • In office • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
Post-training
Apply
$250k – $339k per year • Remote (United States) • Full-Time • 10+ years exp • Master's Degree • Atlanta • Cambridge • San Francisco • Thousand Oaks
Apply
Senior Data Engineer 5 hours ago
$200k – $250k per year • In office • Full-Time • 6+ years exp • San Francisco
Python
SQL
Databases
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Groq
E2B
DevOps
AWS
AWS Lambda
Management
Zapier
Apply
Operating Engineer 5 hours ago
$159k per year • In office • Full-Time • 3+ years exp • San Francisco
Management
Outlook
Microsoft Office
Apply
See all jobs
This is one of many
1,282,209 more open roles from verified company boards, updated every day.