747,188open jobs
44,848companies
107,947added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Jul 17, 2026. Simile scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Simulate how real customers respond to a launch, price change, or campaign — before you ship. Built by the Stanford researchers behind generative agents.

About the Company

Simile is The Simulation Company. We simulate human behavior to keep people at the center of the decisions that shape the world. With AI, anyone can create a product, a campaign, a policy, or a script - the bottleneck has moved upstream. The hard question is no longer whether you can create something, but what to create, for whom, and how to bring it to life. Those are fundamentally human decisions, and they shouldn't be left to chance or handed off to an algorithm. We're building the infrastructure to understand human behavior at scale and to represent humans in an increasingly agentic world. Our mission is to simulate all eight billion people on earth.

We launched five months ago. Since then we've grown revenue 5x, built a new foundation model for human behavior that has run tens of millions of simulations for F100 enterprises, trained a first-of-its-kind confidence model that predicts the accuracy of every simulation, and released the first product that lets organizations verifiably predict the future. The world's leading companies use Simile to make business-critical decisions - from consumer leaders like CVS Health and Wealthfront to professional services organizations like Deloitte and Gallup - strategizing product launches, entering new markets, and forecasting earnings calls.

We've raised over $200M at a $2B post-money valuation led by Greenoaks, with Index Ventures, Hanabi, A*, Bain Capital Ventures, and CVS Health Ventures. We've grown from a small home in Palo Alto to a global team of 50+, and we're building a team of the best researchers, engineers, designers, and operators in the world. The future is too important to be left to chance.

About the Role

As a Member of Technical Staff in Evaluations Engineering, you will build the systems that enable Simile to evaluate whether our simulations of human behavior are accurate, trustworthy, and improving over time.

You will work across data and evaluation infrastructure, evaluation execution workflows, backend services, automation, and internal tooling. Your initial focus will include streamlining how evaluations are run across models; strengthening evaluation versioning, data models, and access controls; and automating customer validations, survey operations, and human data workflows.

Evaluation at Simile presents unusual engineering challenges. Our models predict distributions of human behavior, and the ground truth used to evaluate them can be noisy and heterogeneous. You will partner closely with Evals, Modeling, Product Engineering, and Data Operations to turn complex methods and inputs into systems that are reproducible, scalable, and useful for model development and business decisions.

In this role, you will:

  • Build evaluation execution infrastructure: Develop the services, pipelines, and orchestration needed to run evaluations efficiently across datasets, model versions, populations, and use cases.

  • Strengthen evaluation data systems: Design relational schemas, versioning, provenance, permissions, and quality controls that make evaluation results reproducible and trustworthy.

  • Automate validation and data collection: Partner with Evals and Data Operations to streamline customer validations, survey deployment, response ingestion, and the integration of new ground truth.

  • Build human data workflows: Create labeling and review tools that enable external experts and operators to contribute high-quality judgments to evaluation campaigns.

  • Develop evaluation tooling: Build interfaces that help teams manage evals, compare models, investigate results, and identify regressions.

Requirements

Must Haves

  • Strong Engineering Fundamentals: Several years of experience building and maintaining production-quality software, with sound judgment in system design, testing, debugging, and maintainability.

  • Data and Systems Experience: Experience building backend services, data pipelines, automation workflows, and relational data models.

  • End-to-End Execution: Ability to work across data, backend, and interface layers and take ambiguous projects from technical design through deployment and adoption.

  • Evaluation Judgment: Strong intuition for what makes evaluation infrastructure reliable, including versioning, provenance, reproducibility, holdout integrity, noisy ground truth, and meaningful model comparisons.

  • ML and LLM Fluency: Familiarity with modern model-development and evaluation workflows sufficient to partner effectively with modeling and evaluation researchers.

  • Product and User Judgment: Ability to build clear, efficient tools for researchers, engineers, data operators, and other expert users.

  • Ownership and Communication: A track record of independently driving important technical work and collaborating effectively across engineering, research, and operations.

Nice to Haves

We do not expect one person to have all of these. We are hiring a team with complementary strengths.

  • Model-Evaluation Infrastructure: Experience building LLM or ML evaluation systems, benchmark platforms, regression suites, experiment-tracking tools, or model-quality dashboards.

  • Research and Internal Tools: Experience developing technical surfaces for ML engineers, researchers, data scientists, or operations teams.

  • Human Data Systems: Experience with labeling platforms, expert-review workflows, LLM-as-judge systems, grader calibration, or other human-in-the-loop evaluation methods.

  • Data-Collection Automation: Experience automating surveys, experiments, customer-data ingestion, or other human data collection workflows.

  • Statistical Fluency: Comfort reasoning about sampling error, uncertainty, calibration, confidence intervals, and distributional metrics.

  • Sensitive Data and Access Controls: Experience designing permissions, auditability, and data-governance systems for human or customer data.

  • Agentic Engineering: Experience using modern AI coding tools to accelerate development while independently testing and validating their output.

You might be a great fit if you have worked on ML or evaluation infrastructure, data platforms, backend systems, experiment tracking, research tooling, workflow orchestration, internal tools, or human-data systems. You do not need to have held an “Evals Engineer” title, but you should have several years of experience building reliable production software and be excited to apply that experience to model quality.

You do not need to match every bullet. If you do not perfectly see yourself in this JD but believe you would be exceptional at building the measurement layer for behavioral simulation, we would love to hear from you.

Compensation & Benefits

At Simile, we provide competitive compensation packages that include base salary, equity, and comprehensive benefits.

  • Salary Range: $200,000 - $400,000 USD

    • Note: Final offers are based on experience, specialized skills, interview performance, and relevant training.

  • Equity: Grants are available for eligible roles, subject to board approval.

  • Health & Wellness: Comprehensive medical, dental, and vision coverage.

  • Time Off: Flexible time off policies to support work-life balance.

Our Process

We prioritize thoughtful conversations and clear examples of past work. Our hiring journey is designed to help both sides align on fit, working style, and expectations.

Reapplication Policy: To ensure a fair and thorough evaluation for all applicants, Simile observes a 90-day waiting period before reconsidering candidates for the same role.

Commitment to Diversity & Inclusion

Equal Opportunity: Simile is an equal opportunity workplace. We welcome applicants of all backgrounds and identities, valuing an environment where everyone can contribute authentically.

Accommodations: If you require support or reasonable accommodations during the application process due to a disability, please let us know. We are happy to assist.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
747,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$195k – $205k per year • In office • Master's Degree • San Jose
AI/ML
Fine-tuning
Multimodal AI
Transformers
TensorFlow
PyTorch
Mixture of Experts
Model Distillation
Machine Learning
Apply
Staff AI Engineer 2 months ago
≈ $134k – $275k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Master's Degree • Boca Raton
JavaScript
AI/ML
Claude Code
Model Context Protocol
AI Agents
LLM
RAG
Hybrid Search
LLM Guardrails
Multi-Agent Systems
Frontend
Bootstrap
Apply
≈ $171k – $298k per year (Estimated) • Remote (United States) • 5+ years exp • Bachelor's Degree • San Francisco
Python
Java
Scala
AI/ML
TensorFlow
PyTorch
Anomaly Detection
Machine Learning
Analytics
A/B Testing
Apply
$130k – $220k per year • Hybrid • Full-Time • 3+ years exp • PhD • Santa Clara
Python
C++
AI/ML
Machine Learning
Robotics
ROS
Path Planning
Motion Planning
Imitation Learning
Apply
$116k – $208k per year • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree • United States
Databases
PostgreSQL
Redis
DynamoDB
Apache Kafka
OpenSearch
Amazon Aurora
AI/ML
LangGraph
AutoGen
LangChain
AI Agents
Semantic Kernel
LLM
LLM Guardrails
Multi-Agent Systems
DevOps
Rest API
gRPC
Terraform
Helm
GitHub Actions
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Kubernetes
Platform Engineering
Chaos Engineering
Service Mesh
Amazon EKS
Progressive Delivery
SLI/SLO/SLA
IAM
Amazon ECS
Amazon Kinesis
API Gateway
Cybersecurity
ISO 27001
SOC 2
Zero Trust
SIEM
Chips/EDA
PoC Library
Apply
≈ $20k – $54k per year (Estimated) • In office • Full-Time • Moscow
Python
Databases
PostgreSQL
Milvus
pgvector
Qdrant
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Function Calling
AI Agents
AWS Bedrock
CrewAI
LLM
RAG
OpenRouter
OpenAI
Anthropic
Tool Use
DevOps
Azure
AWS
Docker
Apply
≈ $37k – $78k per year (Estimated) • Hybrid • 4+ years exp • Moscow
Python
SQL
Python
FastAPI
Databases
PostgreSQL
MinIO
AI/ML
llama.cpp
LoRA
Model Context Protocol
vLLM
Fine-tuning
Function Calling
AI Agents
SGLang
PEFT
QLoRA
Transformers
LLM
RAG
SFT
Context Engineering
Tool Use
DevOps
CI/CD
Git
Docker
Linux
Apply
≈ $69k – $171k per year (Estimated) • In office • Bachelor's Degree • Stockholm
JavaScript
TypeScript
AI/ML
AI Agents
Agentic Workflows
Frontend
Next.js
React.js
DevOps
Vercel
Analytics
A/B Testing
Apply
≈ $27k – $56k per year (Estimated) • In office • Full-Time • Moscow
Java
Java
Spring Boot
Databases
PostgreSQL
AI/ML
LLM
RAG
DevOps
CI/CD
Docker
Kubernetes
Management
Telegram
Apply
≈ $24k – $54k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Master's Degree • India
Python
JavaScript
TypeScript
SQL
Node JS
Databases
PostgreSQL
AI/ML
LoRA
Fine-tuning
Embeddings
Quantization
Prompt Engineering
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Hallucination
Hugging Face
LLM Evaluation
LLM Guardrails
Frontend
Vue.js
GraphQL
Angular
React.js
DevOps
Rest API
Ansible
CI/CD
Git
Docker
Kubernetes
Configuration Management
Tekton
Windows
DNS
Management
Agile
Apply
$200k – $400k per year • In office • Full-Time • Bachelor's Degree • San Francisco
Python
AI/ML
vLLM
CUDA Toolkit
Fine-tuning
Quantization
JAX
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Tokenization
CUDA
Triton
NCCL
InfiniBand
NVLink
KV Cache
Apply
$200k – $400k per year • In office • Full-Time • San Francisco • New York
Python
SQL
AI/ML
Cursor
Claude Code
AI Agents
LLM
OpenAI Codex
Post-training
LLM Evaluation
Analytics
A/B Testing
Apply
$200k – $400k per year • In office • Full-Time • San Francisco
Python
AI/ML
Fine-tuning
AI Agents
Apply
$150k – $300k per year • In office • Full-Time • New York • San Francisco
AI/ML
AI Agents
Apply
Brand Designer 1 month ago
$200k – $300k per year • In office • Full-Time • San Francisco • New York
AI/ML
AI Agents
Design
Figma
Apply
≈ $159k – $427k per year (Estimated) • Equity • Hybrid • Internship • 2+ years exp • Bachelor's Degree • San Jose • Milpitas • San Francisco
Python
SQL
AI/ML
Claude
Fine-tuning
Scikit-learn
SciPy
NLP
Llama
Pandas
NumPy
LLM
GPT-4
LLM Guardrails
Machine Learning
DevOps
Docker
Apply
≈ $159k – $427k per year (Estimated) • Equity • Hybrid • Internship • Bachelor's Degree • San Jose • Milpitas • San Francisco
Python
SQL
AI/ML
Claude
Fine-tuning
Scikit-learn
SciPy
NLP
Llama
Pandas
NumPy
LLM
GPT-4
LLM Guardrails
Machine Learning
DevOps
Docker
Apply
≈ $156k – $420k per year (Estimated) • Equity • Hybrid • Internship • Bachelor's Degree • San Jose • Milpitas • San Francisco
Python
Bash
AI/ML
Machine Learning
DevOps
GitHub Actions
CI/CD
AWS
Kubernetes
AIOps
Linux
Apply
$245k – $279k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco • New York • San Jose
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
≈ $132k – $342k per year (Estimated) • Equity • Hybrid • Internship • 2+ years exp • Bachelor's Degree • San Jose • Milpitas • San Francisco
Python
Bash
DevOps
Jenkins
Linux
Apply
See all jobs
This is one of many
747,188 more open roles from verified company boards, updated every day.