435,469open jobs
15,038companies
66,389added this week
Browse all
Salary
$150k – $250k per year
Location
Remote/Hybrid (San Francisco, New York, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Distyl AI deploys generative artificial intelligence systems inside large enterprises. Its teams build and operate workflow specific applications. The company works with Fortune 500 operations groups.

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations - both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations - from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.

What We Are Looking For

At Distyl, we build AI systems using Evaluation-Driven Development -an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.

This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development-and who want to apply that philosophy directly to systems that matter.

Key Responsibilities

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments

  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives

  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself

  • Develop and maintain evaluation pipelines-offline and online-that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone

  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve

  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production

What We Require

  • 5+ years of software engineering experience

  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code

  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes

  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale

  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production

  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration

  • Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so

What We Offer

  • The base salary range for this role is $150K - $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot

  • Complimentary in-office lunches and snacks provided

  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems

  • Ownership of high-impact projects across top enterprises

  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday-Thursday) in-office. .

#LI-Hybrid

We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,469 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Remote/Hybrid • 4+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
Senior AI Engineer 1 day ago
Remote/Hybrid • 9+ years exp
Python
JavaScript
C#
C#
.NET
AI/ML
LangChain
Prompt Engineering
AI Agents
LLM
OpenAI
Hugging Face
Frontend
React.js
Apply
$126k – $265k per year (Estimated) • Remote • 8+ years exp
Python
SQL
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
Remote/Hybrid • 8+ years exp
Python
SQL
C#
C#
.NET
Databases
Snowflake
AI/ML
LangGraph
LangChain
LlamaIndex
dbt
Embeddings
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Vector
Cortex
Apply
Remote/Hybrid
Python
C#
C#
.NET
Cybersecurity
Threat Modeling
Apply
$100k – $140k per year • Remote/Hybrid • Full-Time • San Francisco
Apply
$41k – $84k per year (Estimated) • In office • Full-Time • London
AI/ML
Edge AI
Apply
$170k – $190k per year • Remote/Hybrid • Full-Time • PhD • New York • San Francisco
Apply
$180k – $220k per year • Remote/Hybrid • Full-Time • San Francisco
Apply
Lead AI Strategist 2 months ago
$200k – $250k per year • Remote/Hybrid • Full-Time • 15+ years exp • San Francisco • New York
Apply
$76k – $92k per year • Remote • Internship • Bachelor's Degree • San Francisco
SQL
Analytics
SSIS
SSAS
Management
Microsoft Project
Apply
$70k – $82k per year • Remote • Internship • San Francisco
Analytics
Microsoft Excel
Apply
$70k – $82k per year • In office • Internship • San Francisco
Apply
$79k – $95k per year • Remote • Full-Time • San Francisco
Apply
$213k – $374k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Chicago • New York • Atlanta • San Francisco
AI/ML
AI Agents
Agentforce
Agentic Workflows
Marketing
Salesforce
Apply
See all jobs
This is one of many
435,469 more open roles from verified company boards, updated every day.