659,077open jobs
38,403companies
95,422added this week
Browse all
Salary
$181k – $250k per year
Location
In office (San Francisco)
Seniority
Staff · 8+ years exp
Overview
Company
Impact
Profile match
Aircall is a Paris company founded in 2014 that provides a cloud telephony and call centre platform for small and mid-sized businesses. Its product integrates deeply with customer relationship and helpdesk tools such as Salesforce, HubSpot and Zendesk. The company became a unicorn in 2021 and has offices in Paris, New York, Sydney and Madrid.

Aircall is a unicorn, AI-powered customer communications platform used by 22,000+ companies worldwide to drive revenue, resolve issues faster, and scale customer-facing teams. We’re redefining customer communications by bringing voice, SMS, WhatsApp, and AI together into one seamless workspace.

Our momentum comes from a simple idea: help teams work smarter, not harder. Aircall’s AI Voice Agent automates routine calls, AI Assist streamlines post-call work, and AI Assist Pro delivers real-time guidance so people can do their best work. The result is higher revenue, faster resolutions, and teams that scale with confidence.

Aircall is headquartered in Paris, our European HQ, with a strong North American presence anchored in Seattle, our North American HQ, and teams across Madrid, London, Berlin, San Francisco, New York City, Sydney, and Mexico City. We’ve built a product customers love and a business that’s scaling quickly, backed by world-class investors and driven by rapid AI innovation across multiple product lines.

At Aircall, you’ll join a company in motion. We’re ambitious, product-driven, and execution-focused, with visible impact, fast decisions, and real growth.

How we work at Aircall: We’re customer-obsessed, data-driven, and focused on delivering meaningful outcomes. We value ownership, continuous learning, and thoughtful speed. If you thrive in a collaborative, fast-moving environment where trust and impact matter, you’ll feel at home here.

Aircall's AI suite includes an AI Voice Agent and AI Messaging Agent that autonomously handle calls, WhatsApp, and SMS, plus AI Assist, which delivers real-time coaching, call summaries, and CRM automation for sales and support teams. We are looking for someone that can build out the evaluation foundation across all of these products and other agentic products. You'll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together by establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. You will set up repeatable pipelines for regression testing and benchmarking as models and features evolve so teams can ship confidently without re-inventing evaluation methodology for each product.

Key Responsibilities

  • Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat and messaging.
  • Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency.
  • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies. 
  • Analyze system design decisions and identify strengths, weaknesses, and potential failure points.
  • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time.
  • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers.
  • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window.
  • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates.
  • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics.

Minimum Qualifications

  • BS in Computer Science, Machine Learning, Statistics, or related field
  • 3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience
  • Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models.
  • Hands-on experience in failure analysis and evaluating LLMs
  • Experience building automated evaluation systems
  • Strong communication skills to articulate complex technical concepts across technical and non-technical audiences
  • Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation.

Preferred Qualifications

  • MS / PhD in Computer Science, Machine Learning, Statistics, or related field
  • Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation)
  • Experience with synthetic data generation and prompt engineering
  • Experience training or fine-tuning voice models at scale, with familiarity in synthetic data generation, model distillation, or low-latency inference optimization for production voice agents.

Base salary range:

$181,000—$250,000 USD

Why join us?

Key moment to join Aircall in terms of growth and opportunities

Our people matter, work-life balance is important at Aircall

Fast-learning environment, entrepreneurial and strong team spirit

45+ Nationalities: cosmopolite & multi-cultural mindset

Competitive salary package & benefits

DE&I Statement: 

At Aircall, we believe diversity, equity and inclusion - irrespective of origins, identity, background and orientations - are core to our journey. 

We pride ourselves on promoting active inclusion within our business to foster a strong sense of belonging for all. We’re working to create a place filled with diverse people who can enrich and learn from one another. We’re committed to ensuring that everyone not only has a seat at the table but is valued and respected at it by providing equal opportunities to develop and thrive.  

We will constantly challenge ourselves to make sure that we live up to our ambitions around diversity, equity and inclusion, and keep this conversation open. Above all else, we understand and acknowledge that we have work to do and much to learn.

Want to know more about candidate privacy? Find our  Candidate Privacy Notice here.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
659,077 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$26k – $54k per year (Estimated) • In office • 3+ years exp • Yekaterinburg
Python
Python
FastAPI
Django
Databases
PostgreSQL
Redis
FAISS
Apache Kafka
AI/ML
LangGraph
LangChain
Model Context Protocol
AI Agents
Langfuse
LLM
RAG
A2A
DevOps
Helm
Docker
Kubernetes
Amazon S3
Apply
(USA) Data Analyst II 8 hours ago
$59k – $121k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • United States
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Copilot
AI Agents
Anomaly Detection
QA
Playwright
Apply
$170k per year • In office • Full-Time • Park
Python
Go
Java
C#
AI/ML
Copilot
Claude Code
AI Agents
DevOps
CI/CD
Apply
$197k – $324k per year • In office • Full-Time • 10+ years exp • PhD • San Francisco • Washington
Python
Go
Java
PHP
Ruby
Databases
MySQL
AI/ML
AI Agents
Agentforce
Management
Slack
Apply
$150k – $170k per year • In office
AI/ML
AI Agents
Management
Agile
Apply
$57k – $105k per year (Estimated) • In office • Madrid
Python
JavaScript
TypeScript
Ruby
Node JS
Databases
PostgreSQL
Redis
DynamoDB
OpenSearch
AI/ML
Voice Agents
Frontend
GraphQL
React.js
DevOps
Rest API
AWS
Kubernetes
Amazon EKS
Management
WhatsApp
Apply
In office • Mexico City
AI/ML
Voice Agents
Management
WhatsApp
Marketing
Salesforce
Zendesk
HubSpot
Apply
$180k – $220k per year • In office • Full-Time • 5+ years exp • San Francisco
Python
Databases
Databricks
ClickHouse
ElasticSearch
OpenSearch
AI/ML
LangGraph
LangChain
AWS Bedrock
LLM
AWS Bedrock AgentCore
Red Teaming
LLM Evaluation
LLM Guardrails
Voice Agents
DevOps
Splunk
Terraform
GCP
Datadog
Azure
AWS
Platform Engineering
SLI/SLO/SLA
IAM
Cybersecurity
Wiz
Threat Modeling
Management
WhatsApp
Apply
$58k – $159k per year (Estimated) • In office • Berlin
AI/ML
Voice Agents
Management
WhatsApp
Marketing
Salesforce
Apply
$71k – $147k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Paris
AI/ML
Voice Agents
Management
WhatsApp
Apply
Business Risk Officer 7 hours ago
$163k – $204k per year • Remote • 7+ years exp • San Francisco
Apply
$128k – $192k per year • Equity • In office • Full-Time • 1+ year exp • San Francisco • Pleasanton
Apply
Investment Officer 8 hours ago
$129k – $223k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Apply
$152k – $327k per year (Estimated) • In office • 13+ years exp • San Francisco
AI/ML
OpenAI
Anthropic
Management
Stripe
Apply
$152k – $279k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Python
JavaScript
Node JS
AI/ML
LLM
Frontend
React.js
DevOps
Terraform
AWS CDK
Azure
AWS
Docker
Kubernetes
Platform Engineering
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
HPC
Apply
See all jobs
This is one of many
659,077 more open roles from verified company boards, updated every day.