405,710open jobs
14,091companies
78,515added this week
Browse all
Salary
$120k – $140k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Ellipsis Health analyses speech to measure signs of depression and anxiety, and now runs voice care-management agents for health plans. Its models read vocal biomarkers from a short spoken sample rather than a questionnaire. Payers use the technology to reach members who never complete a screening form.

About the Team

The Forward Deployed Team is our primary customer-facing unit, acting as the crucial bridge between our core product and our enterprise clients. This fast-moving team manages the end-to-end lifecycle of enterprise conversational AI deployments, from pre-sales Statements of Work (SOWs) through go-live.

We maintain daily, high-touch interactions with customers to communicate workflow progress, manage User Acceptance Testing (UAT) deadlines, integrate client feedback, and ensure that deployments are executed rapidly while maintaining rigorous quality standards.

Ellipsis Health is located in the San Francisco Bay Area, but we are open to remote candidates within the United States.

About the Role

As a Forward Deployed QA Engineer, you will occupy a critical, high-impact role dedicated to ensuring the reliability, stability, and quality of our core conversational AI product, Sage, across diverse client workflows.

This role bridges the gap between Quality Assurance, AI Engineering, and Production Operations. You will focus heavily on automated testing, prompt engineering validation, and rapid root cause analysis (RCA) of Large Language Model (LLM)-driven behaviors in fast-paced, real-world deployments.

Responsibilities:

  • Workflow Mapping & Test Case Generation: Deeply analyze assigned client workflows to design robust, comprehensive positive and negative test cases that safeguard system stability.

  • AI-Driven Test Automation: Build and execute automated test scenarios by configuring shadow agents.

  • Prompt Evaluation & Optimization: Apply a strong understanding of prompt awareness to draft, refine, and evaluate prompts used within the testing framework to accurately simulate user behaviors and edge cases.

  • End-to-End Testing Execution: Strategically deploy specific testing methodologies including Sanity, Smoke, Regression, and Functional testing - determining the exact environment (staging, pre-production, production) and timing for each execution.

  • Deployment Cadence & Cross-Functional Collaboration: Partner closely with engineering teams during release cycles to proactively identify, triage, and unblock technical roadblocks, ensuring the product is continuously deployment-ready.

  • Daily LLM Defect RCA: Perform rigorous, daily root cause analysis on LLM-specific failures inherent to generative AI, including hallucinations, high latency, and logic deviations.

  • Live Production Call Debugging: Investigate live customer calls and production incidents in real time to unblock critical production use cases.

  • Audio & Transcription Validation: Query and analyze historical call transcripts, system behaviors, and audio data pipelines to pinpoint where a conversational workflow broke down.Speech-to-Speech (S2S) Pipeline Monitoring: Monitor and evaluate the end-to-end voice AI pipeline. This involves analyzing Automatic Speech Recognition (ASR) accuracy, managing audio-to-text latency issues, and understanding general Speech-to-Speech mechanics alongside the stability of the core Knowledge Base feeding the AI.

  • Advanced Evaluation Frameworks: Maintain a strong conceptual understanding of advanced LLM evaluation paradigms and tools such as LLM-as-a-judge - to remain aware of how AI response quality and accuracy are programmatically graded at scale.

  • Telephony & Call Flow Awareness: Possess a foundational understanding of real-world call management and telephony routing concepts, including how the system is expected to navigate warm transfers, blind transfers, and voicemail detection workflows.

Qualifications:

  • Experience in QA Engineering: Strong background in software quality assurance, with a proven track record of designing, executing, and managing end-to-end test strategies (Smoke, Sanity, Regression, Functional).

  • LLM & Generative AI Expertise: Hands-on experience or deep technical familiarity with troubleshooting LLM behaviors, diagnosing hallucinations, managing latency, and using LLM call-tracing tools.

  • Technical & Logging Proficiency: Ability to comfortably write SQL queries (specifically PostgreSQL) to pull data logs and navigate cloud infrastructure logs (such as GCP) to perform rapid root-cause analysis.

  • Voice & Conversational AI Domain Knowledge: Foundational understanding of Speech-to-Speech pipelines, including Automatic Speech Recognition (ASR), audio-to-text workflows, and core knowledge base integrations.

  • Telephony Foundations: Basic familiarity with enterprise telephony routing, call management mechanics (warm/blind transfers), and voicemail detection systems.

  • Client-Facing Capability: Strong communication skills and the professional agility required to manage UAT timelines, coordinate with client stakeholders, and support rapid production deployments.

Salary and Benefits

We offer competitive salary and benefits, including 401(k) matching, health, vision, and dental insurance, and very flexible paid time off.

The typical salary range for this role is $120,000 to $140,000 USD, depending on skills, qualifications, and relevant experience.

Background Checks

As a health technology company, we reserve the right to run background checks on candidates to whom we extend offers, in compliance with applicable laws. We evaluate candidates holistically and comply with all “ban the box” regulations.

Assistance

If you have a disability or require accommodations during the application or recruitment process, please contact [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
405,710 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$150k – $180k per year • Equity 0.5–0.5% • In office • Full-Time • San Francisco
TypeScript
Databases
PostgreSQL
Supabase
DevOps
Vercel
QA
Playwright
Apply
$77k – $201k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sydney
JavaScript
Kotlin
Objective-C
TypeScript
Frontend
React.js
Mobile
Expo
React Native
DevOps
AWS
Azure
GCP
Apply
$27k – $71k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Hyderabad
Java
TypeScript
JavaScript
Java
Spring Boot
Frontend
Angular
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$138k – $250k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Los Angeles
C++
JavaScript
Python
Python
Alembic
Django
FastAPI
Flask
Databases
Neo4j
PostgreSQL
Redis
DevOps
CI/CD
Docker
Git
Podman
Game Dev
Houdini
Design
Maya
Apply
AI Concept artist 1 day ago
$30k – $113k per year (Estimated) • In office • 3+ years exp • Los Angeles
AI/ML
ChatGPT
Edge AI
Fine-tuning
Midjourney
Prompt Engineering
Runway
Stable Diffusion
Game Dev
Unity
Unreal Engine
Design
Adobe Photoshop
Blender
Procreate
Apply
$180k – $230k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
Python
Pydantic
Databases
Databricks
AI/ML
Edge AI
Langfuse
LLM
NLP
NumPy
Pandas
Scikit-learn
DevOps
AWS
Azure
GCP
Analytics
A/B Testing
Apply
$123k – $231k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • San Francisco
SQL
Databases
Databricks
Analytics
ETL/ELT
Apply
$160k – $250k per year • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
Edge AI
LLM
Text-to-Speech
DevOps
Atlantis
AWS
Azure
CI/CD
Datadog
Docker
GCP
Git
GitLab
GitOps
IAM
Kubernetes
New Relic
OpenTelemetry
PagerDuty
Prometheus
SLI/SLO/SLA
Terraform
Cybersecurity
HIPAA
Snyk
SOC 2
Management
Confluence
Jira
Apply
$150k – $190k per year • Remote • Full-Time • 8+ years exp • San Francisco
AI/ML
AI Agents
Cybersecurity
HIPAA
Marketing
Salesforce
Apply
$140k – $160k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
TypeScript
AI/ML
AI Agents
Edge AI
DevOps
CI/CD
Datadog
Splunk
QA
Playwright
Selenium
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
Coordinator: Docket 2 hours ago
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
405,710 more open roles from verified company boards, updated every day.