368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$140k – $250k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Automate every customer interaction with AI Phone Agents built specifically for enterprise.

Machine Learning Researcher, Audio

Location: San Francisco, CA or Remote

About Bland

At Bland.com, our mission is to empower enterprises to build AI phone agents at scale. Based in San Francisco, we are a fast-growing team reimagining how customers interact with businesses through voice. We have raised $100 million from leading Silicon Valley investors, including Emergence Capital, Scale Venture Partners, Y Combinator, and founders of Twilio, Affirm, and ElevenLabs.

Voice is quickly becoming the primary interface between businesses and their customers. We are building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human.

The Role: Machine Learning Researcher, Audio

As a Machine Learning Researcher at Bland, you'll be working on foundational research and development across the core components of our voice stack: speech-to-text, large language models, neural audio codecs, and text-to-speech. Your work will define how our agents understand, reason, and speak in real time at enterprise scale.

This is not a narrow research role. You will take ideas from theory to large-scale training to production inference systems serving millions of calls per day. You will design new modeling approaches, validate them with rigorous experimentation, and collaborate with engineering teams to deploy them into real customer environments.

What You Will Do

Build and Scale Next-Generation TTS Systems

  • Design and train large scale text-to-speech models capable of expressive, controllable, human-sounding output.

  • Develop neural audio codec-based TTS architectures for efficient, high-fidelity generation.

  • Improve prosody modeling, question inflection, emotional expression, and multi-speaker robustness.

  • Optimize for real-time, low-latency inference in production.

Advance Speech-to-Text Modeling

  • Build and fine-tune large scale ASR systems robust to accents, noise, telephony artifacts, and code switching.

  • Leverage self-supervised pretraining and large-scale weak supervision.

  • Improve transcription accuracy for real-world enterprise scenarios, including structured extraction and conversational nuance.

Pioneer Neural Audio Codecs

  • Research and implement neural audio codecs that achieve extreme compression with minimal perceptual loss.

  • Explore discrete and continuous latent representations for scalable speech modeling.

  • Design codec architectures that enable downstream generative modeling and controllable synthesis.

Develop Scalable Training Pipelines

  • Curate and process massive audio datasets across languages, speakers, and environments.

  • Design staged training curricula and data filtering strategies.

  • Scale training across distributed GPU clusters focusing on cost, throughput, and reliability.

Run Rigorous Experiments

  • Design ablation studies that isolate the impact of architectural changes.

  • Measure improvements using both objective metrics and perceptual evaluations.

  • Validate ideas quickly through focused experiments that confirm or eliminate hypotheses.

What Makes You a Great Fit

Deep Research Foundations

  • Experience with self-supervised learning, multimodal modeling, or generative modeling.

  • Ability to derive new formulations and implement them efficiently.

Expertise in Voice Modeling

  • Hands-on experience building or scaling TTS, STT, or neural audio codec systems.

  • Familiarity with large scale speech datasets and real-world audio variability.

  • Strong intuition for audio quality, prosody, and conversational dynamics.

Systems and Hardware Awareness

  • Experience training and serving large models on modern accelerators.

  • Knowledge of inference optimization techniques, including quantization, kernel optimization, and memory efficiency.

  • Understanding of real-time constraints in telephony or streaming environments.

Experimental Rigor

  • Track record of designing controlled experiments and meaningful ablations.

  • Comfortable working with both offline benchmarks and live production metrics.

  • Ability to move quickly from hypothesis to validation.

Builder Mentality

  • Comfortable in fast-moving startup environments.

  • Strong ownership mindset from research through deployment.

  • Excited by ambiguous, unsolved problems.

How You Show Up

  • You treat unsolved problems as opportunities to invent new paradigms.

  • You identify the single experiment that can validate an idea in days, not months.

  • You measure everything and let data drive decisions.

  • You are obsessed with making voice agents sound truly human.

  • You use AI tools aggressively to amplify your own impact and accelerate research cycles.

Bonus Points

  • Experience with large scale distributed training.

  • Research publications or open source contributions in speech or language AI.

  • Background in real-time speech systems or telephony.

  • PhD in ML, AI, or a related field, or equivalent research impact.

Benefits and Compensation

  • Healthcare, dental, vision, all the good stuff

  • Meaningful equity in a fast-growing company

  • Every tool you need to succeed

  • Beautiful office in Levi's Plaza, SF with rooftop views

  • Competitive salary: $160,000 to $250,000

If you are energized by building and scaling TTS models, pioneering neural audio codecs, and pushing the boundaries of speech-to-text systems, we would love to hear from you.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$140k – $292k per year (Estimated) • In office • Full-Time • 9+ years exp • Master's Degree • Peachtree Corners
Python
AI/ML
Computer Vision
Multimodal AI
Post-training
PyTorch
Ray
Time Series Forecasting
DevOps
Docker
Kubernetes
SLURM
Robotics
Isaac Lab
Isaac Sim
MuJoCo
Sim-to-Real
Apply
In office • Full-Time • 1+ year exp
JavaScript
Node JS
C
C
FFmpeg
AI/ML
Claude
ElevenLabs
OpenAI
Recommender Systems
DevOps
Rest API
Design
Adobe After Effects
Management
Airtable
n8n
Slack
Zapier
Marketing
Google Ads
Apply
$305k per year • In office • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Anthropic
Claude
LLM
Multimodal AI
Apply
$115k – $228k per year (Estimated) • Remote/Hybrid • 6+ years exp • Bachelor's Degree • Scottsdale
MATLAB
SystemVerilog
TCL Scripting
Verilog
VHDL
MATLAB
Simulink
AI/ML
Claude
Claude Code
Copilot
Cursor
Quantization
DevOps
Git
GitHub
Chips/EDA
HDL Coder
Xilinx Vivado
Apply
$110k – $219k per year (Estimated) • Equity • Remote/Hybrid • 6+ years exp • Bachelor's Degree • Boston
MATLAB
SystemVerilog
TCL Scripting
Verilog
VHDL
MATLAB
Simulink
AI/ML
Claude
Claude Code
Copilot
Cursor
Quantization
DevOps
Git
GitHub
Chips/EDA
HDL Coder
Xilinx Vivado
Apply
Product Designer 7 days ago
$150k – $200k per year • In office • Full-Time • 4+ years exp • San Francisco
AI/ML
AI Agents
Claude
LLM
Mobile
Twilio
Design
Figma
Apply
Brand Designer 17 days ago
$120k – $170k per year • In office • Full-Time • 3+ years exp • San Francisco
AI/ML
Claude
Claude Code
Mobile
Twilio
Apply
$140k – $250k per year • In office • Full-Time • San Francisco
AI/ML
ElevenLabs
Fine-tuning
LLM
Multimodal AI
Mobile
Twilio
Apply
$120k – $180k per year • In office • Full-Time • San Francisco
SQL
Mobile
Twilio
DevOps
Incident Management
Apply
$120k – $180k per year • In office • Full-Time • San Francisco
SQL
Mobile
Twilio
DevOps
Incident Management
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$185k – $385k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
JavaScript
Python
Databases
MySQL
PostgreSQL
AI/ML
OpenAI
Frontend
React.js
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.