368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$125k – $165k per year
Location
In office
Employment
Full-Time
Overview
Company
Impact
Profile match
Cantina is a social artificial intelligence technology company based in San Francisco, California, and founded in 2023. The company provides a platform where users can create, interact with, and share multimodal AI characters that feature unique personalities and the ability to generate video content. It operates primarily through a mobile application and web interface, focusing on the intersection of generative AI and social media for a global audience of creators and consumers.

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are looking for an MLOps Engineer to build and scale the inference infrastructure for our generative audio models, including Text-to-Speech (TTS), voice conversion, and Automatic Speech Recognition (ASR). You will be responsible for designing and deploying high-performance systems that ensure low-latency, reliable, and scalable model serving for both streaming and batch inference. This role is central to bridging the gap between research and production, ensuring our audio models are optimized for performance and cost-efficiency as we scale.

What You’ll Do:

  • Design and maintain inference infrastructure for generative audio model architectures.

  • Implement and manage high-performance inference engines.

  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.

  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.

  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.

  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.

  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring:

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.

  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.

  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.

  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.

  • Experience with GPU-accelerated inference and performance profiling techniques.

  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation:

The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account - $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$32k – $78k per year (Estimated) • Remote/Hybrid • Contractor • 5+ years exp • Saint Petersburg
DevOps
AWS
Azure
CI/CD
Docker
GCP
IAM
Kubernetes
Cybersecurity
ISO 27001
Least Privilege
NIST CSF
SOC 2
Management
Google Workspace
Apply
$20k – $49k per year (Estimated) • Remote/Hybrid • Moscow
JavaScript
Python
PHP
Python
Django
PHP
Drupal
Databases
PostgreSQL
Frontend
HTMX
DevOps
CI/CD
Docker
Git
Apply
$19k – $48k per year (Estimated) • In office • Full-Time • Tver
Python
Python
Celery
Django
Django REST Framework
Databases
PostgreSQL
RabbitMQ
DevOps
GitLab
Rest API
Management
Jira
Apply
$27k – $72k per year (Estimated) • In office • Full-Time • Pune
JavaScript
Node JS
TypeScript
Frontend
React.js
DevOps
CI/CD
Docker
Git
Rest API
Apply
$23k – $61k per year (Estimated) • In office • Full-Time • Pune
JavaScript
Node JS
TypeScript
Frontend
React.js
DevOps
CI/CD
Docker
Git
Rest API
Apply
$175k – $250k per year • In office • Full-Time • 6+ years exp • San Francisco
SQL
Mobile
Braze
Deep Linking
Analytics
A/B Testing
Marketing
Amplitude
Iterable
Mixpanel
Apply
$180k – $270k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sunnyvale • San Francisco
C++
Go
Node JS
JavaScript
AI/ML
Speech Recognition
DevOps
WebRTC
Apply
$200k – $220k per year • In office • Full-Time
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
Knowledge Distillation
Multimodal AI
PyTorch
Synthetic Data
Tokenization
DPO
FSDP
GRPO
LLM Guardrails
Text-to-Speech
Speech Recognition
Apply
$200k – $240k per year • In office • Full-Time • 8+ years exp
Go
DevOps
AWS
Terraform
Apply
$200k – $220k per year • In office • Full-Time
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
Knowledge Distillation
Multimodal AI
PyTorch
Quantization
Synthetic Data
Tokenization
DPO
FSDP
GRPO
LLM Guardrails
Text-to-Speech
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.