368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$200k – $260k per year
Location
In office
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Cantina is a social artificial intelligence technology company based in San Francisco, California, and founded in 2023. The company provides a platform where users can create, interact with, and share multimodal AI characters that feature unique personalities and the ability to generate video content. It operates primarily through a mobile application and web interface, focusing on the intersection of generative AI and social media for a global audience of creators and consumers.

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are looking for a new Member of Technical Staff to build and scale the data pipelines behind our large video generation models. This role is focused on collecting large amounts of relevant video data, preparing high-quality training samples, and developing robust preprocessing, filtering, and parsing workflows. You'll orchestrate annotation pipelines across platforms such as MTurk and own the full lifecycle of training data, from raw ingestion to clean, model-ready samples that directly drive quality improvements. This role sits at the intersection of data engineering and ML research, making it central to how we turn messy real-world data into the fuel that moves our models forward.

What You’ll Do:

  • Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and DynamoDB.

  • Design and run annotation workflows across platforms such as MTurk, Prolific, and Mechanical Turk, including task design, quality control, and label validation.

  • Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.

  • Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.

  • Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.

  • Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.

  • Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.

  • Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.

  • Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.

What You’ll Bring:

  • 3+ years of experience in machine learning, applied ML, data pipelines, or related engineering roles, ideally working on large-scale multimodal, video, or vision-based systems.

  • Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.

  • Hands-on experience preparing training data for ML models, including parsing, filtering, dataset curation, quality control, and large-scale data handling using tools such as AWS S3 and DynamoDB.

  • Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.

  • Experience working with Kubernetes for orchestrating distributed workloads, including data preprocessing, pipeline execution, and dataset delivery to training clusters.

  • Comfort working across cloud and on-demand compute environments such as AWS and RunPod, with the ability to port and optimize pipelines across infrastructure.

  • Familiarity with distributed data processing frameworks and experience designing systems that operate reliably at scale across many nodes or workers.

  • Working knowledge of PyTorch and the broader deep learning stack, with the ability to read, debug, and optimize research model inference code for use in production preprocessing pipelines.

  • Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.

  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Engineering, Mathematics, or a related technical field; experience in generative video, computer vision, or multimodal ML is strongly preferred.

  • Bonus: Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.

Compensation:

The anticipated annual base salary range for this role is between $200,000-$260,000 (€170,000-€225,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account - $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$96k – $227k per year (Estimated) • In office • Full-Time • Melbourne • Sydney
PowerShell
Python
DevOps
AWS
Azure
CI/CD
GCP
Incident Management
Kubernetes
Platform Engineering
Service Mesh
Terraform
Apply
$121k – $182k per year • Remote/Hybrid • Full-Time • 8+ years exp • Rutherford • Ward
Java
Python
SQL
Databases
ElasticSearch
AI/ML
Devin
DevOps
AWS
Azure
Azure DevOps
CI/CD
CircleCI
Docker
GCP
GitLab
GitLab CI
Grafana
Jenkins
Kibana
Kubernetes
Logstash
Prometheus
Rest API
Splunk
Apply
$34k – $82k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bengaluru • Hyderabad
Python
AI/ML
AI Agents
Anomaly Detection
LangChain
LlamaIndex
LLM
LLM Guardrails
RAG
DevOps
AWS
Azure
CI/CD
GCP
Platform Engineering
Vector
Management
ServiceNow
Apply
$28k – $116k per year (Estimated) • Remote/Hybrid • Full-Time • Bengaluru • Hyderabad
Python
SQL
Python
FastAPI
Pydantic
Databases
FAISS
OpenSearch
Pinecone
SAP HANA
Snowflake
AI/ML
A2A
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
Claude
Claude Code
Cline
Copilot
Cursor
DeepEval
Embeddings
Gemini
Hallucination
LangChain
LangGraph
Llama
LLM
LLM Guardrails
Model Context Protocol
OCR
OpenAI
Promptfoo
RAG
DevOps
AWS
Azure
GCP
GitHub
Management
Power Automate
Marketing
Salesforce
Apply
$95k – $224k per year (Estimated) • In office • Full-Time • Sydney
PowerShell
Python
DevOps
AIOps
Amazon ECS
Ansible
AWS
Grafana
Prometheus
Terraform
VMWare
Cybersecurity
CyberArk
Apply
$175k – $250k per year • In office • Full-Time • 6+ years exp • San Francisco
SQL
Mobile
Braze
Deep Linking
Analytics
A/B Testing
Marketing
Amplitude
Iterable
Mixpanel
Apply
$180k – $270k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sunnyvale • San Francisco
C++
Go
Node JS
JavaScript
AI/ML
Speech Recognition
DevOps
WebRTC
Apply
$200k – $220k per year • In office • Full-Time
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
Knowledge Distillation
Multimodal AI
PyTorch
Synthetic Data
Tokenization
DPO
FSDP
GRPO
LLM Guardrails
Text-to-Speech
Speech Recognition
Apply
$200k – $240k per year • In office • Full-Time • 8+ years exp
Go
DevOps
AWS
Terraform
Apply
$200k – $220k per year • In office • Full-Time
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
Knowledge Distillation
Multimodal AI
PyTorch
Quantization
Synthetic Data
Tokenization
DPO
FSDP
GRPO
LLM Guardrails
Text-to-Speech
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.