599,254open jobs
27,958companies
85,677added this week
Browse all
Salary
$180k – $220k per year
Location
In office (New York)
Seniority
Senior
Overview
Company
Impact
Profile match
Udio is a music generation company founded in 2023 by former Google DeepMind researchers. Its models produce full songs with vocals and instrumentation from a text description, and the product includes editing tools for extending and remixing tracks. After litigation with major labels the company moved towards licensed deals with rights holders.

About the Role

We are looking for a Senior Backend Engineer to lead the unification of large, highly rich, and heterogeneous datasets sourced from a wide range of external providers. These datasets are used to power our generative audio models. 

Your work will create the foundational dataset that powers our research by building robust, scalable systems for linking, deduplicating, reconciling, and enriching data at massive scale. This role centers on high-impact bulk ingestion and advanced data linkage. You will design the logic, algorithms, and strategies that transform many independent datasets into a unified, high-quality canonical asset used throughout the company.

You will collaborate closely with ML researchers and product teams, working with tools such as BigQuery, Dataflow/Beam, TFRecords, and-where beneficial-distributed systems frameworks like Ray. Familiarity with ML workflows using JAX or multihost training is a plus, as the datasets you produce will directly support that ecosystem.

What You'll Do

  • Build high-throughput bulk ingestion workflows  to integrate datasets from multiple external providers. 
  • Design and implement scalable entity-resolution  solutions, including record linking, deduplication, clustering, and conflict arbitration. 
  • Create and refine matching logic, decision rules, and similarity functions  to align datasets with high accuracy and strong coverage. 
  • Define and track data quality indicators, such as overlap metrics, match precision/recall, duplicate rates, and completeness. 
  • Prepare training-ready datasets in formats such as TFRecords, and structure data to meet ML research requirements. 
  • Develop processing components using  Dataflow (Beam) and manage large analytical workloads in BigQuery
  • Leverage frameworks like Ray  to accelerate large-scale experiments, feature extraction, and research-oriented data preparation. 
  • Collaborate with ML researchers to anticipate downstream requirements and evolve linkage strategies as new sources and use cases emerge. 

What We're Looking For 

  • Experience working with large, heterogeneous datasets from multiple providers or domains. 
  • Strong background in entity resolution, deduplication, data unification, or related large-scale data integration techniques. 
  • Proficiency in Python, with an emphasis on efficient, scalable data processing. 
  • Experience with BigQuery, Google Dataflow/Apache Beam, or similar batch-processing frameworks. 
  • Familiarity with data validation, normalization, reconciliation, and building consistent views across diverse data sources. 
  • Ability to craft well-structured matching and decision strategies  that balance accuracy, completeness, and computational efficiency. 
  • Comfortable iterating quickly on pragmatic solutions, balancing correctness with time-to-delivery. 
  • Clear communication skills and the ability to collaborate closely with ML and research teams. 

 Nice to Have

  • Knowledge of architecting Google Cloud Platform systems at scale
  • Experience with distributed compute frameworks such as Ray, Spark, or Flink
  • Understanding of JAX-based ML pipelinesmultihost training setups,  or large-scale data preparation for accelerator-backed workflows. 
  • Familiarity with TFRecords  or other high-volume training data formats. 
  • Exposure to ranking, clustering, or statistical similarity modeling. 
  • Experience with Go, NextJS, and/or React Native to contribute to full-stack development

Why Join Us

  • You will design the  core dataset  that underpins our research, product development, and generative audio models. 
  • You'll work on large-scale data challenges that require creativity, algorithmic thinking, and engineering excellence.
  • You'll join a small, fast-moving team where your decisions shape the direction of our data and research capabilities.

Benefits

  • Highly competitive salary and equity 
  • Quarterly productivity budget
  • Flexible time off
  • Fantastic office location in Manhattan
  • Productivity package, including ChatGPT Plus, Claude Code, and Copilot
  • Top notch private health, dental, and vision insurance for you and your dependents
  • 401(k) plan options with employer matching 
  • Concierge medical/primary care through One Medical and Rightway
  • Mental health support from Spring Health
  • Personalized life insurance, travel assistance, and many other perks

Udio’s success hinges on hiring great people and creating an environment where we can be happy, feel challenged, and do our best work. 

Udio provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

This role is eligible for a compensation package of base salary, equity, and benefits. The starting base salary range for this role is $180,000 - $220,000. Actual salary may vary based on level, work experience, performance, and other factors evaluated during the hiring process.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
599,254 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$90k – $140k per year • Remote • Full-Time • 1+ year exp • Bachelor's Degree • United States
Python
JavaScript
TypeScript
SQL
Databases
Google BigQuery
BigQuery
AI/ML
LangGraph
LangChain
Model Context Protocol
Vertex AI
Embeddings
Prompt Engineering
Function Calling
AI Agents
Gemini
LLM
RAG
Google ADK
Hallucination
LLM Evaluation
LLM Guardrails
Multi-Agent Systems
Tool Use
DevOps
GCP
Azure
CI/CD
Git
AWS
Google Cloud Run
Vector
Cybersecurity
MITRE ATT&CK
Management
Agile
Apply
$39k – $89k per year (Estimated) • In office • Full-Time • 4+ years exp • Buenos Aires
Python
SQL
Databases
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Scikit-learn
Pandas
NumPy
Time Series Forecasting
Analytics
Tableau
Power BI
Seaborn
Matplotlib
Looker
Apply
$52k – $64k per year • In office • Full-Time • 4+ years exp • Boulogne-Billancourt
Python
SQL
Databases
ElasticSearch
Google BigQuery
BigQuery
DevOps
Terraform
Puppet
Ansible
GCP
Helm
GitHub Actions
Istio
Kibana
Logstash
Prometheus
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
Spinnaker
Service Mesh
Google GKE
Google Cloud Run
IAM
Cybersecurity
ISO 27001
Management
Agile
Apply
$49k – $98k per year • Remote/Hybrid • Full-Time • Vilnius
Python
JavaScript
Node JS
Databases
Google BigQuery
BigQuery
Frontend
Vue.js
React.js
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Vector
Management
Slack
Confluence
Jira
Apply
Data Engineer 1 day ago
Remote/Hybrid
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
MS SQL
Google BigQuery
BigQuery
AI/ML
Spark
Pandas
DevOps
Terraform
GCP
GitHub Actions
CI/CD
Jenkins
Git
Analytics
SSIS
Dimensional Modeling
Apply
$150k – $225k per year • Remote • 5+ years exp • New York
AI/ML
Copilot
ChatGPT
Claude Code
Udio
Design
Figma
Apply
$250k – $350k per year • Remote • 5+ years exp • PhD • New York
AI/ML
Copilot
ChatGPT
Claude Code
JAX
TensorFlow
Udio
Post-training
TPU
DevOps
GCP
Kubernetes
Apply
$160k – $350k per year • In office • Los Angeles
AI/ML
Copilot
ChatGPT
Claude Code
Udio
Apply
$180k – $220k per year • In office • New York
Python
Go
JavaScript
SQL
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Spark
ChatGPT
Claude Code
Udio
Frontend
Next.js
React.js
Mobile
React Native
Analytics
ETL/ELT
Apply
Remote • Full-Time • Associate's Degree • Washington • New York
Management
Agile
Apply
$101k – $263k per year (Estimated) • In office • Full-Time • 3+ years exp • New York
AI/ML
Claude
ChatGPT
Gemini
DevOps
Vercel
Design
Adobe Photoshop
Figma
Adobe After Effects
Apply
$140k – $196k per year • Equity • In office • 8+ years exp • New York
Marketing
Reddit
Apply
$63k – $140k per year • Equity • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Boston • New York
Apply
$121k – $338k per year • In office • Full-Time • 8+ years exp • Washington • Atlanta • Hartford • Boston • Miami
SQL
Management
Confluence
Jira
Agile
Scrum
Apply
See all jobs
This is one of many
599,254 more open roles from verified company boards, updated every day.