386,234open jobs
10,126companies
50,599added this week
Browse all
Salary
$117k – $230k per year (Estimated)
Location
Remote (United States)
Seniority
Staff
Overview
Company
Impact
Profile match
Domino Data Lab is an enterprise data science platform company headquartered in San Francisco, California, and founded in 2013. The company provides an environment where research teams manage compute, environments, experiments, models, and governance across cloud and on-premise infrastructure. It sells mainly to regulated industries such as pharmaceuticals, financial services, and insurance, where reproducibility and model audit trails are requirements.

Who we are

At Domino, we build software that helps the largest, AI-driven organizations build and operate advanced data science and AI solutions at scale. Our platform integrates a streamlined model development environment, MLOps capabilities, and novel features for collaboration, reuse, and reproducibility - all of which make data science teams more productive, reduce time to value, and ensure compliance. Our customers - like Johnson & Johnson, GSK, Bristol Myers, UBS, FINRA and the US Navy - are using our software to solve some of the most important challenges in the world, such as developing new medicines, securing our financial markets, or protecting our country. Backed by Sequoia Capital, Coatue Management, NVIDIA, Snowflake and other leading investors, we have been in business for a decade but are still a small team operating with the spirit of a startup. Especially in the world of AI today, we believe that the future is still being invented - and we want to be the ones building it. For more information, visit www.domino.ai

What we are building

The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently. A core part of this mission is Tempest, an in-house platform that orchestrates realistic, long-duration workloads against live Kubernetes clusters and validates the results against real observability data. Today, when scale testing surfaces a bottleneck, a resource misconfiguration, or a regression in system behavior, the team can identify and report the issue - but we need someone who can take the next step: profiling services, tracing root causes through Prometheus and New Relic data, and partnering with platform engineers to drive durable fixes. Focused on iteration and continuous improvement, the team looks for targeted enhancements that create outsized impact, and this role will close the gap between detection and resolution at the infrastructure level.

What your impact will be

In your first year, you will:

  • Serve as the technical owner of Tempest, Domino's scale and reliability platform, ensuring it remains reliable, extensible, and aligned with evolving infrastructure needs
  • Diagnose and drive resolution of performance bottlenecks and resource misconfigurations surfaced by scale testing - working directly with platform and infrastructure teams to ship fixes, not just file tickets
  • Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on rigorous empirical testing across deployment sizes
  • Strengthen observability across scale testing by improving Prometheus and New Relic instrumentation, making it faster to pinpoint root causes during and after multi-day load runs
  • Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance for this increasingly divergent product line
  • Partner with platform teams to enable effective scale and reliability testing across additional cloud providers, helping position Domino for future multi-cloud success
  • Increase the efficiency and leverage of a small team by building infrastructure automation that scales operationally as the product and customer base grow

What we look for in this role

  • Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments
  • Strong proficiency in Python and comfort working in a large, modular codebase that spans orchestration, infrastructure automation, and systems integration
  • Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) - writing queries, building dashboards, and using metrics to diagnose performance and reliability issues at the systems level
  • Demonstrated ability to go beyond detection to resolution: profiling services, identifying resource bottlenecks, and working with engineering teams to ship durable fixes
  • Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar) as part of a broader infrastructure or reliability practice
  • Clear ownership mindset - self-directed, accountable, and able to communicate priorities and status effectively in a remote, async environment

What we value

  • We value a growth mindset. High-performing creative individuals who dig into problems and see the opportunities for success
  • We believe in individuals who seek truth and speak the truth and can be their whole selves at work
  • We value all of you that believe improving is always possible At Domino Everything is a work in progress - we can do better at everything
  • We emphasize an environment of teaching and learning to equip employees with the tools needed to be successful in their function and the company
  • We strongly believe in the value of growing a diverse team and encourage people of all backgrounds, genders, ethnicities, abilities, and sexual orientations to apply

The annual US base salary range for this role is listed below. For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. This salary range will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location. Additional benefits for this role may include: equity, company bonus or sales commissions/bonuses; 401(k) plan; medical, dental, and vision benefits; and wellness stipends.

Compensation Range

$185,000—$210,000 USD

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,234 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$99k – $132k per year • In office • Full-Time • Heilbronn
SQL
DevOps
Azure
Grafana
Kubernetes
Prometheus
Apply
$17k – $75k per year (Estimated) • In office • Bachelor's Degree • Noida
DevOps
Alertmanager
Blue-Green Deployment
cert-manager
CI/CD
Git
GitOps
Grafana
Helm
Jaeger
Jenkins
JFrog Artifactory
Kubernetes
Loki
Nginx
OpenStack
Prometheus
Sealed Secrets
Traefik
Cybersecurity
kube-bench
kube-hunter
SonarQube
Trivy
Apply
$191k – $397k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
AI/ML
CUDA
CUDA Toolkit
NCCL
DevOps
Kubernetes
Apply
In office • 6+ years exp
C++
Java
Databases
MySQL
PostgreSQL
Redis
DevOps
Docker
Git
Kubernetes
WebSockets
Apply
$108k – $115k per year • In office • Full-Time • 7+ years exp • Ashburn
Java
SQL
TypeScript
JavaScript
Java
Spring Boot
Frontend
Angular
DevOps
AWS
Docker
Git
Kubernetes
Apply
$48k – $114k per year (Estimated) • Remote • Bachelor's Degree
Python
Databases
Snowflake
DevOps
AWS
Azure
GCP
kubectl
Kubernetes
Management
Jira
Apply
$36k – $86k per year (Estimated) • Remote • Bachelor's Degree • Barcelona
Python
Databases
Snowflake
DevOps
AWS
Azure
GCP
kubectl
Kubernetes
Management
Jira
Apply
IT Support Engineer 3 days ago
$12k – $30k per year (Estimated) • Remote • Contractor • 2+ years exp
Databases
Snowflake
DevOps
Azure
Cybersecurity
Okta
Management
Google Workspace
Jira
Slack
Apply
$131k – $268k per year (Estimated) • In office
Python
SQL
Databases
Snowflake
DevOps
Amazon EKS
AWS
Azure
Azure AKS
Docker
GCP
Google GKE
Kubernetes
Apply
$56k – $133k per year (Estimated) • Remote • Bachelor's Degree
Python
Databases
Snowflake
DevOps
AWS
Azure
GCP
kubectl
Kubernetes
Management
Jira
Apply
See all jobs
This is one of many
386,234 more open roles from verified company boards, updated every day.