Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Jun 10, 2026. Roku scores B on the Alion truth index.
Join Roku's Advertising Performance team as a Senior Machine Learning Engineer (DevOps/SRE). In this role, you will support and scale our Machine Learning infrastructure, streamline the end-to-end ML lifecycle, and lead the design and operation of scalable cloud infrastructure for ML workloads. You will also define and enforce observability standards for ML systems and participate in on-call rotation for critical ML training and serving infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling, with a passion for building platforms that accelerate ML experimentation and deployment at internet scale.
Missions
- Lead the design and operation of scalable, production-grade cloud infrastructure for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environments.
- Architect and improve CI/CD systems for ML models and platform services to enable fast, reliable, and safe production releases.
- Define and enforce observability standards for ML systems, including model performance monitoring, drift detection, capacity planning, and pipeline health metrics.
Profil recherché
- The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling - with a passion for building platforms that accelerate ML experimentation and deployment at internet scale- Strong infrastructure-as-code experience with Terraform or similar tooling
- Experience in the Advertising domain is a plus
- Excellent communication and cross-functional collaboration skills
- Strong programming skills in Python, and/or Scala, or Java for platform automation and tooling
- 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML or AI systems
- Deep experience with Kubernetes and container orchestration on GCP (GKE) and/or AWS (EKS)
- Expertise with NoSQL or low-latency data stores such as Aerospike or similar technologies
- BS or MS in Computer Science, Engineering, or a related quantitative field
- Experience building and maintaining CI/CD systems using tools such as Jenkins or GitLab Runner
- Hands-on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka
- Familiarity with feature engineering platforms such as Chronon and model lifecycle tools such as MLflow
- Experience with observability platforms such as Prometheus, Grafana, and Datadog

