690,383open jobs
40,512companies
97,848added this week
Browse all
Salary
$70k – $172k per year (Estimated)
Location
Remote/Hybrid (South San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Roche is a Swiss healthcare group founded in Basel in 1896 and is unusual in operating two large divisions of comparable importance: prescription medicines and in vitro diagnostics. The pharmaceutical business is built on oncology, neurology, immunology, ophthalmology and haemophilia, with products such as Ocrevus, Hemlibra, Perjeta and Vabysmo, while the diagnostics division supplies laboratory analysers, molecular tests and sequencing systems to hospitals worldwide. The group owns the American biotechnology company Genentech outright and the Japanese firm Chugai in majority, is controlled by a family shareholder pool, and spends among the largest research budgets in the industry.

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments.

In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche’s compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully.

Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.

The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute-helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.

The Opportunity:

SLURM Architecture & Ecosystem Leadership

  • Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution.

  • Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.

  • Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures.

Hybrid Workload & Kubernetes Integration

  • Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads.

  • Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.

Technical Leadership, Mentorship & Governance

  • Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.

  • Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter.

  • Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly.

Who You Are:

  • Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.

  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.

  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.

  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.

  • SLURM Architecture & Optimization: Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources.

  • SLURM Operations, Accounting & Observability: Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.

  • Kubernetes & Container Knowledge:Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.

  • AI Infrastructure & Interconnects:Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).

  • Automation:Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.

  • Broad Platform Expertise: Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems spanning multiple technology domains.

  • Domain Expertise & Problem Solving:Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.

  • Coaching & Collaboration:Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.

  • Cross-Organizational Coordination:Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.

  • Strategic Vision:Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.

Where pay transparency applies, details are provided based on the primary posting location. For this role, the primary location is Kaiseraugst. If you are interested in additional locations where the role may be available, we will provide the relevant compensation details later in the hiring process.

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.

Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
690,383 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
South San Francisco
$63k – $162k per year (Estimated) • In office • Full-Time • Brussels
Python
Java
Rust
Bash
Databases
RabbitMQ
Apache Kafka
DevOps
Splunk
Terraform
Ansible
GCP
VMWare
Azure
CI/CD
AWS
Docker
Kubernetes
Linux
Windows
TCP/IP
VPN
Cybersecurity
Nessus
SIEM
Management
Jira
ServiceNow
Apply
Software Engineer 3 hours ago
$72k – $173k per year (Estimated) • In office • Full-Time • Canberra
JavaScript
Java
TypeScript
C#
C#
.NET
Databases
Apache Kafka
Frontend
Angular
DevOps
Rest API
Ansible
CI/CD
Git
Docker
Kubernetes
GitLab
Linux
Windows
Analytics
Apache NiFi
Management
Confluence
Jira
Agile
Apply
$22k – $45k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bengaluru
Python
JavaScript
Java
TypeScript
SQL
Python
FastAPI
Java
Spring Boot
Databases
Apache Kafka
AI/ML
Spark
LLM
Frontend
React.js
Mobile
Clean Architecture
DevOps
Rest API
Terraform
AWS CDK
CI/CD
AWS
Docker
Kubernetes
Apply
$80k – $90k per year • Remote • Full-Time • 5+ years exp • United States
Python
JavaScript
SQL
Bash
Databases
ElasticSearch
OpenSearch
DevOps
Ansible
Red Hat
Zabbix
OpenShift
Kibana
Loki
Fluent Bit
Fluentd
Logstash
Prometheus
CI/CD
Kubernetes
Nginx
Grafana
SLI/SLO/SLA
Linux
Apache HTTP Server
Apply
$90k – $132k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • United States
Python
Go
Java
Java
Flyway
Liquibase
DevOps
Terraform
Helm
GitHub Actions
GitLab CI
CI/CD
ArgoCD
Jenkins
Git
AWS
Kubernetes
SRE
Platform Engineering
JFrog Artifactory
Amazon EKS
GitLab
Cybersecurity
SonarQube
Management
Agile
Apply
$68k – $127k per year • In office • Full-Time • 5+ years exp • High School Diploma • Indianapolis
Management
Google Workspace
Microsoft Office
Apply
AWS Cloud Engineer 4 hours ago
$52k – $94k per year (Estimated) • Remote/Hybrid • Full-Time • Warsaw
AI/ML
AWS Bedrock
DevOps
GitHub Actions
AWS CDK
GitLab CI
CI/CD
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon CloudWatch
API Gateway
DNS
Cybersecurity
Least Privilege
Apply
$111k – $222k per year (Estimated) • Remote/Hybrid • Full-Time • San José
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Hyderabad
Databases
DynamoDB
Amazon Redshift
AI/ML
Prompt Engineering
AI Agents
AWS Bedrock
RAG
Amazon SageMaker
DevOps
AWS
Kubernetes
Platform Engineering
Amazon EKS
AWS Lambda
Vector
Amazon S3
Amazon ECS
Amazon EventBridge
API Gateway
Analytics
ETL/ELT
AWS Glue
Apply
$26k – $60k per year (Estimated) • In office • Full-Time • Hyderabad
Apply
$215k – $245k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • South San Francisco
DevOps
SLI/SLO/SLA
Analytics
Microsoft Excel
Management
Agile
Apply
Scientist 1 day ago
$129k – $175k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • South San Francisco
AI/ML
Image Segmentation
Apply
$123k – $166k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • South San Francisco
Apply
$90k – $115k per year • In office • 2+ years exp • South San Francisco
DevOps
Vector
Apply
$90k – $100k per year • Remote • Contractor • Bachelor's Degree • South San Francisco
Apply
See all jobs
This is one of many
690,383 more open roles from verified company boards, updated every day.