368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$61k – $120k per year (Estimated)
Location
Remote (Poland)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
ALTER GPU CENTER is a cloud computing infrastructure company. It provides GPU-based computing resources and related services for high-performance workloads.

About the role

We are looking for a Lead DevOps Engineer to provide technical leadership for DevOps and Site Reliability Engineering practices supporting large-scale GPU infrastructure used for AI training and inference workloads.

This role combines hands-on engineering with team leadership. You will be responsible for shaping automation standards, improving platform reliability, and leading a team working on software-defined infrastructure, high-performance networking, observability, and operational excellence across complex production environments.

Responsibilities

  • Lead, mentor, and support a team of DevOps and SRE engineers working across the full lifecycle of GPU infrastructure platforms

  • Design and implement Infrastructure as Code solutions for provisioning and managing bare-metal GPU servers, networking, storage, and cluster orchestration components

  • Build and improve CI/CD pipelines for infrastructure, platform services, and internal tooling

  • Develop and maintain monitoring, logging, alerting, and observability solutions for large-scale GPU environments

  • Define and track SLIs/SLOs, improve incident response processes, and contribute to post-incident reviews and long-term reliability improvements

  • Work closely with Infrastructure, Networking, Facilities, and AI/ML teams to ensure stable and scalable platform operations

  • Automate operational processes such as cluster scaling, firmware and BIOS updates, hardware diagnostics, and capacity planning

  • Support DevSecOps practices, including infrastructure hardening, vulnerability management, and compliance automation

  • Identify operational inefficiencies and reduce repetitive manual work through automation

  • Evaluate and introduce new tools and solutions related to GPU infrastructure, orchestration, and cloud-native operations

Requirements

  • 8+ years of experience in DevOps, SRE, Platform Engineering, or a similar area

  • At least 3 years of experience in a technical lead, lead engineer, or team leadership role

  • Strong practical experience with infrastructure automation in large-scale or complex production environments

  • Very good knowledge of Terraform, Ansible, Pulumi, Crossplane, or similar Infrastructure as Code tools

  • Experience with GitOps, configuration management, and CI/CD practices

  • Hands-on experience with Kubernetes

  • Experience working with GPU-related technologies such as NVIDIA GPU Operator, device plugins, MIG, or time-slicing

  • Good scripting or programming skills in Python, Go, or Bash

  • Experience with bare-metal provisioning, infrastructure automation, or data center environments

  • Good knowledge of observability tools such as Prometheus, Grafana, Loki, and OpenTelemetry

  • Good understanding of distributed systems reliability and production incident management

  • Experience with high-performance networking technologies such as RDMA, InfiniBand, or RoCE will be a strong advantage

  • Ability to lead technical discussions, support team development, and communicate effectively with both technical and business stakeholders

  • English proficiency at least at a communicative level is required, as you will be working in an international team

Nice to have

  • Experience in AI infrastructure, HPC environments, hyperscale infrastructure, or data center operations

  • Familiarity with orchestration and scheduling tools such as Slurm, Ray, Run:ai, KServe, or Kubernetes-based schedulers

  • Experience integrating telemetry from power, cooling, or environmental systems

  • Experience building internal platforms or self-service tools for engineering or research teams

  • Understanding of security, compliance, and audit requirements in regulated or security-sensitive environments

What we offer

  • Benefits package

  • Opportunity to shape the DevOps and SRE foundation for advanced GPU infrastructure supporting AI workloads

  • Real impact on the scalability, reliability, and operational standards of next-generation compute environments

  • Collaboration with experienced engineers across infrastructure, platform, and AI domains

  • A dynamic environment with space for ownership, technical leadership, and professional growth

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Łódź
$17k – $42k per year (Estimated) • Remote • Moscow
SQL
Databases
MS SQL
PostgreSQL
DevOps
Ansible
CI/CD
Jenkins
Prometheus
Zabbix
GitLab
Apply
$27k – $65k per year (Estimated) • Remote • Contractor • 3+ years exp • Moscow
Databases
MySQL
Redis
DevOps
Ansible
CI/CD
Docker
Grafana
HAProxy
Istio
Prometheus
SLI/SLO/SLA
Terraform
VMWare
Zabbix
GitLab
Apply
$54k – $64k per year • Remote/Hybrid • Internship • Gdańsk
Python
Scala
SQL
AI/ML
Spark
Amazon SageMaker
DevOps
AWS
CI/CD
Apply
$32k – $47k per year • In office • Full-Time • 3+ years exp • Bengaluru
JavaScript
Node JS
Python
Node JS
Nest.JS
Databases
Memcached
MySQL
PostgreSQL
Redis
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
QA
Jest
Mocha
Apply
$14k – $30k per year (Estimated) • In office • Internship • Chelyabinsk
PHP
Databases
MySQL
DevOps
CI/CD
Docker
Git
GitHub
Apply
$29k – $74k per year (Estimated) • Remote • Full-Time • Poland
Python
SQL
Python
Django
Databases
PostgreSQL
DevOps
GitHub
GitLab
Apply
$44k – $84k per year (Estimated) • Remote • Full-Time • 4+ years exp • Łódź
Bash
Python
AI/ML
CUDA Toolkit
cuDNN
InfiniBand
NCCL
DevOps
Configuration Management
Debian
Grafana
Kubernetes
Prometheus
Red Hat
SLURM
Ubuntu
HPC
Apply
DevOps Engineer 1 month ago
$46k – $87k per year (Estimated) • Remote • Full-Time • 4+ years exp • Łódź
Python
AI/ML
KServe
Ray
DevOps
Ansible
CI/CD
GitOps
Grafana
Kubernetes
Loki
OpenTelemetry
Platform Engineering
Prometheus
SLURM
Terraform
HPC
Cybersecurity
Crowdstrike
Snyk
Apply
$61k – $120k per year (Estimated) • Remote • Full-Time • 7+ years exp • Łódź
Bash
Python
AI/ML
CUDA Toolkit
cuDNN
InfiniBand
NCCL
DevOps
Ansible
Configuration Management
Kubernetes
SLURM
Terraform
HPC
Apply
Front-end Developer 1 month ago
$52k – $87k per year (Estimated) • Remote • Part-Time • Łódź
JavaScript
TypeScript
Frontend
React.js
DevOps
Docker
GitHub
GitLab
Apply
$69k – $111k per year (Estimated) • Remote/Hybrid • Full-Time • Łódź
Python
SQL
Databases
Snowflake
AI/ML
Streamlit
DevOps
Azure
Analytics
Power BI
Apply
UX Designer 6 days ago
In office • Full-Time • Łódź
AI/ML
Copilot
Midjourney
Marketing
Hotjar
Apply
$68k – $122k per year (Estimated) • Remote • Full-Time • Łódź
Cybersecurity
ISO 27001
Apply
$54k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Łódź
C#
SQL
C#
.NET
Entity Framework Core
AI/ML
Semantic Kernel
OpenAI
DevOps
Azure
Azure DevOps
CI/CD
Grafana
OpenTelemetry
Rest API
Apply
$32k – $38k per year • Remote • Łódź
PHP
TypeScript
JavaScript
PHP
Symfony
AI/ML
Claude
Claude Code
Frontend
Angular
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.