658,489open jobs
38,337companies
96,793added this week
Browse all
Salary
$77k – $127k per year (Estimated)
Location
Remote (Poland)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead - GPU Infrastructure based in Poland.

This is a hands-on technical leadership role responsible for building and delivering a full-stack GPU infrastructure platform.

You will own architecture and implementation across bare-metal GPU scheduling, Kubernetes, managed inference, and platform observability.

The role combines deep infrastructure engineering with leadership of a distributed team spanning backend, frontend, DevOps, QA, and documentation.

You will help operate GPU compute environments for research, model training, and inference workloads at scale.

A key focus will be reliable GPU fleet operations, high-performance networking, multi-tenant infrastructure, and production-grade platform services.

You will also serve as the primary technical interface for infrastructure partners while translating internal workload requirements into robust platform capabilities.

The position is fully remote and offers the opportunity to shape a critical AI infrastructure platform from architecture through production delivery.

Accountabilities

    • Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, technical reviews, and maintaining the architecture baseline.
    • Lead and line-manage a distributed engineering team across backend, frontend, DevOps, QA, and documentation, setting engineering standards and overseeing code reviews, design reviews, release gates, one-to-ones, and performance development.
    • Design, build, and operate a managed Slurm service for research users, covering controllers, accounting, partitions, login nodes, node onboarding, NVIDIA drivers and CUDA baselines, job and node-health monitoring, autohealing, storage visibility, identity, and workload isolation.
    • Own Kubernetes cluster bootstrap and lifecycle on bare-metal infrastructure, including NVIDIA GPU and Network Operators, KubeVirt and VFIO-based GPU isolation, upgrades, backup and recovery, and node replacement.
    • Define and deliver managed inference architecture covering serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable infrastructure.
    • Establish observability and operational practices across the control plane, GPU fleet, and application tiers, including metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model.
    • Act as the primary technical interface with infrastructure partners and vendors, translating requirements into specifications and acceptance tests, managing escalations, and contributing to capacity planning and hardware sourcing.
    • Work directly with research, model-training, and product teams to understand workloads, translate requirements into platform capabilities, and manage capacity constraints.
    • Hire additional members of the platform team and establish the technical standards and expectations for future engineering hires.
    • Maintain a hands-on contribution to architecture, technical reviews, implementation decisions, and infrastructure delivery rather than operating solely in a management capacity.
    • Requirements:

      • Bring 8+ years of hands-on engineering experience, including at least 3 years leading teams that build and operate infrastructure platforms used by other teams.
      • Hold a Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or demonstrate equivalent practical experience.
      • Have hands-on experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node-health scripting, and upgrades while jobs remain active; experience with HPC or GPU training clusters is highly desirable.
      • Demonstrate deep experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle management, Fabric Manager, NVSwitch on SXM systems, DCGM-based health and utilization, MIG, node burn-in, and acceptance processes.
      • Have strong knowledge of high-performance interconnects, including InfiniBand fabrics, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues.
      • Possess advanced Linux systems knowledge covering kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning for compute-intensive workloads.
      • Have production Kubernetes operations experience beyond deployment, including control planes, upgrades, CNI and CSI, operators, custom controllers, and multi-tenancy architecture.
      • Understand HPC storage and large-scale data movement, including shared filesystems such as VAST, Lustre, or NFS, node-local NVMe caching, and distributing large model weights and datasets across multiple nodes.
      • Be experienced with observability and production operations using Prometheus, Grafana, Loki, or equivalent technologies, together with SLO management, incident response, and post-incident reviews.
      • Have working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions, without requiring feature-development specialization.
      • Have shipped a platform with real users, such as a multi-tenant IaaS/PaaS, research computing service, or comparable infrastructure platform involving resource isolation, quotas, usage metering, and user-facing APIs or CLI surfaces.
      • Demonstrate leadership that remains technically engaged, including people management across time zones, cross-track technical reviews, documented architecture decisions, and the confidence to challenge partners or executives with clear technical reasoning.
      • Have excellent written and spoken English, particularly for technical, partner, and leadership communication.
      • Be based within the UTC to UTC+5:30 time-zone range to provide working-hour overlap with teams and partners in Europe and India, with availability for occasional travel to partner sites and team events.
      • Experience with Slurm operators on Kubernetes, Kubernetes-native schedulers, modern AI serving stacks such as vLLM, SGLang, or TensorRT-LLM, and GPU parallelism or quantization strategies is a plus.
      • Experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, or GPU autohealing technologies is desirable.
      • Background working on GPU cloud platforms, university or national HPC centers, AI lab infrastructure teams, peer-to-peer or distributed systems, or hardware-provider relationships is also considered an advantage.
      • Benefits:

        • Fully remote work within the specified UTC to UTC+5:30 time-zone range.
        • The opportunity to lead architecture and delivery of a full-stack GPU infrastructure platform spanning bare-metal compute, Kubernetes, Slurm, and managed AI inference.
        • A hands-on technical leadership position combining engineering, architecture, people management, and infrastructure-partner engagement.
        • Collaboration with a distributed international team across Europe and India.
        • Exposure to advanced GPU infrastructure, high-performance networking, AI workloads, multi-tenant compute, and production inference systems.
        • Opportunities to shape engineering standards, platform architecture, hiring, and long-term infrastructure capabilities.
        • Occasional travel opportunities to partner sites and team events.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
658,489 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Remote/Hybrid • Full-Time • 5+ years exp • Moscow
Python
JavaScript
Java
TypeScript
SQL
Python
pySpark
Java
Spring Framework
Databases
RabbitMQ
Apache Kafka
AI/ML
Spark
Frontend
Angular
DevOps
Jaeger
Rancher
Prometheus
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Grafana
Analytics
Data Vault
Management
Confluence
Jira
Agile
Scrum
Kanban
QA
Cypress
Playwright
Apply
$21k – $49k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad
PHP
Prolog
Apply
In office • 9+ years exp • Bachelor's Degree • Bengaluru
Python
Java
Databases
Apache Kafka
DevOps
gRPC
GCP
Azure
CI/CD
Jenkins
Git
AWS
Kubernetes
Management
Agile
Apply
$142k – $254k per year (Estimated) • Remote • Secret • Full-Time • 5+ years exp • Master's Degree
Python
Java
C++
Scala
DevOps
Docker
Kubernetes
Apply
$135k – $155k per year • Remote • Full-Time • 7+ years exp
Python
JavaScript
TypeScript
Python
Django
Django REST Framework
Databases
PostgreSQL
Frontend
Zustand
Redux
React.js
Vite
npm
DevOps
GCP
GitHub Actions
CI/CD
Kubernetes
Management
Agile
QA
Playwright
Jest
Pytest
Vitest
Apply
$88k – $173k per year (Estimated) • Remote • Full-Time • 5+ years exp
SQL
DevOps
IAM
Cybersecurity
GDPR
Least Privilege
Apply
$28k – $64k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
SQL
Mobile
Clean Architecture
DevOps
Rest API
Terraform
CloudFormation
Pulumi
Azure
CI/CD
AWS
Bicep
IAM
Apply
$104k – $217k per year (Estimated) • Remote • Full-Time • 5+ years exp
SQL
DevOps
IAM
Cybersecurity
GDPR
Least Privilege
Apply
$69k – $147k per year (Estimated) • Remote • Full-Time • 5+ years exp
SQL
DevOps
IAM
Cybersecurity
GDPR
Least Privilege
Apply
$54k – $123k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
SQL
Mobile
Clean Architecture
DevOps
Rest API
Terraform
CloudFormation
Pulumi
Azure
CI/CD
AWS
Bicep
IAM
Apply
See all jobs
This is one of many
658,489 more open roles from verified company boards, updated every day.