682,105open jobs
39,475companies
99,595added this week
Browse all
Location
Remote (Georgia)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
We provide powerful solutions that will help your business grow globally. Try our superior performance for free.

The world’s digital experiences run on something invisible: the infrastructure and software that keep them fast, reliable, and secure. At Gcore, you’ll help design and deliver that foundation for an AI-driven world.

We’re a global provider of infrastructure and software solutions for AI, cloud, network, and security, powering everything from real-time communication and streaming to enterprise AI and secure web applications. With 210+ edge locations, 50+ cloud regions, and thousands of GPUs, your work here can reach users and businesses across the globe.

You’ll collaborate with leading technology partners such as Intel, NVIDIA, Dell, and Equinix, and work on platforms that power digital products used around the world. Our vision is simple: to connect the world to AI, anywhere, anytime.

Want to work on technology that goes beyond a single product or industry? Join a global team of 550+ professionals building infrastructure and software that supports the entire digital ecosystem.

What You’ll Do

  • Design and build a managed Slurm service on Kubernetes
  • Write clean, reliable, and maintainable Go code
  • Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads
  • Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, DCGM

What We're Looking For

  • Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
  • Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
  • Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
  • Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
  • A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
  • Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges

Nice to Have

  • Experience operating large-scale HPC or GPU clusters for external customers
  • Experience with PyTorch distributed training and other large-scale AI/ML frameworks
  • Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
  • Experience building unified job-submission workflows across Kubernetes and Slurm
  • Experience in GPU-cloud or HPC product engineering environments
  • Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects

Benefits

At Gcore, we want you to do your best work and enjoy the journey. Our benefits are designed to support your growth, well-being, and life beyond work:

  • Competitive compensation
  • Flexible working hours and hybrid or remote options, depending on your role
  • Work from anywhere in the world for up to 45 days per year
  • Private medical insurance for you and your family*
  • Extra paid vacation and sick leave days*
  • Support for life’s important moments and celebrations
  • Language courses to help you connect and grow
  • Modern, welcoming offices with snacks, drinks, and entertainment*
  • Team sports and social activities*

*Benefits may vary depending on your location.

Equal Opportunity Employer

We provide equal opportunity to all applicants without regard to race, color, religion, sex, sexual orientation, age, gender identity, gender expression, national origin, disability, or any other legally protected characteristics.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
682,105 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Cyprus
$43k – $90k per year (Estimated) • In office • 10+ years exp • Master's Degree • Pune
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
Redis
Google BigQuery
BigQuery
Frontend
Angular
React.js
DevOps
GCP
GitHub Actions
New Relic
OpenTelemetry
Datadog
Azure
CI/CD
AWS
Kubernetes
Google GKE
Google Cloud Run
API Gateway
Management
Agile
Scrum
Apply
Platform Engineer 1 day ago
$108k – $239k per year (Estimated) • Remote/Hybrid • Tysons
DevOps
Terraform
GCP
CloudFormation
Pulumi
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Configuration Management
Cybersecurity
Zero Trust
Apply
$84k – $125k per year • In office • 8+ years exp • Bachelor's Degree • Tysons
Python
Java
C#
C#
.NET
Databases
MS SQL
DevOps
GCP
Kubernetes
Management
Jira
Agile
QA
Selenium
Apply
$145k – $250k per year • Remote/Hybrid • Top Secret • Tysons
Python
JavaScript
TypeScript
Databases
PostgreSQL
AI/ML
Model Context Protocol
NLP
Frontend
Angular
DevOps
GCP
OpenShift
Azure
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
Google GKE
Azure AKS
GitLab
Apply
$118k – $238k per year (Estimated) • Remote/Hybrid • Secret • Colorado Springs
JavaScript
TypeScript
Node JS
Frontend
React.js
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
Management
Agile
Scrum
Kanban
Apply
$32k – $84k per year (Estimated) • Remote • Full-Time • Serbia
Python
Python
FastAPI
Django
Databases
PostgreSQL
ClickHouse
DevOps
Ansible
Prometheus
CI/CD
Docker
Grafana
Configuration Management
Apply
$19k – $45k per year (Estimated) • Remote • Full-Time • 1+ year exp • Poland
Apply
DCIM Engineer 7 days ago
$62k – $120k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • United States
Apply
$44k – $108k per year (Estimated) • Remote • Full-Time • 2+ years exp • Poland
AI/ML
Edge AI
Management
Jira
Kanban
Apply
$40k – $100k per year (Estimated) • Remote • Full-Time • 5+ years exp • Serbia
Python
AI/ML
LoRA
vLLM
CUDA Toolkit
Quantization
Multimodal AI
SGLang
TensorRT
PEFT
TensorRT-LLM
PyTorch
CUDA
Triton
Speculative Decoding
DevOps
Docker
Kubernetes
Apply
$28k – $62k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Manila • Dubai • Amman • São Paulo • Santo Domingo
AI/ML
AI Agents
NLP
NER
RAG
Hallucination
Semantic Search
Time Series Forecasting
Semantic Search
Knowledge Graph
LLM Guardrails
Agentic Workflows
DevOps
Azure
AWS
Vector
FinOps
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Apply
Engineering Manager 10 days ago
Remote • Full-Time • 4+ years exp • Cyprus
Python
Rust
Databases
ClickHouse
DevOps
Prometheus
Kubernetes
Nginx
Grafana
Apply
$26k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Cyprus
SQL
AI/ML
LLM
Analytics
A/B Testing
QA
Postman
Apply
Financial Analyst 14 days ago
$24k – $71k per year (Estimated) • In office • Full-Time • Cyprus
SQL
Analytics
Power BI
Apply
B2B Marketing Lead 17 days ago
$39k – $92k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Cyprus
AI/ML
Claude
ChatGPT
Apply
See all jobs
This is one of many
682,105 more open roles from verified company boards, updated every day.