368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$80k – $203k per year (Estimated)
Location
Remote (Ireland)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure) based in Ireland.

This is a highly hands-on infrastructure role focused on deploying, commissioning, and operating GPU cloud environments across regional and core datacenters.

You will turn validated architectures and bills of materials into production-ready infrastructure spanning hardware, networking, storage, Linux, and platform software.

The role sits at the intersection of datacenter operations, GPU infrastructure, network engineering, and cloud platform operations.

You will work with high-density NVIDIA GPU systems, advanced networking, storage platforms, Kubernetes, virtualization, and observability tooling.

As a practical technical escalation point, you will troubleshoot complex issues across physical and software layers and drive incidents through resolution.

You will also help establish deployment standards, validation procedures, documentation, and operational practices for a rapidly evolving AI infrastructure environment.

The role offers broad technical ownership in an international, fast-moving setting where hands-on execution and operational excellence are essential.

Accountabilities

    • Datacenter deployment: Coordinate deployments with datacenter providers, integrators, logistics teams, vendors, and internal engineering; validate rack layouts, power, cooling, airflow, cabling, labeling, and physical readiness.
    • Rack and infrastructure commissioning: Support rack-and-stack activities for GPU and CPU servers, storage, switches, routers, firewalls, PDUs, serial/OOB systems, and supporting infrastructure.
    • Cabling and connectivity: Validate fiber and copper cabling, optics, transceivers, breakout cables, port mappings, link speeds, redundancy, and management, storage, north-south, and east-west connectivity.
    • Hardware bring-up: Commission GPU servers, storage nodes, and platform infrastructure while validating BIOS, BMC, firmware, NICs, DPUs, GPUs, NVMe, RAID/HBA, PCIe topology, NUMA, thermals, power, and hardware health.
    • Hardware validation: Execute burn-in, stress, network, storage, and acceptance testing before production handover; troubleshoot issues involving GPUs, DPUs, NICs, optics, memory, disks, firmware, and BIOS.
    • Network deployment support: Work with network engineering to validate switch configurations, routing, VLAN/VRF segmentation, BGP, ECMP, EVPN/VXLAN, OVS/OVN, VyOS, firewalls, WAF infrastructure, and customer connectivity.
    • AI networking: Support validation of RoCE/RDMA fabrics for distributed AI workloads and troubleshoot issues such as link flaps, MTU mismatches, route errors, packet loss, PFC/ECN problems, and congestion.
    • Platform installation: Install and validate Ubuntu/Linux environments, NVIDIA drivers, CUDA, OFED or inbox drivers, Docker/containerd, KVM/QEMU, platform agents, and GPU infrastructure components.
    • Cloud and Kubernetes environments: Support CloudStack, Kubernetes, KubeVirt, GPU Operator, CSI/CNI integrations, GPU passthrough, SR-IOV, BlueField DPUs, VM networking, and container networking.
    • Storage integration: Support integration and validation of StorPool, Weka, local NVMe, and other supported storage platforms.
    • Operational readiness: Execute acceptance testing, produce deployment readiness reports, maintain runbooks, and ensure infrastructure is fully operational before customer or production handover.
    • Day-2 operations: Perform controlled firmware, OS, driver, BIOS, switch, and hardware maintenance while supporting production incidents and infrastructure escalations.
    • Incident management: Investigate operational failures, perform root-cause analysis, distinguish temporary workarounds from permanent fixes, and work with engineering to eliminate recurring issues.
    • Observability: Validate telemetry and monitoring across hosts, GPUs, DPUs, switches, storage, and platform components using tools such as Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, and NVIDIA NetQ or equivalents.
    • Performance validation: Establish baselines for GPU, network, storage, and host performance and support benchmarking and infrastructure validation.
    • Documentation: Maintain accurate as-built records covering rack elevations, cable maps, port mappings, serial numbers, asset records, IP allocations, changes, and operational procedures.
    • Cross-functional coordination: Partner with infrastructure, networking, storage, platform, fleet automation, observability, product engineering, sales engineering, and service delivery teams.
    • Vendor management: Coordinate with datacenter providers, system integrators, server and storage vendors, NVIDIA, and networking suppliers to resolve deployment and infrastructure issues.
    • Continuous improvement: Feed field experience back into reference architectures, BOMs, rack designs, cabling standards, deployment playbooks, validation processes, and automation.
    • Requirements

      • Datacenter infrastructure: Strong hands-on experience deploying and maintaining datacenter infrastructure, ideally within GPU, HPC, AI cloud, private cloud, or high-density compute environments.
      • Bare-metal deployment: Proven ability to bring servers from physical installation and bare metal through validation and production readiness.
      • GPU infrastructure: Experience with NVIDIA GPU servers, drivers, firmware, PCIe topology, hardware validation, and high-performance compute environments.
      • Next-generation AI infrastructure: Familiarity with NVL72-style rack-scale architectures, NVLink/NVSwitch domains, in-rack networking, high-density power delivery, and OEM/NVIDIA validation requirements.
      • Datacenter readiness: Ability to assess power density, cooling, rack dimensions, floor loading, containment, serviceability, maintenance access, and other physical requirements for AI infrastructure.
      • Linux: Strong Linux troubleshooting capabilities and experience managing operating systems, kernels, drivers, and hardware interfaces.
      • Networking: Practical knowledge of VLANs, VRFs, BGP, ECMP, OVS/OVN, routing, OOB management, and high-speed datacenter connectivity.
      • GPU networking: Familiarity with NVIDIA/Mellanox networking, RoCE/RDMA, SR-IOV, BlueField DPUs, and high-performance east-west infrastructure.
      • Virtualization and containers: Experience with KVM/QEMU, VFIO, PCI passthrough, Docker/containerd, Kubernetes, and/or KubeVirt.
      • Storage: Experience integrating or troubleshooting local NVMe, storage nodes, and enterprise or distributed storage platforms.
      • Automation: Familiarity with Terraform, Ansible, Bash, and/or Python for deployment, validation, configuration, or operational automation.
      • Observability: Experience with infrastructure monitoring, telemetry, logs, metrics, health checks, and performance dashboards.
      • Documentation: Strong attention to detail and discipline in producing accurate as-built documentation, runbooks, validation records, and handover materials.
      • Troubleshooting: Strong systems-thinking ability across physical infrastructure, hardware, firmware, networking, Linux, storage, and platform layers.
      • Operational mindset: Comfortable supporting production environments, deployment windows, operational escalations, and customer-impacting incidents.
      • Communication: Able to clearly explain technical issues, risks, workarounds, and permanent solutions to engineering teams, vendors, and leadership.
      • Personal qualities: Highly practical, detail-oriented, calm under pressure, autonomous, and comfortable working both inside datacenters and remotely with smart-hands teams.
      • Benefits

        • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.
        • Full-time or contract engagement, depending on the agreed arrangement.
        • Europe-based remote working environment with flexibility.
        • Opportunity to work on cutting-edge GPU cloud and AI infrastructure at significant scale.
        • Hands-on exposure to NVIDIA GPU platforms, high-density datacenter environments, RoCE/RDMA networking, Kubernetes, virtualization, storage, and advanced observability.
        • Broad cross-functional scope spanning hardware, datacenter operations, networking, storage, Linux, and cloud platforms.
        • High-impact role within a fast-growing international scale-up.
        • Strong opportunities for technical growth and career development as the infrastructure platform expands.
        • Friendly, diverse, flexible, and international working environment.
        • Inclusive workplace committed to equal opportunity and respect for all qualified candidates.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$133k – $181k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Cooper
Bash
PowerShell
Python
DevOps
Ansible
Azure
Configuration Management
Docker
Kubernetes
OpenShift
Ubuntu
Windows Server
Cybersecurity
CIS Benchmarks
Defense in Depth
Nessus
Qualys Cloud Platform
Zero Trust
Apply
$11k – $14k per year • In office • 2+ years exp • Almaty
Bash
Python
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
DevOps
Ansible
ArgoCD
AWS
Azure
CentOS Stream
CI/CD
Docker
FluxCD
Git
GitHub Actions
GitLab CI
GitOps
Grafana
HAProxy
Hyper-V
Jenkins
Kubernetes
Nginx
OpenStack
Prometheus
Proxmox VE
Terraform
Ubuntu
VMWare
Yandex Cloud
Zabbix
GitHub
GitLab
Apply
In office • Minsk
Bash
Python
Databases
Cassandra
PostgreSQL
DevOps
Ansible
CI/CD
Docker
GitHub Actions
GitLab CI
Helm
Kubernetes
Nginx
TeamCity
Terraform
GitHub
GitLab
Apply
up to $60k per year (gross) • Remote • Full-Time • 5+ years exp • Almaty
SQL
Databases
MariaDB
MySQL
PostgreSQL
DevOps
Ansible
Chef
Grafana
Prometheus
Puppet
Zabbix
Apply
Data & AI Engineer 2 days ago
In office • Full-Time • 3+ years exp • Luxembourg City
Java
Python
TypeScript
C#
Java
Spring Boot
Python
FastAPI
C#
.NET
AI/ML
Computer Vision
LLM
Mistral
NLP
Ollama
ONNX
RAG
Semantic Search
Transformers
Edge AI
OCR
Semantic Search
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GitHub Actions
Rest API
Terraform
Vector
GitHub
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.