414,183open jobs
14,210companies
59,991added this week
Browse all
Salary
$93k – $207k per year (Estimated)
Location
Remote (Germany)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated infrastructure reliably in production.

Your role involves guiding partners beyond the initial cluster deployment and validation phase into advanced Day 2 operations. These operations cover ongoing infrastructure health, observability, lifecycle management, quick remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and the broader external customer environments of the partners. This is a highly technical, hands-on role at the intersection of NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations.

What you'll be doing:

  • Lead NCP Day 2 operational readiness efforts. Collaborate directly with NVIDIA Cloud Partners to set up the systems, procedures, automation, and operational methods necessary to consistently manage NVIDIA accelerated infrastructure following initial deployment and activation.

  • Build continuous infrastructure validation. Develop and implement methods to continuously validate GPU, CPU, storage, and network health. Do this across large-scale AI clusters to identify degraded infrastructure before it impacts critical training or inference workloads.

  • Establish observability and operational telemetry. Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.

  • Develop automated detection and remediation. Build workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service while minimizing disruption to customer workloads.

  • Refine fleet lifecycle administration. Implement scalable strategies for managing sizable GPU fleets, including NVIDIA driver and firmware lifecycle administration, Kubernetes node maintenance, OS patching, configuration management, upgrades, and configuration drift identification.

  • Operationalize NVIDIA reference architectures. Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable operational standards.

  • Define operational health and readiness. Develop health signals, SLOs, important metrics, acceptance criteria, and ongoing validation mechanisms that provide NVIDIA and NCPs with clear insight into infrastructure reliability and service readiness.

  • Build reusable operational frameworks. Develop tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations that can be applied consistently across multiple NCP environments.

What we need to see:

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.

  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.

  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.

  • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.

  • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements.

  • Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.

  • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers.

  • Programming and automation experience using Python, Go, shell scripting, or similar languages.

Ways to stand out from the crowd:

  • Experience managing extensive GPU or accelerated computing infrastructure that supports AI training and inference workloads.

  • Experience with NVIDIA technologies including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.

  • Proven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency.

  • Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines and translating reference architectures or infrastructure requirements into repeatable production operating models across multiple customer or partner environments.

  • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features necessary to consistently support extended training and production inference.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and dedicated people in the world working for us.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
414,183 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Germany
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
Python
Java
DevOps
GCP
Istio
OpenTelemetry
Consul
Linkerd
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Chaos Engineering
Service Mesh
Apply
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
Python
PowerShell
Databases
DynamoDB
Amazon Redshift
Amazon Aurora
AI/ML
Ray
DevOps
Terraform
GCP
GitHub Actions
Istio
OpenTelemetry
AWS CDK
CloudFormation
Linkerd
Prometheus
GitLab CI
Azure
CI/CD
GitOps
Jenkins
AWS
Kubernetes
Grafana
Service Mesh
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
eBPF
FinOps
GitHub
GitLab
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Amazon Kinesis
API Gateway
Cybersecurity
ISO 27001
PCI DSS
SOC 2
HIPAA
FedRAMP
Zero Trust
Least Privilege
Apply
$145k – $205k per year • Remote • 6+ years exp • PhD
Rust
SQL
C++
Rust
Actix Web
Axum
Databases
MySQL
PostgreSQL
Redis
NATS
DynamoDB
RabbitMQ
Apache Kafka
Kafka
DevOps
Rest API
gRPC
Terraform
GCP
Istio
OpenTelemetry
Linkerd
Prometheus
Pulumi
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Service Mesh
Apply
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
Python
Java
DevOps
Splunk
Loki
New Relic
OpenTelemetry
Datadog
Prometheus
CI/CD
Grafana
Platform Engineering
Thanos
Mimir
eBPF
Cortex
Incident Management
Apply
In office • 5+ years exp
Python
DevOps
Zabbix
VMWare
Prometheus
Docker
Grafana
Platform Engineering
Cybersecurity
Wireshark
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
LangChain
LlamaIndex
Function Calling
AI Agents
LLM
OpenAI
Agentic Workflows
Tool Use
DevOps
CI/CD
Git
Apply
$14k – $23k per year (Estimated) • In office • Internship • PhD • Shanghai
Python
C
C++
C
MPI
AI/ML
CUDA Toolkit
OpenCL
CUDA
DevOps
HPC
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
C#
C++
AI/ML
Model Context Protocol
AI Agents
Apply
In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
AI Agents
Apply
$13k – $21k per year (Estimated) • In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$88k – $212k per year (Estimated) • Remote • Full-Time • Bachelor's Degree • Nanterre
Marketing
Salesforce
Apply
$86k – $130k per year • Remote • Full-Time • 10+ years exp • Bachelor's Degree • Madrid • Milan • Lisbon • Paris
Apply
$17k – $94k per year (Estimated) • In office • Internship • Germany
MATLAB
MATLAB
Simulink
Apply
$59k – $140k per year (Estimated) • In office • Full-Time • 3+ years exp • Germany
Apply
$77k – $149k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Munich • Hamburg
Python
SQL
Databases
Databricks
Delta Lake
AI/ML
Spark
dbt
DevOps
Azure
CI/CD
AWS
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
414,183 more open roles from verified company boards, updated every day.