368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$105k – $265k per year (Estimated)
Location
In office (Tel Aviv, Yokneam)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is looking for an experienced HPC DevOps Engineer to help us build the supercomputers and HPC clusters of the future. As a Senior HPC DevOps Engineer, you'll be a key player in groundbreaking advancements in artificial intelligence and GPU computing. Your expertise will drive the latest breakthroughs, providing insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.

What you’ll be doing:

  • Innovate and Implement: Design, implement, and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging, and alerting systems.

  • Infrastructure as Code (IaC): Utilize and develop tools to manage infrastructure as code, ensuring scalable and repeatable deployments.

  • Streamline CI/CD Pipelines: Develop and maintain continuous integration and continuous delivery (CI/CD) pipelines to automate and streamline deployment processes.

  • Automate Everything: Develop automation scripts and tools to automate deployment, configuration management, and operational monitoring.

  • Develop complex Networking automations.

  • Troubleshoot Complex Issues: Perform comprehensive troubleshooting from bare metal to application level, ensuring system reliability and efficiency.

  • Lead and Educate: Serve as a technical resource, developing and sharing best practices with internal teams.

  • Drive Innovation: Support R&D activities and engage in proof of concepts (POCs) and proof of values (POVs) for future improvements.

What we need to see:

  • B.Sc. in Computer Science, Engineering, or a related field

  • 5+ years of experience

  • Advanced proficiency in programming and scripting languages, with a solid understanding of object-oriented programming principles.

  • Familiarity with Jenkins, Ansible, Puppet/Chef.

  • Deep understanding of Kubernetes and container-related microservice technologies.

  • Hands-on experience with event streaming or message queue technologies e.g. Apache Kafka.

  • Experience with multiple storage solutions like Lustre, GPFS, ZFS, and XFS. Expertise with virtual systems (VMware, Hyper-V, KVM, Citrix).

  • Familiarity with cloud platforms (AWS, Azure, Google Cloud).

Ways to stand out from the crowd:

  • Proven networking experience or strong knowledge through professional networking training.

  • Architectural Insight: Knowledge of CPU and/or GPU architecture.

  • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes.

At NVIDIA, we value diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We provide reasonable accommodations to ensure all individuals can participate in the job application or interview process, perform essential job functions, and receive other benefits and privileges of employment. Join us and be part of a team that's pushing the boundaries of technology and making a real impact in the world.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Tel Aviv
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$71k – $170k per year (Estimated) • In office • Full-Time • Netanya
Python
TypeScript
AI/ML
Accelerate
Fine-tuning
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Edge AI
Hugging Face
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$24k – $55k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Moscow
DevOps
AWS
Azure
SLI/SLO/SLA
Cybersecurity
GDPR
Management
Jira
ServiceNow
Apply
$55k – $157k per year (Estimated) • Remote • Full-Time • Sydney
C++
Go
Lua
Python
C++
CMake
Databases
ActiveMQ
Aerospike
Apache Kafka
Cassandra
DevOps
Docker
gRPC
Kubernetes
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$156k – $348k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United Kingdom
AI/ML
CUDA
CUDA Toolkit
AI Agents
NVIDIA NeMo
Apply
Product Manager 8 hours ago
$91k – $210k per year (Estimated) • In office • 4+ years exp • Tel Aviv
AI/ML
LLM
Apply
$68k – $224k per year (Estimated) • Remote/Hybrid • 7+ years exp • Bachelor's Degree • Tel Aviv
JavaScript
Frontend
React.js
Apply
$159k – $348k per year (Estimated) • In office • Tel Aviv
AI/ML
AI Agents
Human-in-the-Loop
LLM Guardrails
Model Context Protocol
RAG
DevOps
CI/CD
Kubernetes
Cybersecurity
Least Privilege
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tel Aviv
C++
DevOps
Platform Engineering
RTOS
Apply
Equity • Remote/Hybrid • Part-Time • Bachelor's Degree • Tel Aviv
Python
SQL
AI/ML
AI Agents
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.