415,615open jobs
14,704companies
74,833added this week
Browse all
Salary
$115k – $304k per year (Estimated)
Location
In office (Israel, Tel Aviv)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is powering the world's most advanced AI Factories. To ensure their seamless operation, we are building a mission-critical Observability and Prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for our largest enterprise customers.

We are looking for a Senior Software Engineer to join the AIOps platform team and help build the core distributed systems that ingest massive telemetry streams from GPU clusters and operationalize predictive AI models at scale. You will work at the intersection of high-performance data engineering and production ML, turning research algorithms into reliable, mission-critical software.

What you'll be doing:

  • Architect and build an agentic AIOps system that autonomously monitors GPU fleet health, aggregates and correlates massive telemetry streams, surfaces intelligent alerts, and orchestrates multi-step diagnostic workflows and corrective actions - powering real-time dashboards, automated root-cause analysis, and proactive incident response.

  • Research, evaluate, and prototype data storage strategies and data representations across diverse database technologies and modalities, ensuring AI models are trained on high-quality, well-structured data that improves predictive accuracy and generalization.

  • High-Scale Engineering: Design distributed systems to handle the extreme telemetry density of large-scale AI clusters, ensuring efficient data ingestion, processing, and real-time analysis.

  • Instrument services with deep observability (metrics, logs, traces) to support rapid debugging and continuous performance improvement.

  • Build and own the model-serving infrastructure that operationalizes predictive algorithms at scale - packaging, versioning, deploying, and monitoring AI models in both SaaS and on-premises environments.

  • Contribute to the platform's core libraries and abstractions that accelerate development across the broader AIOps engineering team.

What we need to see:

  • B.Sc./M.Sc. in Computer Science, Computer Engineering, or a related technical field.

  • 8+ years of software engineering experience building production distributed systems.

  • Core Systems Programming: Expert-level proficiency in languages such as Go, C++, or Rust, with a focus on high-performance, concurrent architectures.

  • Solid understanding of Kubernetes and container-based deployments for production services.

  • Experience deploying, monitoring, and maintaining ML models or data-intensive services in a production environment.

  • Comfort working in ambiguous, fast-moving environments where the product is still being shaped.

Ways to stand out from the crowd:

  • Experience building ML model-serving platforms or MLOps tooling (model registries, A/B rollout frameworks, feature stores) at scale.

  • A track record of taking systems from prototype to stable, production-grade platform serving real enterprise customers.

  • A "Systems" Thinker: You don't just write software; you understand the full stack, from how data moves across the wire to how it’s processed in a distributed cluster.

  • Practical Innovation: The ability to simplify complex problems and build internal tools or frameworks that empower other engineering teams to move faster.

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are passionate about building mission-critical systems at the frontier of AI infrastructure, we want to hear from you.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
415,615 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Tel Aviv
$120k – $180k per year • Equity 0.2–1.5% • In office • Full-Time • 1+ year exp • San Francisco
Python
AI/ML
AI Agents
ChatGPT
Human-in-the-Loop
DevOps
Error Budget
Apply
$120k – $180k per year • Equity 0.2–1.5% • In office • Full-Time • 1+ year exp • San Francisco
Python
AI/ML
AI Agents
ChatGPT
DevOps
Error Budget
Apply
$55k – $75k per year • Equity 0–0.2% • In office • Full-Time • Albuquerque
Python
AI/ML
AI Agents
ChatGPT
Apply
$120k – $180k per year • Equity 0.2–1.5% • In office • Full-Time • 1+ year exp • San Francisco
Python
AI/ML
AI Agents
ChatGPT
DevOps
Error Budget
Apply
$275k – $325k per year • In office • Full-Time • 3+ years exp • San Francisco
AI/ML
AI Agents
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
$144k – $230k per year • In office • Full-Time • PhD • Santa Clara
Apex
Apex
MuleSoft
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
$24k – $58k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Shanghai
C++
Python
AI/ML
AI Agents
Copilot
Cursor
LLM
Reinforcement Learning
Robotics
Isaac Lab
Isaac Sim
Perception
Reinforcement Learning
ROS
Sim-to-Real
Teleoperation
Apply
$53k – $135k per year (Estimated) • In office • 3+ years exp • Tel Aviv
Apply
$55k – $121k per year (Estimated) • In office • Full-Time • 3+ years exp • Tel Aviv
DevOps
Ansible
Azure
Management
Jira
Slack
Apply
Game Engineer 1 day ago
$29k – $104k per year (Estimated) • In office • Full-Time • 4+ years exp • Tel Aviv
C#
Go
Node JS
JavaScript
DevOps
AWS
Azure
CI/CD
GCP
WebSockets
Game Dev
Unity
Analytics
A/B Testing
Apply
SecOps Engineer 1 day ago
Remote/Hybrid • Full-Time • Tel Aviv
Python
AI/ML
Anomaly Detection
DevOps
AWS
Azure
GCP
SLI/SLO/SLA
Splunk
Apply
$82k – $217k per year (Estimated) • In office • 5+ years exp • Tel Aviv
C#
C++
Go
Java
Kotlin
Scala
TypeScript
JavaScript
Databases
PostgreSQL
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
See all jobs
This is one of many
415,615 more open roles from verified company boards, updated every day.