703,472open jobs
41,487companies
102,442added this week
Browse all
Salary
$140k – $224k per year
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA’s DGX Cloud organization is seeking a Senior Data Engineer to become part of its data team! We develop the reliable data foundation that supports fleet health, capacity, utilization, cost, reliability, and operational decision-making throughout DGX Cloud. Our platform supports engineering, operations, finance, and product teams managing and expanding large GPU fleets across cloud service providers and NVIDIA Cloud Partners. We are looking for a practical engineer and technical lead to take charge of a key part of the Navigator data platform. We develop the systems that transform distributed infrastructure telemetry and operational data into dependable, managed data products that support fleet health, capacity, utilization, cost, and operational decisions.

We are seeking a hands-on, platform-minded engineer to build and evolve the systems that turn distributed infrastructure telemetry and operational data into reliable, governed data products. You will work across ingestion, transformation, data quality, platform architecture, security, observability, and self-service consumption to help make Navigator and the DGXC data platform a dependable source of truth. We do expect strong engineering fundamentals, experience operating production systems, and the ability to learn new platforms and domains quickly.

What you'll be doing:

  • Own systems end to end. For example, work from ambiguous customer and operational needs through architecture, implementation, deployment, observability, incident response, and ongoing support.

  • Construct data pipelines and products. Such as designing and maintain batch and streaming ingestion, transformation, reconciliation, and serving paths for fleet, capacity, utilization, cost, scheduling, and operational telemetry.

  • Build shared libraries, workflow and DAG or equivalent experience abstractions to evolve the data platform. Develop deployment tooling, data contracts, and paved-road patterns that improve team speed and safety.

  • Engineer reliable distributed workloads. As well as diagnose correctness and performance issues across applications, SQL engines, Spark jobs, storage systems, networks, and cloud services. Build for retries, idempotency, backfills, schema evolution, and partial failure.

  • Treat security as part of the build. For example, applying least privilege, service identities, secrets management, access controls, environment isolation, auditability, and safe operational practices throughout the system lifecycle.

  • Improve quality and operations: Establish automated tests, data-quality checks, lineage, freshness and completeness monitoring, actionable alerting, SLOs, and clear ownership.

  • Deliver consumption experiences. Such as making trusted data usable through well-modeled tables, APIs, automation, dashboards, and focused internal applications-not only through one-off queries.

  • Raise the engineering bar. Lead build reviews, communicate tradeoffs, mentor other engineers, and improve the team's architecture, testing, debugging, and operational practices.

What we need to see:

  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.

  • 5+ years of experience building and operating production software, data platforms, backend infrastructure, databases, or distributed systems.

  • Strong software-engineering fundamentals and production proficiency in Python or another backend or systems language, with the ability and willingness to work primarily in Python and SQL.

  • Deep hands-on experience in at least one of the following areas: Distributed data processing using Spark or a comparable compute framework, Relational, distributed, or analytical database architecture and operation at scale, Production ETL, change-data-capture, streaming, or event-processing systems, Backend or cloud-platform systems that process, transform, or serve substantial data volumes, Strong SQL and data-modeling skills, including a practical understanding of query performance, schema evolution, incremental processing, consistency, and analytical consumption patterns.

  • Demonstrated ability to debug unfamiliar systems across multiple layers using logs, metrics, traces, query plans, profiles, and controlled experiments to find root causes.

  • Experience operating services or pipelines in a cloud or similarly complex production environment, including testing, CI/CD, monitoring, alerting, rollback, and incident response.

  • Working knowledge of secure platform development, including identity and access management, least privilege, secret handling, trust boundaries, and safe multi-environment deployments.

  • Ability to make sound architectural tradeoffs, own work through ambiguity, and communicate effectively with users, partner teams, and engineers from different fields.

  • A track record of learning unfamiliar technologies and domains and turning that learning into maintainable systems and reusable team practices.

  • Experience with AI agents and LLM-supported workflow automation, particularly as applied to engineering and operational activities.

Ways to stand out from the crowd:

  • Experience with Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, or another modern lakehouse or distributed-compute platform.

  • Experience with Kafka or another streaming platform, change-data capture, event development, partitioning, consumer groups, offset management, or other high-volume event systems.

  • Experience with scaling, migrating, or performance-tuning relational, distributed, time-series, object-storage, or search-focused data systems, including Elasticsearch or OpenSearch.

  • Background working with AWS, Azure, GCP, Kubernetes, Slurm, compute clusters, GPU-accelerated infrastructure, or fleet-scale telemetry.

  • Experience developing agentic systems, LLM-enabled workflow automation, harness engineering, or dependable evaluation and operational tooling for AI agents.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 140,000 USD - 224,250 USD for Level 3, and 168,000 USD - 270,250 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 25, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
703,472 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
$77k – $174k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Canada
Python
Java
Bash
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
MS SQL
AI/ML
Copilot
Claude Code
AI Agents
Edge AI
DevOps
Ansible
GCP
GitHub Actions
Istio
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Grafana
Configuration Management
Linux
QA
Selenium
Apply
Data Analyst 2 hours ago
$90k – $94k per year • In office • Contractor • Bachelor's Degree
Python
SQL
Analytics
Power BI
Microsoft Excel
Management
Agile
Apply
$118k – $147k per year • Remote/Hybrid • Full-Time • 4+ years exp • Boston
Python
JavaScript
SQL
AI/ML
Copilot
Cursor
Claude Code
AI Agents
DevOps
CI/CD
Jenkins
AWS
Platform Engineering
GitHub
Management
Confluence
Jira
Agile
Apply
$116k – $184k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
Qwen
DeepSeek
vLLM
Gemma
Fine-tuning
SGLang
Ollama
Llama
PyTorch
OpenAI
Hugging Face
Apply
$135k – $339k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
Python
C++
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
RAG
Semantic Search
Semantic Search
Agentic Workflows
Tool Use
DevOps
Rest API
Apply
$116k – $184k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
Qwen
DeepSeek
vLLM
Gemma
Fine-tuning
SGLang
Ollama
Llama
PyTorch
OpenAI
Hugging Face
Apply
$216k – $345k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
DevOps
AWS
Marketing
Salesforce
Apply
$135k – $339k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
Python
C++
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
RAG
Semantic Search
Semantic Search
Agentic Workflows
Tool Use
DevOps
Rest API
Apply
$184k – $288k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara • Austin • Durham • Seattle
Verilog
C++
VHDL
AI/ML
CUDA Toolkit
CUDA
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
AI/ML
CUDA Toolkit
Triton Inference Server
Fine-tuning
RLHF
Quantization
JAX
Multimodal AI
Knowledge Distillation
TensorRT
TensorRT-LLM
Transformers
PyTorch
Synthetic Data
CUDA
Post-training
Megatron-LM
NVIDIA NeMo
NCCL
InfiniBand
NVLink
Speculative Decoding
Model Distillation
RLAIF
Machine Learning
Apply
In office • Full-Time • High School Diploma • United States
Apply
$41k – $71k per year (Estimated) • In office • Internship • Bachelor's Degree • United States
Apply
$41k – $71k per year (Estimated) • In office • Internship • Bachelor's Degree • United States
Apply
$37k – $63k per year (Estimated) • Equity • Remote • Full-Time • 2+ years exp • Bachelor's Degree • United States
Management
Agile
Apply
$79k – $152k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • United States
Apply
See all jobs
This is one of many
703,472 more open roles from verified company boards, updated every day.