368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$90k – $228k per year (Estimated)
Location
Remote (Germany)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You’ll collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated computing for AI.

What you’ll be doing:

  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation.

  • Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization; build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization.

  • Define and build inference benchmarking methodologies and tools; contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite.

  • Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.

  • Conduct and publish original research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products.

What we need to see:

  • Bachelor’s degree (or equivalent expeience) in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE) with 7+ years of experience; alternatively, Master’s degree in CS/CE/SE with 5+ years of experience; or PhD degree with the thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.

  • Strong programming skills in Python and C/C++; experience with Go or Rust is a plus; solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, distributed systems, deep learning theories.

  • Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).

  • Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).

  • Experience with containers and orchestration (Docker, Kubernetes, Slurm); familiarity with Linux namespaces and cgroups.

  • Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.

Ways to stand out from the crowd

  • Experience building and optimizing LLM inference engines (e.g., vLLM, SGLang).

  • Hands-on work with ML compilers and DSLs (e.g., Triton, TorchDynamo/Inductor, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).

  • Experience contributing to containerization/virtualization technologies such as containerd/CRI-O/CRIU.

  • Experience with cloud platforms (AWS/GCP/Azure), infrastructure as code, CI/CD, and production observability.

  • Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.

At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential. Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Germany
$42k – $107k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Mexico City
C#
Go
Java
Rust
AI/ML
AI Agents
LLM
Model Context Protocol
Cybersecurity
Delinea
Zero Trust
Apply
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
$70k – $105k per year • In office • Full-Time • 3+ years exp
Python
SQL
TypeScript
AI/ML
LLM
RAG
Function Calling
LLM Guardrails
Cybersecurity
GDPR
Management
n8n
Apply
Founding Engineer 1 day ago
$93k – $140k per year • In office • Full-Time • 3+ years exp • Munich
Python
Python
FastAPI
AI/ML
Fine-tuning
LLM
VLM
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$156k – $348k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United Kingdom
AI/ML
CUDA
CUDA Toolkit
AI Agents
NVIDIA NeMo
Apply
$62k – $127k per year (Estimated) • In office • Full-Time • Mainz • Mannheim • Dresden • Ludwigsburg • Berlin
Apply
$59k – $121k per year (Estimated) • In office • Full-Time • Ludwigsburg • Mainz • Dresden
AI/ML
Copilot
Apply
$64k – $143k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Germany
Apply
$31k per year • In office • Internship • Bachelor's Degree • Germany
Analytics
Power BI
Management
Power Automate
Apply
$71k – $142k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Germany
C#
C++
C++
Qt
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.