368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$31k – $78k per year (Estimated)
Location
In office (Pune)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has continuously reinvented itself for more than two decades. The invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. GPU-powered deep learning helped ignite the era of modern AI,

establishing

GPUs as the foundation of intelligent applications across productivity, gaming, and creative workflows, and reinforcing NVIDIA’s position as a leading AI computing company.

More recently, there is a growing focus on running AI models locally, closer to where data is generated. This approach reduces latency, enables real-time processing, and addresses privacy concerns by minimizing the need to send data to centralized servers. As technology continues to evolve, client-side AI will play an increasingly

important role

in shaping the digital landscape.

The

LocalAI

team is seeking a Senior Systems Software Engineer to develop efficient on-device AI software for RTX and DGX-class systems. The role focuses on delivering high-performance local inference with low latency, optimized memory

utilization

, robust infrastructure, and practical deployment on resource-constrained platforms.

What

You’ll

Be Doing:

  • Partner with NVIDIA’s software, research, architecture, and product teams to align technical requirements and strategic priorities, fostering the AI ecosystem on RTX and DGX PCs.

  • Build and

    optimize

    the local AI inference stack for RTX, RTX Pro, and DGX GPUs, with a focus on performance, stability, and scalability across diverse hardware architectures.

  • Design and develop modern inference runtimes and execution stacks using frameworks such as llama.cpp,

    vLLM

    ,

    PyTorch

    ,

    WinML

    , DXCGC, and TensorRT-RTX, supporting LLM, vision-language, TTS, ASR, and diffusion-based AI workloads.

  • Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures. Apply model optimization techniques, including quantization, pruning, sparsity, and distillation, to enable efficient deployment of large models on local and edge devices.

  • Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; develop infrastructure for performance and accuracy sweeps;

    analyse

    results to

    identify

    gaps and drive fixes; and

    establish

    engineering guidelines to accelerate bring-up and ensure production readiness of new models and inference backends.

What we need to

see:

  • 5+ Years of experience with Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience.

  • Excellent C++ programming and debugging skills, with

    a strong foundation

    in data structures, algorithms, and machine learning.

  • Proven experience developing and

    optimizing

    AI inference pipelines and applications using ML/DL frameworks such as Llama.cpp,

    vLLM

    ,

    PyTorch

    ,

    Windows ML

    , DXCGC, and TensorRT.

  • Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache

    behaviour

    , graph execution, quantization, and hardware-aware optimization.

  • Strong analytical and problem-solving skills, with the ability to manage multiple priorities effectively in a fast-paced environment.

  • Excellent written and verbal communication skills, enabling effective collaboration across engineering teams and management.

Ways to stand out from the crowd:

  • Understanding of modern machine learning, deep neural network, and generative AI techniques, with relevant contributions to major open-source projects.

  • Consistent

    track record

    of delivering end-to-end products in multinational companies with geographically distributed teams.

  • Proficiency

    in low-level system and GPU programming, CUDA, and the development of high-performance systems.

  • Contributions to open-source inference runtimes, model tooling, or performance infrastructure.

  • Hands-on experience building applications using frameworks and APIs such as llama.cpp,

    PyTorch

    , TensorRT, Vulkan, DirectX, and

    vLLM

    .

We're

a top employer recognized for innovation, growth, and a commitment to diversity as an equal-opportunity workplace. We offer competitive salaries, a generous benefits package, and the opportunity to work alongside some of the technology industry's most talented and forward-thinking professionals. As our engineering teams continue to grow rapidly,

we're

looking for creative, self-driven engineers with a passion for technology to join us.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Pune
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
$20k – $48k per year (Estimated) • In office • Moscow
C#
C++
Java
Python
SQL
DevOps
CI/CD
Management
Draw.io
Jira
Apply
$21k – $55k per year (Estimated) • In office • Full-Time • Bengaluru
C++
Go
Java
C++
Protobuf
Databases
Apache Ignite
ElasticSearch
PostgreSQL
RabbitMQ
Redis
DevOps
CI/CD
Docker
gRPC
Kibana
OpenTelemetry
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tel Aviv
C++
DevOps
Platform Engineering
RTOS
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
C++
Java
Python
SQL
C#
TypeScript
JavaScript
Java
Maven
C#
.NET
Databases
Apache Kafka
MySQL
AI/ML
ChatGPT
Copilot
Frontend
Angular
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Prometheus
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$156k – $348k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United Kingdom
AI/ML
CUDA
CUDA Toolkit
AI Agents
NVIDIA NeMo
Apply
$13k – $29k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Pune
JavaScript
Apex
Apex
MuleSoft
AI/ML
AI Agents
Edge AI
DevOps
AWS
Azure
Management
Draw.io
Marketing
Salesforce
Apply
$11k – $42k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Pune
ABAP
Apply
Data Architect 2 hours ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$23k – $62k per year (Estimated) • In office • Full-Time • 3+ years exp • Navi Mumbai • Pune
Python
Apply
$27k – $71k per year (Estimated) • In office • Full-Time • 7+ years exp • Pune
C#
C++
Java
Python
DevOps
Azure
CI/CD
Git
QA
Pytest
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.