Location
In office (Bengaluru)
Overview
Company
Impact
Profile match
ZenteiQ is a deep technology company headquartered in Bengaluru, India, and founded in 2022 out of research at the Indian Institute of Science. The company works on scientific machine learning, combining numerical simulation with neural networks for engineering and industrial modelling problems, and builds training programmes around those methods. It sells to industrial and research customers in India while running education initiatives that push scientific computing skills into the wider engineering workforce.
We are looking for a Performance Engineer, Inference, to understand and improve the systems that serve our foundation models. Inference is a tightly coupled system spanning model execution, serving runtimes, distributed systems, accelerators, scheduling, memory, and reliability. You will measure the system end-to-end, identify the highest-leverage performance gaps, and work across teams to close them while preserving correctness.
Responsibilities:
- Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability, and cost.
- Build profiling, benchmarking, and observability tools that make inference performance measurable and explainable.
- Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems, and accelerators.
- Partner with model, platform, and infrastructure teams to prioritise and land high-impact optimisations.
- Validate that performance improvements preserve model quality and numerical correctness.
Requirements:
- Hands-on experience profiling and optimising ML systems or other performance-critical production systems.
- Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
- Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems.
- Strong Python skills and the ability to read, instrument, and modify large production codebases.
- Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs, and accelerator performance fundamentals.
Preferred Qualifications:
- Experience with large-scale or multi-node inference, including tensor, pipeline, data, or expert parallelism.
- Experience with GPUs, TPUs, NPUs, or other ML accelerators and associated profiling tools.
- Experience with quantisation, low-precision inference, KV-cache optimisation, speculative decoding, or long-context serving.
- Experience contributing to or modifying inference runtimes, kernels, compilers, or distributed serving components.
- Experience optimising inference for constrained or on-device environments.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
431,915 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Bengaluru
Associate Analyst
1 hour ago
≈ $13k – $29k per year (Estimated) • In office • 1+ year exp • Chennai
Python
SQL
Scala
Databases
Delta Lake
AI/ML
Spark
Analytics
Power BI
Apply
Senior Software Engineer
1 hour ago
≈ $27k – $63k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
JavaScript
TypeScript
SQL
Node JS
Databases
Neo4j
Frontend
Angular
Bootstrap
React.js
Sass
DevOps
CI/CD
Docker
QA
Jest
Mocha
Apply
Sr. Data Scientist
1 hour ago
≈ $25k – $51k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
Python
pySpark
AI/ML
Hadoop
Spark
Scikit-learn
TensorFlow
Pandas
NumPy
PyTorch
Recommender Systems
Apply
≈ $14k – $34k per year (Estimated) • Equity • In office • Full-Time • Bachelor's Degree • Kraków
Python
Java
AI/ML
Copilot
Claude
Claude Code
DevOps
CI/CD
Apply
Python-разработчик
1 hour ago
≈ $18k – $48k per year (Estimated) • Remote • Full-Time • Novosibirsk
Python
Rust
Python
Django
Ruff
Databases
PostgreSQL
Redis
DevOps
Rest API
Ansible
Git
Cybersecurity
Keycloak
QA
Pytest
Apply
AI Data Engineer (LLM Training Data)
21 day ago
In office • 2+ years exp • Bengaluru
Python
Rust
Python
Beautiful Soup
Dask
AI/ML
Spark
LLM
Ray
Synthetic Data
Hugging Face
NVIDIA NeMo
DevOps
GCP
Analytics
ETL/ELT
Apply
Mobile Systems Engineer - Device AI
22 days ago
In office • 3+ years exp • Bengaluru
JavaScript
Kotlin
C++
Dart
Swift
Objective-C
AI/ML
Edge AI
Frontend
React.js
Mobile
Flutter
React Native
Firebase
Kotlin Multiplatform
Offline-First
Apply
Tech Lead - Full Stack
23 days ago
≈ $27k – $67k per year (Estimated) • In office • 6+ years exp • Bengaluru
Python
JavaScript
TypeScript
SQL
Python
FastAPI
Pydantic
Databases
PostgreSQL
Redis
Apache Kafka
AI/ML
Embeddings
AI Agents
RAG
Frontend
Next.js
React.js
DevOps
Rest API
GCP
Prometheus
CI/CD
Git
Docker
Kubernetes
Google GKE
Cybersecurity
SonarQube
QA
Swagger
Jest
Pytest
Apply
Distributed ML Infrastructure Engineer
24 days ago
≈ $23k – $62k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
JAX
PyTorch
LLM
TensorBoard
Ray
TPU
NCCL
InfiniBand
DevOps
SLURM
Git
Docker
Kubernetes
HPC
Apply
ML / AI Engineer
2 months ago
In office • Bengaluru
Python
Python
pySpark
AI/ML
Spark
DeepSpeed
CUDA Toolkit
Fine-tuning
RLHF
Computer Vision
NLP
PyTorch
DPO
FSDP
TPU
DevOps
Platform Engineering
HPC
Apply
≈ $23k – $50k per year (Estimated) • In office • Full-Time • PhD • Bengaluru
Apply
Computational Data Science Researcher
1 day ago
Remote/Hybrid • Full-Time • PhD • Bengaluru
AI/ML
Reinforcement Learning
Time Series Forecasting
Interpretability
Apply
Business Operations Analyst
1 day ago
≈ $17k – $38k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru • Gurgaon
Management
Outlook
Apply
Senior Governance Analyst
1 day ago
$91k – $151k per year • In office • Full-Time • 1+ year exp • Bachelor's Degree • Atlanta • Bengaluru • Hyderabad
Analytics
Tableau
Power BI
Microsoft Excel
Apply
Apply
This is one of many
431,915 more open roles from verified company boards, updated every day.

