Bengaluru
5+ year exp
Python
Rust
Bachelor's Degree
Apply now
Save
Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs
IDFC First Bank
PERSONAL BANKING. Maximise savings with 6.50% p.a. interest.
We are looking for an experienced AI Engineer to build and optimize high-performance LLM inference infrastructure for production-scale AI systems. This role focuses on inference performance, GPU optimization, Kubernetes-native infrastructure, and large-scale model serving rather than application-level GenAI development.
The ideal candidate has hands-on experience deploying and optimizing LLM inference workloads, understands GPU architecture and performance engineering, and has worked on production AI infrastructure. Candidates whose experience is limited to RAG pipelines, chatbots, LangChain, or OpenAI API integrations without systems-level expertise will not be considered.
Responsibilities:
- Design, build, and optimize production-grade LLM inference infrastructure.
- Deploy, configure, and operate LLM serving frameworks such as vLLM in production environments.
- Optimize inference latency, throughput, and GPU utilization across large-scale workloads.
- Develop Kubernetes-native infrastructure including Operators, Helm Charts, monitoring, and GPU observability.
- Analyze and improve inference performance using profiling and benchmarking tools.
- Design scalable serving architectures with batching, KV cache management, tensor parallelism, and pipeline parallelism.
- Work closely with platform, ML, and infrastructure teams to improve inference efficiency and production reliability.
- Build tooling for benchmarking, monitoring, debugging, and performance validation.
- Contribute to production-ready ML infrastructure, CI/CD pipelines, and cloud-native deployments.
Requirements:
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field (or equivalent experience).
- 5-8 years of experience building performance-critical software, ML infrastructure, distributed systems, or AI platforms.
- Strong programming skills in Python (Go or Rust is a plus).
- Excellent understanding of algorithms, data structures, operating systems, computer architecture, parallel programming, distributed systems, and deep learning fundamentals.
- Strong understanding of modern LLM inference architecture and serving concepts.
- LLM Inference: Hands-on experience with vLLM deployment and optimization, TTFT (Time to First Token), ITL (Inter-Token Latency), Prefill and Decode stages, KV Cache optimization, Speculative Decoding, Production benchmarking, and reproducible performance measurements.
- LLM Serving: Experience with: Continuous batching, Chunked prefill, KV cache management, Tensor parallelism, Pipeline parallelism, and Production inference optimization.
- GPU and Performance Engineering: Strong understanding of CUDA fundamentals, GPU memory hierarchy, CUDA Streams, NCCL, Tensor Cores, GPU utilization optimization, Roofline modeling, Nsight Systems, and Nsight Compute.
- Infrastructure and Platform: Hands-on experience with Kubernetes, Helm, GPU observability, Production monitoring, Distributed systems, Cloud platforms (AWS/GCP/Azure), CI/CD pipelines, Infrastructure as Code.
- ML Frameworks: Experience optimizing PyTorch, vLLM, SGLang, Production inference engines.
- Good to Have: NVIDIA Dynamo, llm-d, Triton, TorchDynamo, MLIR, LLVM, XLA, CUTLASS, CUDA Graphs, Tensor Core optimization, Building custom inference engines, Production-scale AI infrastructure.
Preferred Skills:
- Strong preference will be given to candidates from companies such as NVIDIA, Microsoft, Qualcomm, and Simplismart AI (closest fit).
- Other organizations building large-scale AI infrastructure, GPU systems, or LLM inference platforms.
Recommended for you based on this role
Similar stack
Same company
In your city
Backend Engineer - Gen AI
13 days ago
Mumbai
Go
JavaScript
Node JS
Python
Python
FastAPI
Frontend
React.js
Apply
Senior Data Scientist (Gen AI)
16 days ago
Bengaluru
Python
SQL
AI/ML
BERT
Fine-tuning
LLM
NLP
Prompt Engineering
PyTorch
TensorFlow
Tokenization
Transformers
Analytics
Matplotlib
Seaborn
Tableau
Apply
Sr. Data Scientist (Chatbots)
16 days ago
Bengaluru
Python
SQL
AI/ML
BERT
Keras
LangGraph
NLP
NLTK
NumPy
PyTorch
Scikit-learn
Sentiment Analysis
spaCy
Speech Recognition
TensorFlow
LangChain
Analytics
Matplotlib
Seaborn
Tableau
Apply
Lead Data Scientist (Gen AI)
20 days ago
Bengaluru
Python
SQL
AI/ML
BERT
Gensim
Keras
NLP
NLTK
NumPy
PyTorch
Scikit-learn
Sentiment Analysis
spaCy
TensorFlow
DevOps
AWS
Azure
GCP
Apply
Senior Data Scientist
20 days ago
3+ year exp • Bengaluru
Python
SQL
AI/ML
Computer Vision
Fine-tuning
Hallucination
LangChain
LlamaIndex
LLM
Multimodal AI
NLP
Prompt Engineering
PyTorch
RAG
Speech Recognition
TensorFlow
DevOps
AWS
Azure
GCP
Apply
Data Engineer
20 days ago
2+ year exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Python
pySpark
Databases
Apache Kafka
AI/ML
Airflow
Spark
DevOps
Amazon EC2
AWS
Azure
Bitbucket
GCP
Management
Jira
Apply
Apply
Backend Developer
23 days ago
3+ year exp • Chennai
Go
Java
SQL
Java
Spring Boot
Databases
MySQL
PostgreSQL
Frontend
GraphQL
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Jaeger
Jenkins
Kubernetes
Apply
Data Scientist (Fraud & Ops Analytics)
26 days ago
3+ year exp • Master's Degree • Mumbai
Python
SQL
AI/ML
Anomaly Detection
NumPy
Pandas
PyTorch
Scikit-learn
TensorFlow
Apply
AI Engineer
27 days ago
Bengaluru
Python
AI/ML
AI Agents
Computer Vision
Fine-tuning
Hallucination
LLM
LoRA
Multimodal AI
NLP
PEFT
QLoRA
RAG
Spark
Unsloth
Transformers
DevOps
AWS
Azure
CI/CD
Docker
Kubernetes
Vector
Cybersecurity
GDPR
PCI DSS
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Free forever • Less than a minute • No credit card
Work setup
Location
Bengaluru
