368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$26k – $72k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Middle · 3+ years exp
Overview
Company
Impact
Profile match
ZenteiQ is a deep technology company headquartered in Bengaluru, India, and founded in 2022 out of research at the Indian Institute of Science. The company works on scientific machine learning, combining numerical simulation with neural networks for engineering and industrial modelling problems, and builds training programmes around those methods. It sells to industrial and research customers in India while running education initiatives that push scientific computing skills into the wider engineering workforce.

We are looking for a Distributed ML Infrastructure Engineer to build and optimise the large-scale distributed training systems behind our foundation models. You will own the performance, scalability and reliability of training infrastructure spanning multi-node GPU and TPU clusters. This is a hands-on systems role at the core of how BrahmAI gets trained, working closely with our research teams to turn raw compute into efficient, dependable training throughput.

Responsibilities:

  • Build and maintain distributed training infrastructure for large-scale AI workloads.
  • Optimise training performance through profiling, memory, communication, and compute optimisation.
  • Implement distributed training strategies including data, tensor, pipeline and sequence parallelism.
  • Improve training throughput, scalability and reliability across multi-node clusters.
  • Develop automation, monitoring, profiling and debugging tools for production training systems.
  • Collaborate with research and engineering teams to deliver efficient and scalable training infrastructure.

Requirements:

  • 3+ years of experience in Distributed Systems, HPC or ML Infrastructure.
  • Strong proficiency in Python, Linux and shell scripting.
  • Experience with multi-node GPU/TPU clusters, CUDA, PyTorch Distributed, NCCL, MPI/OpenMPI and distributed communication.
  • Strong understanding of GPU/TPU architecture, distributed computing, profiling and performance optimisation.
  • Experience with Docker, Kubernetes, Git and containerised development.
  • Strong debugging, profiling and performance analysis skills.

Good to Have / Bonus Points:

  • JAX, XLA, MaxText, XPK, PJRT and TPU training infrastructure.
  • Triton, Pallas or custom kernel optimisation.
  • Slurm, Ray, InfiniBand/RDMA and distributed storage systems.
  • Experience with TensorBoard Profiler, Nsight Systems/Compute or XProf.
  • Experience supporting large-scale LLM pretraining and distributed training at scale.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
Platform Engineer 1 day ago
$87k – $139k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
$123k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas • Denver • Birmingham
Java
SQL
Java
Gradle
Hibernate
Maven
Spring Boot
Spring Framework
Databases
Apache Kafka
MySQL
Redis
DevOps
CI/CD
Dynatrace
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Apply
$58k – $171k per year (Estimated) • In office • Full-Time • 1+ year exp • Munich
Python
AI/ML
Computer Vision
ONNX
PyTorch
DevOps
Docker
SLURM
Apply
Staff Engineer 1 day ago
$105k – $232k per year • In office • Full-Time • 8+ years exp • Munich
Python
TypeScript
JavaScript
AI/ML
LLM
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Terraform
Apply
$22k – $64k per year (Estimated) • In office • 2+ years exp • Bengaluru
Python
Rust
Python
Beautiful Soup
Dask
AI/ML
LLM
Ray
Spark
Synthetic Data
Hugging Face
NVIDIA NeMo
DevOps
GCP
Analytics
ETL/ELT
Apply
$15k – $54k per year (Estimated) • In office • 3+ years exp • Bengaluru
C++
Dart
Kotlin
Objective-C
Swift
JavaScript
AI/ML
Edge AI
Frontend
React.js
Mobile
Firebase
Flutter
Kotlin Multiplatform
Offline-First
React Native
Apply
$30k – $79k per year (Estimated) • In office • 6+ years exp • Bengaluru
JavaScript
Python
SQL
TypeScript
Python
FastAPI
Pydantic
Databases
Apache Kafka
PostgreSQL
Redis
AI/ML
Embeddings
RAG
AI Agents
Frontend
Next.js
React.js
DevOps
CI/CD
Docker
GCP
Git
Google GKE
Kubernetes
Prometheus
Rest API
Cybersecurity
SonarQube
QA
Jest
Pytest
Swagger
Apply
ML / AI Engineer 1 month ago
$31k – $128k per year (Estimated) • In office • Bengaluru
Python
Python
pySpark
AI/ML
Computer Vision
CUDA Toolkit
DeepSpeed
Fine-tuning
NLP
PyTorch
RLHF
Spark
DPO
FSDP
TPU
DevOps
Platform Engineering
HPC
Apply
AI / ML Engineer 1 month ago
$31k – $128k per year (Estimated) • In office • Bengaluru
Python
AI/ML
Fine-tuning
NLP
PyTorch
RLHF
TPU
DevOps
Platform Engineering
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 1 hour ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
$26k – $69k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Bengaluru • Hyderabad • Chennai • Noida
Databases
Db2
IMS
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.