Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Sep 4, 2026.
Distributed ML Infrastructure Engineer (GPU/TPU)
Full-Time (FTE) | Distributed ML Infrastructure Engineer (GPU/TPU) | Bangalore |
Experience: 3+ Years
About ZenteiQ
ZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence
Infrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national
systems. Rather than wrapping general-purpose language models, we train foundation models
(BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials
domains, and put them to work through industrial platforms (KogneX) and a talent OS for
engineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work
closely with IISc, ARTPARK and a national AI Hub Network - building the sovereign AI
infrastructure that India's engineering and industrial systems will run on.
About the Role
We are looking for a Distributed ML Infrastructure Engineer to build and optimize the large-scale
distributed training systems behind our foundation models. You will own the performance,
scalability and reliability of training infrastructure spanning multi-node GPU and TPU clusters.
This is a hands-on systems role at the core of how BrahmAI gets trained, working closely with
our research teams to turn raw compute into efficient, dependable training throughput.
What You'll Do
- Build and maintain distributed training infrastructure for large-scale AI workloads.
- Optimize training performance through profiling, memory, communication and compute
optimization.
- Implement distributed training strategies including data, tensor, pipeline and sequence
parallelism.
- Improve training throughput, scalability and reliability across multi-node clusters.
- Develop automation, monitoring, profiling and debugging tools for production training
systems.
- Collaborate with research and engineering teams to deliver efficient and scalable training
infrastructure.
What We're Looking For
- 3+ years of experience in Distributed Systems, HPC or ML Infrastructure.
- Strong proficiency in Python, Linux and shell scripting.
- Experience with multi-node GPU/TPU clusters, CUDA, PyTorch Distributed, NCCL,
MPI/OpenMPI and distributed communication.
- Strong understanding of GPU/TPU architecture, distributed computing, profiling and
performance optimization.
- Experience with Docker, Kubernetes, Git and containerized development.
- Strong debugging, profiling and performance analysis skills.
Good to Have / Bonus Points
- JAX, XLA, MaxText, XPK, PJRT and TPU training infrastructure.
- Triton, Pallas or custom kernel optimization.
- Slurm, Ray, InfiniBand/RDMA and distributed storage systems.
- Experience with TensorBoard Profiler, Nsight Systems/Compute or XProf.
- Experience supporting large-scale LLM pretraining and distributed training at scale.
Why ZenteiQ
- Build real physics-native foundation models from scratch, not another LLM wrapper -
deep, defensible technical work.
- Be part of a nationally recognised mission: one of 8 startups selected under the
government's IndiaAI Mission to build a sovereign foundation model.
- Work alongside IISc-trained scientists and researchers, in a company founded by an IISc
professor.
- See your work land in real industry pilots across automotive, mobility, defence and
industrial R&D, with measurable impact.
- Join a lean, high-caliber team at an early, high-ownership stage.

