Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Sep 4, 2026.
About ZenteiQ
ZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence
Infrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national
systems. Rather than wrapping general-purpose language models, we train foundation models
(BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials
domains, and put them to work through industrial platforms (KogneX) and a talent OS for
engineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work
closely with IISc, ARTPARK and a national AI Hub Network - building the sovereign AI
infrastructure that India's engineering and industrial systems will run on.
About the Role
You will build the systems that let our researchers train, evaluate, and serve large foundation
models reliably at scale. This role sits at the intersection of model research and infrastructure,
with a focus on accelerator-based (TPU) training and inference, performance, reproducibility,
and researcher velocity. You will own the paved path that turns expensive, long-running model
runs into a repeatable, observable, and cost-efficient process.
What You'll Do
- Build and improve the paved path for distributed training, evaluation, experiment tracking,
checkpoint management, model release, and inference on Cloud TPUs.
- Operate long-running ML workloads with strong observability, failure detection, automated
recovery, and practical operational tooling.
- Profile and remove bottlenecks across compute, HBM and host memory, data input,
networking and collectives, XLA compilation, checkpointing, and serving.
- Build validation, representative-scale testing, CI/CD, reproducibility, and lineage systems
that catch problems before expensive model runs or production releases.
- Create reusable APIs, abstractions, and self-service tooling that help researchers move
quickly while preserving useful low-level controls.
- Own accelerator capacity workflows, including TPU provisioning, quotas, reservations,
priorities, topology-aware placement, utilization, and cost efficiency.
- Partner with research teams to debug model and systems failures, translate recurring
pain points into durable platform improvements and define operational standards.
What We're Looking For
- Strong software engineering skills and experience owning systems used by researchers
or engineers in production or research-critical environments.
- Hands-on experience with ML training or inference infrastructure at meaningful scale,
including large accelerator or distributed compute workloads.
- Strong distributed-systems fundamentals and experience with Kubernetes, GKE, or
comparable orchestration for long-running compute jobs.
- Ability to debug across model code, data pipelines, runtimes, cluster services, storage,
networking, and accelerator behaviour.
- Strong observability, reliability, and performance-engineering fundamentals, including
incident response and root-cause analysis.
- Production-quality Python and the ability to build durable platform abstractions rather than
one-off scripts.
- High ownership, pragmatic judgment, and clear collaboration with fast-moving research
teams.
Good to Have / Bonus Points
- Experience operating large Cloud TPU clusters, including TPU VMs or Pods, multislice
jobs, topology, quotas, scheduling, and failure recovery.
- Experience with JAX, PyTorch/XLA, TensorFlow, XLA/HLO, PJRT, MaxText, Pallas, or
another large-scale training stack.
- Experience serving LLMs on TPUs, including batching, partitioning, compilation caching,
autoscaling, latency, and throughput optimization.
- Familiarity with GCP/GKE, Terraform, Vertex AI, XPK, Ray, Argo, Airflow, MLflow, Weights
& Biases, or comparable internal platforms.
- Experience with model registries, distributed checkpointing, data and artifact lineage, and
large artifact distribution.
- Proficiency in a systems language such as Go, Rust, C++, or Java.
Why ZenteiQ
- Build real physics-native foundation models from scratch, not another LLM wrapper -
deep, defensible technical work.
- Be part of a nationally recognised mission: one of 8 startups selected under the
government's IndiaAI Mission to build a sovereign foundation model.
- Work alongside IISc-trained scientists and researchers, in a company founded by an IISc
professor.
- See your work land in real industry pilots across automotive, mobility, defence and
industrial R&D, with measurable impact.
- Join a lean, high-caliber team at an early, high-ownership stage.

