{"id":1121183,"url":"https://alion.io/job/zenteiq-ml-systems-engineer-training-inference","title":"ML Systems Engineer – Training & Inference","company":{"id":1338,"name":"ZenteiQ","domain":"zenteiq.ai","url":"https://alion.io/company/zenteiq-ai","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":21000,"max_usd":58000,"period":"year","method":"role_country_seniority_unknown","sample_n":136},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"CI/CD","optional":false},{"name":"Google GKE","optional":false},{"name":"Kubernetes","optional":false},{"name":"Python","optional":false},{"name":"TPU","optional":false},{"name":"XLA","optional":false},{"name":"C++","optional":true},{"name":"GCP","optional":true},{"name":"Java","optional":true},{"name":"JAX","optional":true},{"name":"LLM","optional":true},{"name":"MLFlow","optional":true},{"name":"PyTorch","optional":true},{"name":"PyTorch C++","optional":true},{"name":"Ray","optional":true},{"name":"Rust","optional":true},{"name":"TensorFlow","optional":true},{"name":"TensorFlow C++","optional":true},{"name":"Terraform","optional":true},{"name":"Vertex AI","optional":true},{"name":"Weights & Biases","optional":true}],"status":"live","first_seen_at":"2026-09-04T06:58:38Z","employer_posted_date":"2026-09-04","last_verified_at":"2026-10-07T23:41:13Z","board_verified":true,"closed_at":null,"days_open":33,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":33},"description":"About ZenteiQ\nZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence\nInfrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national\nsystems. Rather than wrapping general-purpose language models, we train foundation models\n(BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials\ndomains, and put them to work through industrial platforms (KogneX) and a talent OS for\nengineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work\nclosely with IISc, ARTPARK and a national AI Hub Network - building the sovereign AI\ninfrastructure that India's engineering and industrial systems will run on.\nAbout the Role\nYou will build the systems that let our researchers train, evaluate, and serve large foundation\nmodels reliably at scale. This role sits at the intersection of model research and infrastructure,\nwith a focus on accelerator-based (TPU) training and inference, performance, reproducibility,\nand researcher velocity. You will own the paved path that turns expensive, long-running model\nruns into a repeatable, observable, and cost-efficient process.\nWhat You'll Do\nBuild and improve the paved path for distributed training, evaluation, experiment tracking,\ncheckpoint management, model release, and inference on Cloud TPUs.\nOperate long-running ML workloads with strong observability, failure detection, automated\nrecovery, and practical operational tooling.\nProfile and remove bottlenecks across compute, HBM and host memory, data input,\nnetworking and collectives, XLA compilation, checkpointing, and serving.\nBuild validation, representative-scale testing, CI/CD, reproducibility, and lineage systems\nthat catch problems before expensive model runs or production releases.\nCreate reusable APIs, abstractions, and self-service tooling that help researchers move\nquickly while preserving useful low-level controls.\nOwn accelerator capacity workflows, including TPU provisioning, quotas, reservations,\npriorities, topology-aware placement, utilization, and cost efficiency.\nPartner with research teams to debug model and systems failures, translate recurring\npain points into durable platform improvements and define operational standards.\nWhat We're Looking For\nStrong software engineering skills and experience owning systems used by researchers\nor engineers in production or research-critical environments.\nHands-on experience with ML training or inference infrastructure at meaningful scale,\nincluding large accelerator or distributed compute workloads.\nStrong distributed-systems fundamentals and experience with Kubernetes, GKE, or\ncomparable orchestration for long-running compute jobs.\nAbility to debug across model code, data pipelines, runtimes, cluster services, storage,\nnetworking, and accelerator behaviour.\nStrong observability, reliability, and performance-engineering fundamentals, including\nincident response and root-cause analysis.\nProduction-quality Python and the ability to build durable platform abstractions rather than\none-off scripts.\nHigh ownership, pragmatic judgment, and clear collaboration with fast-moving research\nteams.\nGood to Have / Bonus Points\nExperience operating large Cloud TPU clusters, including TPU VMs or Pods, multislice\njobs, topology, quotas, scheduling, and failure recovery.\nExperience with JAX, PyTorch/XLA, TensorFlow, XLA/HLO, PJRT, MaxText, Pallas, or\nanother large-scale training stack.\nExperience serving LLMs on TPUs, including batching, partitioning, compilation caching,\nautoscaling, latency, and throughput optimization.\nFamiliarity with GCP/GKE, Terraform, Vertex AI, XPK, Ray, Argo, Airflow, MLflow, Weights\n& Biases, or comparable internal platforms.\nExperience with model registries, distributed checkpointing, data and artifact lineage, and\nlarge artifact distribution.\nProficiency in a systems language such as Go, Rust, C++, or Java.\nWhy ZenteiQ\nBuild real physics-native foundation models from scratch, not another LLM wrapper -\ndeep, defensible technical work.\nBe part of a nationally recognised mission: one of 8 startups selected under the\ngovernment's IndiaAI Mission to build a sovereign foundation model.\nWork alongside IISc-trained scientists and researchers, in a company founded by an IISc\nprofessor.\nSee your work land in real industry pilots across automotive, mobility, defence and\nindustrial R&D, with measurable impact.\nJoin a lean, high-caliber team at an early, high-ownership stage.","description_format":"text","description_chars":4498,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Machine Learning","STEM Education","Simulation & Digital Twin Software","AI for Science"],"lifecycle":[{"event":"open","at":"2026-09-22T19:09:34Z"}],"visa":[],"liveness":{"score":22,"band":"cold","label":"Long shot","p_open":1,"p_active":0.491,"p_room":0.45,"age_days":32,"expected_fill_days":21,"reasons":["conf:1","win:tail"],"computed_at":"2026-10-07T05:47:15Z"},"pay":null,"html_url":"https://alion.io/job/zenteiq-ml-systems-engineer-training-inference","json_url":"https://alion.io/job/zenteiq-ml-systems-engineer-training-inference.json","meta":{"generated_at":"2026-10-08T00:57:13Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1571,"day_limit":5000,"remaining_today":3429,"minute_limit":60,"resets_at":"2026-10-09T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":1338},"rest":"https://alion.io/mcp/rest/get_company?id=1338"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fzenteiq-ml-systems-engineer-training-inference"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fzenteiq-ml-systems-engineer-training-inference"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fzenteiq-ml-systems-engineer-training-inference"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/zenteiq-ml-systems-engineer-training-inference\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fzenteiq-ml-systems-engineer-training-inference"}]}