845,403open jobs
53,832companies
143,151added this week
Browse all
Salary
≈ $25k – $56k per year (Estimated)
Location
In office (Pune)
Seniority
Senior · 6+ years exp

First seen by Alion on Sep 18, 2026.

Overview
Company
Impact
Profile match
With. Consulting Pandits is a team of Senior Recruiters, HR Strategists and Tech Consultants with 10 - 20 Years of hiring experience in Domestic and Global Markets.

Key Responsibilities :

- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.

- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.

- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.

- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.

- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.

- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.

- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.

- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.

- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.

- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.

- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.

- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.

- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.

- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.

- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each change.

Minimum Requirements :

- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.

- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.

- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.

- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.

- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.

- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.

- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.

- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.

- Scripting in Python and Bash, and comfort with YAML-heavy configuration.

- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call rotation.

Preferred Requirements :

- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.

- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.

- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.

- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments.

Skills

Kubernetes, IAC Terraform, Ansible, Linux, Python, Prometheus, Grafana, PostgreSQL, Redis, Cloud Infrastructure, DevOps

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
845,403 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Pune
≈ $57k – $143k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Luxembourg City
C++
AI/ML
Anomaly Detection
DevOps
CI/CD
AWS
Platform Engineering
Apply
In office • Full-Time • Taipei
C#
C#
.NET
DevOps
CI/CD
Docker
Kubernetes
Apply
≈ $17k – $42k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Kuala Lumpur
PHP
PowerShell
PHP
Laravel
DevOps
Terraform
GitHub Actions
Azure
CI/CD
GitHub
Apply
≈ $89k – $165k per year (Estimated) • In office • Full-Time • London • Glasgow
Python
Java
Groovy
DevOps
Terraform
CI/CD
Jenkins
AWS
Kubernetes
Platform Engineering
GitHub
Apply
In office • 2+ years exp
DevOps
Ansible
OpenShift
OpenStack
Apply
≈ $29k – $60k per year (Estimated) • In office • 6+ years exp • Pune
SQL
AI/ML
Anomaly Detection
Apply
≈ $23k – $62k per year (Estimated) • In office • 6+ years exp • Gandhinagar • Mumbai • Pune • Bengaluru
Apply
≈ $11k – $28k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Pune
Python
SQL
Databases
MySQL
Oracle
DevOps
GCP
AWS
Incident Management
Unix
SOAP
Apply
≈ $21k – $44k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Pune • Chennai • Bengaluru
JavaScript
Java
TypeScript
SQL
Java
Maven
Spring Boot
Databases
MySQL
PostgreSQL
Oracle
Frontend
React.js
DevOps
Rest API
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
In office • Full-Time • 3+ years exp • Pune
Apply
See all jobs
This is one of many
845,403 more open roles from verified company boards, updated every day.