{"id":1230652,"url":"https://alion.io/job/consulting-pandits-platform-engineer","title":"Platform Engineer","company":{"id":3800739,"name":"Consulting Pandits","domain":"consultingpandits.com","url":"https://alion.io/company/consulting-pandits","size_band":null,"is_staffing_agency":true,"employer_type":"agency","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":25000,"max_usd":56000,"period":"year","method":"role_seniority_country_cell","sample_n":8},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"Bash","optional":false},{"name":"DNS","optional":false},{"name":"etcd","optional":false},{"name":"GitOps","optional":false},{"name":"Grafana","optional":false},{"name":"Helm","optional":false},{"name":"HPC","optional":false},{"name":"IAM","optional":false},{"name":"K3s","optional":false},{"name":"Kubernetes","optional":false},{"name":"Kyverno","optional":false},{"name":"Linux","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Redis","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-09-18T04:13:49Z","employer_posted_date":null,"last_verified_at":"2026-09-18T04:13:49Z","board_verified":false,"closed_at":null,"days_open":10,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":10},"description":"Key Responsibilities :\n\n- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.\n\n- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.\n\n- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.\n\n- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.\n\n- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.\n\n- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.\n\n- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.\n\n- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.\n\n- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.\n\n- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.\n\n- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.\n\n- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.\n\n- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.\n\n- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.\n\n- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each change.\n\nMinimum Requirements :\n\n- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.\n\n- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.\n\n- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.\n\n- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.\n\n- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.\n\n- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.\n\n- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.\n\n- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.\n\n- Scripting in Python and Bash, and comfort with YAML-heavy configuration.\n\n- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call rotation.\n\nPreferred Requirements :\n\n- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.\n\n- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.\n\n- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.\n\n- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments.\nSkills\nKubernetes, IAC Terraform, Ansible, Linux, Python, Prometheus, Grafana, PostgreSQL, Redis, Cloud Infrastructure, DevOps","description_format":"text","description_chars":4583,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T14:00:00Z"}],"liveness":{"score":43,"band":"fade","label":"Fading","p_open":1,"p_active":0.455,"p_room":0.945,"age_days":9,"expected_fill_days":24,"reasons":["seen:9","agency","velocity","win:mid"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/consulting-pandits-platform-engineer","json_url":"https://alion.io/job/consulting-pandits-platform-engineer.json","meta":{"generated_at":"2026-09-28T04:41:19Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2997,"day_limit":5000,"remaining_today":2003,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}