{"id":1248034,"url":"https://alion.io/job/aivar-innovations-senior-kubernetes-platform-engineer","title":"Senior Kubernetes Platform Engineer","company":{"id":2036,"name":"Aivar Innovations","domain":"aivar.tech","url":"https://alion.io/company/aivar-innovations","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":18500,"max_usd":42000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":42},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"ArgoCD","optional":false},{"name":"CI/CD","optional":false},{"name":"Crossplane","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"DNS","optional":false},{"name":"Envoy","optional":false},{"name":"Helm","optional":false},{"name":"kubectl","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"Loki","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Prometheus","optional":false},{"name":"Rancher","optional":false},{"name":"WebSockets","optional":false},{"name":"AWS","optional":true},{"name":"Grafana","optional":true}],"status":"closed","first_seen_at":"2026-08-20T15:20:19Z","employer_posted_date":"2026-08-20","last_verified_at":"2026-09-30T07:10:04Z","board_verified":false,"closed_at":"2026-09-30T07:10:04Z","days_open":40,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":40},"description":"About Us\nAivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes-from intelligent customer interactions to enterprise knowledge systems.\nExperience: 5-9 years | 4+ years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure\nThe Role: You build the Kubernetes substrate that makes Kubogent possible.\nKubogent is a Kubernetes-native AI infrastructure and MLOps platform. A central control plane manages multiple workload clusters, allocates infrastructure to tenants and projects, deploys platform capabilities on demand, runs GPU and ML workloads, exposes remote operational access, and continuously reconciles desired state with what is actually running.\nThis role is for someone who understands Kubernetes as a distributed system and programmable control plane, not just as a deployment target.\nWhat you'll do\nOwn Kubernetes control loops. Design and build CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines using Go and controller-runtime.\nBuild multi-cluster control. Help design and implement how Kubogent onboards, authenticates, observes, upgrades, and controls workload clusters that may only initiate outbound connections to the control plane.\nBuild the workload-cluster agent. Own registration, heartbeat, inventory, command execution, reconnect behaviour, versioning, rollout, credential rotation, and failure recovery.\nTranslate platform intent into Kubernetes state. Convert concepts such as project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services into safe, idempotent Kubernetes reconciliation.\nOwn resource isolation and placement. Work with namespaces, quotas, limits, priority, scheduling, taints/tolerations, affinity, topology, GPU resources, gang scheduling, and workload placement.\nBuild GPU and accelerator support. Integrate with NVIDIA GPU Operator, device plugins, MIG where appropriate, node feature discovery, topology-aware scheduling, and accelerator-specific runtime requirements.\nBuild secure remote operations. Design mechanisms for browser-based kubectl/exec/log access, tunneled cluster connectivity, least-privilege credentials, mTLS, authorization, and auditable command execution.\nOwn cluster capability deployment. Build mechanisms that install only the operators and services required by capabilities enabled on projects or clusters, rather than treating every cluster as identical.\nDesign for unreliable environments. Clusters disappear, links break, agents restart, watches expire, APIs throttle, upgrades partially fail, and reconciliation gets repeated. Your systems must remain correct anyway.\nWork deeply with Kubernetes API machinery. Informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.\nOwn platform upgrades. Design safe version skew, agent upgrades, CRD evolution, migration, backward\ncompatibility, and rollout/rollback strategies.\nDrive observability for the platform itself. Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.\nSet the bar. Review designs and code, mentor engineers, and establish patterns for building reliable Kubernetes-native systems.\nWhat we're looking for\n5+ years in software, platform, infrastructure, SRE, or cloud engineering, with substantial hands-on Kubernetes experience.\nStrong Go. You should be comfortable designing production services, concurrency, interfaces, testing, profiling, and failure handling in Go.\nDeep Kubernetes internals. Controllers, CRDs, reconciliation, API machinery, watches/informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.\nYou have built Kubernetes software, not only operated clusters. We especially value experience with Kubebuilder, controller-runtime, Operator SDK, custom schedulers, admission webhooks, or Kubernetes-integrated platforms.\nStrong distributed-systems instincts. Idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence should be familiar ideas.\nExperience operating Kubernetes across environments. Cloud-managed Kubernetes and on-prem/private-cloud experience are both valuable.\nComfort with networking. TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress/gateway, and debugging connectivity failures.\nProduction troubleshooting ability. You can move from symptom to root cause across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.\nStrong security fundamentals. Service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.\nFluency with agentic coding tools. We expect AI to accelerate implementation and investigation while you remain responsible for architecture, correctness, failure handling, and operational quality.\nStrong pluses:\nMulti-cluster management platforms.\nKubernetes API aggregation or extension patterns.\nCluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, or similar systems.\nEnvoy, reverse tunnels, relay systems, or secure remote cluster access.\nGPU scheduling, NVIDIA GPU Operator, MIG, CUDA-aware workloads, or distributed training infrastru\nPrometheus, OpenTelemetry, Loki, or Kubernetes observability stacks.\nKubernetes conformance, upgrade testing, chaos testing, or large-scale fleet management.\nThis role may not be the best match if:\nYour Kubernetes experience is primarily writing manifests, Helm charts, and maintaining CI/CD pipelines. We arebuilding Kubernetes-native control-plane software.\nYou expect infrastructure to be reliable and synchronous. Disconnected clusters, duplicate commands, partial\nupgrades, stale state, and repeated reconciliation are normal operating conditions here.\nYou prefer solving platform problems by adding manual operational procedures. Kubogent must turn those procedures into productized, automated control loops.","description_format":"text","description_chars":6365,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Speech & Audio AI","AI Consulting & Integration","Cloud Consulting & Migration"],"lifecycle":[{"event":"open","at":"2026-09-25T18:07:15Z"},{"event":"close","at":"2026-09-30T07:10:04Z"}],"liveness":null,"pay":null,"html_url":"https://alion.io/job/aivar-innovations-senior-kubernetes-platform-engineer","json_url":"https://alion.io/job/aivar-innovations-senior-kubernetes-platform-engineer.json","meta":{"generated_at":"2026-10-01T18:43:18Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":995,"day_limit":5000,"remaining_today":4005,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}