791,964open jobs
50,470companies
123,351added this week
Browse all
Salary
≈ $30k – $66k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Aug 20, 2026.

Overview
Company
Impact
Profile match
Aivar Innovations is a tech company delivering IT services with a focus on AI, ML, and cloud solutions to help businesses drive digital transformation by developing AI-powered platforms and voice AI solutions.

About Us

Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes-from intelligent customer interactions to enterprise knowledge systems.

Experience: 5-9 years | 4+ years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure

The Role: You build the Kubernetes substrate that makes Kubogent possible.

Kubogent is a Kubernetes-native AI infrastructure and MLOps platform. A central control plane manages multiple workload clusters, allocates infrastructure to tenants and projects, deploys platform capabilities on demand, runs GPU and ML workloads, exposes remote operational access, and continuously reconciles desired state with what is actually running.

This role is for someone who understands Kubernetes as a distributed system and programmable control plane, not just as a deployment target.

What you'll do

  • Own Kubernetes control loops. Design and build CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines using Go and controller-runtime.
  • Build multi-cluster control. Help design and implement how Kubogent onboards, authenticates, observes, upgrades, and controls workload clusters that may only initiate outbound connections to the control plane.
  • Build the workload-cluster agent. Own registration, heartbeat, inventory, command execution, reconnect behaviour, versioning, rollout, credential rotation, and failure recovery.
  • Translate platform intent into Kubernetes state. Convert concepts such as project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services into safe, idempotent Kubernetes reconciliation.
  • Own resource isolation and placement. Work with namespaces, quotas, limits, priority, scheduling, taints/tolerations, affinity, topology, GPU resources, gang scheduling, and workload placement.
  • Build GPU and accelerator support. Integrate with NVIDIA GPU Operator, device plugins, MIG where appropriate, node feature discovery, topology-aware scheduling, and accelerator-specific runtime requirements.
  • Build secure remote operations. Design mechanisms for browser-based kubectl/exec/log access, tunneled cluster connectivity, least-privilege credentials, mTLS, authorization, and auditable command execution.
  • Own cluster capability deployment. Build mechanisms that install only the operators and services required by capabilities enabled on projects or clusters, rather than treating every cluster as identical.
  • Design for unreliable environments. Clusters disappear, links break, agents restart, watches expire, APIs throttle, upgrades partially fail, and reconciliation gets repeated. Your systems must remain correct anyway.
  • Work deeply with Kubernetes API machinery. Informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.
  • Own platform upgrades. Design safe version skew, agent upgrades, CRD evolution, migration, backward
  • compatibility, and rollout/rollback strategies.
  • Drive observability for the platform itself. Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.
  • Set the bar. Review designs and code, mentor engineers, and establish patterns for building reliable Kubernetes-native systems.

What we're looking for

  • 5+ years in software, platform, infrastructure, SRE, or cloud engineering, with substantial hands-on Kubernetes experience.
  • Strong Go. You should be comfortable designing production services, concurrency, interfaces, testing, profiling, and failure handling in Go.
  • Deep Kubernetes internals. Controllers, CRDs, reconciliation, API machinery, watches/informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.
  • You have built Kubernetes software, not only operated clusters. We especially value experience with Kubebuilder, controller-runtime, Operator SDK, custom schedulers, admission webhooks, or Kubernetes-integrated platforms.
  • Strong distributed-systems instincts. Idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence should be familiar ideas.
  • Experience operating Kubernetes across environments. Cloud-managed Kubernetes and on-prem/private-cloud experience are both valuable.
  • Comfort with networking. TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress/gateway, and debugging connectivity failures.
  • Production troubleshooting ability. You can move from symptom to root cause across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.
  • Strong security fundamentals. Service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.
  • Fluency with agentic coding tools. We expect AI to accelerate implementation and investigation while you remain responsible for architecture, correctness, failure handling, and operational quality.

Strong pluses:

  • Multi-cluster management platforms.
  • Kubernetes API aggregation or extension patterns.
  • Cluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, or similar systems.
  • Envoy, reverse tunnels, relay systems, or secure remote cluster access.
  • GPU scheduling, NVIDIA GPU Operator, MIG, CUDA-aware workloads, or distributed training infrastru
  • Prometheus, OpenTelemetry, Loki, or Kubernetes observability stacks.
  • Kubernetes conformance, upgrade testing, chaos testing, or large-scale fleet management.

This role may not be the best match if:

  • Your Kubernetes experience is primarily writing manifests, Helm charts, and maintaining CI/CD pipelines. We arebuilding Kubernetes-native control-plane software.
  • You expect infrastructure to be reliable and synchronous. Disconnected clusters, duplicate commands, partial
  • upgrades, stale state, and repeated reconciliation are normal operating conditions here.
  • You prefer solving platform problems by adding manual operational procedures. Kubogent must turn those procedures into productized, automated control loops.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
791,964 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Bengaluru
Looker Developer 3 months ago
≈ $24k – $52k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
SQL
Databases
PostgreSQL
Snowflake
LookML
Google BigQuery
Amazon Redshift
BigQuery
DevOps
GCP
CI/CD
Git
Cybersecurity
ISO 27001
Analytics
Tableau
Power BI
ETL/ELT
Looker
Dimensional Modeling
Management
Agile
Scrum
Apply
AWS DevOps Engineer 1 month ago
≈ $16k – $48k per year (Estimated) • In office • Full-Time • 4+ years exp • Hyderabad
Python
Go
Bash
Databases
PostgreSQL
Redis
DynamoDB
RabbitMQ
Apache Kafka
OpenSearch
DevOps
Terraform
Ansible
Helm
GitHub Actions
Istio
Terragrunt
CircleCI
CloudFormation
FluxCD
Linkerd
Prometheus
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
Blue-Green Deployment
Service Mesh
Configuration Management
Amazon EKS
AWS Lambda
Amazon EC2
FinOps
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Linux
Cybersecurity
SonarQube
Trivy
ISO 27001
Checkov
Management
Agile
Apply
≈ $23k – $50k per year (Estimated) • In office • 6+ years exp • Hyderabad
DevOps
Rest API
SOAP
Apply
Apps DBA 5 months ago
≈ $27k – $60k per year (Estimated) • Remote (United States, India) • 5+ years exp • Bachelor's Degree • Hyderabad
Assembly
Databases
Oracle
DevOps
Azure
AWS
Linux
Apply
Product Architect 2 months ago
≈ $39k – $86k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
JavaScript
Java
Node JS
Java
Spring Boot
Databases
MySQL
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
AI Agents
AWS Bedrock
RAG
OpenAI
Anthropic
Frontend
React.js
DevOps
Rest API
CI/CD
AWS
Docker
Kubernetes
Self-Healing
QA
Selenium
Cypress
Playwright
Appium
Apply
Associate AI Developer 10 hours ago
$40k per year • Hybrid • 2+ years exp • Bachelor's Degree
Python
Python
Pydantic
Databases
Weaviate
Pinecone
FAISS
AI/ML
LangChain
Model Context Protocol
CUDA Toolkit
MLFlow
Computer Vision
AI Agents
AWS Bedrock
TensorFlow
PyTorch
RAG
CUDA
Triton
OpenAI
Amazon SageMaker
Google AI Studio
OCR
Tool Use
Machine Learning
DevOps
OpenTelemetry
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$140k – $200k per year • Equity 0.2–1% • Remote (United States) • Full-Time • 3+ years exp • New York
Python
Go
Rust
DevOps
gRPC
Terraform
Helm
Kustomize
Kubernetes
Platform Engineering
GitHub
Cybersecurity
Snyk
Apply
≈ $104k – $225k per year (Estimated) • In office • Contractor • PhD • Lausanne
Python
Rust
C++
Assembly
OCaml
C++
PyTorch C++
Assembly
Keystone Engine
AI/ML
vLLM
CUDA Toolkit
AI Agents
SGLang
PyTorch
LLM
CUDA
Triton
Agentic Workflows
DevOps
CI/CD
Apply
$76k – $88k per year • In office
Python
JavaScript
TypeScript
SQL
Databases
MySQL
PostgreSQL
MariaDB
AI/ML
Machine Learning
Frontend
Next.js
React.js
DevOps
Rest API
Helm
CI/CD
Git
Docker
Kubernetes
Nginx
Grafana
Analytics
Power BI
ETL/ELT
Superset
Management
Slack
Miro
Confluence
Notion
Jira
Google Workspace
QA
Cypress
Playwright
Apply
$151k – $251k per year • In office • 10+ years exp • Bachelor's Degree • Akron
AI/ML
Model Context Protocol
AI Agents
Guardrails AI
NeMo Guardrails
LLM
RAG
A2A
Human-in-the-Loop
Red Teaming
LLM Guardrails
EU AI Act
NIST AI RMF
DevOps
Helm
Kubernetes
Cybersecurity
OWASP Top 10
CVE
Zero Trust
SBOM
SLSA
Sigstore
Cosign
Kyverno
PKI
Apply
≈ $29k – $63k per year (Estimated) • In office • Full-Time • 4+ years exp • Coimbatore
Python
Java
DevOps
Terraform
GCP
GitHub Actions
CloudFormation
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Linux
Apply
DevOps Engineer 1 month ago
In office • Full-Time • 3+ years exp • Coimbatore
Python
Java
DevOps
Terraform
GCP
GitHub Actions
CloudFormation
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Linux
Apply
≈ $30k – $65k per year (Estimated) • In office • 5+ years exp • Bengaluru
AI/ML
Llama
AWS Trainium
DevOps
Terraform
GCP
Helm
Istio
Kustomize
Prometheus
CI/CD
ArgoCD
AWS
Kubernetes
Grafana
Karpenter
Amazon EKS
Amazon EC2
AIOps
IAM
Linux
Cybersecurity
Trivy
Falco
Calico
Apply
≈ $29k – $62k per year (Estimated) • In office • Full-Time • 6+ years exp • Coimbatore
Python
Go
Java
Kotlin
Databases
PostgreSQL
Neo4j
pgvector
DynamoDB
ElasticSearch
Apache Kafka
OpenSearch
ArangoDB
Amazon Neptune
AI/ML
Model Context Protocol
LLM
RAG
DevOps
Rest API
gRPC
Terraform
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon Kinesis
Analytics
Master Data Management
Apply
≈ $29k – $73k per year (Estimated) • In office • Full-Time • 4+ years exp • Coimbatore
Python
AI/ML
Claude
LoRA
Fine-tuning
Prompt Engineering
Knowledge Distillation
AI Agents
NLP
NER
PEFT
QLoRA
AWS Bedrock
LLM
RAG
Hallucination
Anthropic
Amazon SageMaker
GraphRAG
Structured Outputs
LLM Evaluation
Model Distillation
DevOps
AWS
Analytics
A/B Testing
Apply
≈ $34k – $82k per year (Estimated) • In office • Bengaluru
Python
TypeScript
SQL
AI/ML
Function Calling
Speech Recognition
LLM
RAG
Text-to-Speech
LLM Guardrails
DevOps
GCP
AWS
Apply
≈ $41k – $91k per year (Estimated) • In office • Full-Time • 8+ years exp • Bengaluru
AI/ML
LLM
RAG
Agentic Workflows
DevOps
Azure
AWS
Apply
DevOps Engineer 1 day ago
In office • 4+ years exp • Bengaluru
Python
PowerShell
Perl
DevOps
Terraform
Ansible
CloudFormation
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Configuration Management
Amazon EKS
Amazon ECS
TCP/IP
Management
Kanban
ITIL
Apply
DevOps Engineer 4 hours ago
$15k – $25k per year • In office • Full-Time • 3+ years exp • PhD • Bengaluru
Databases
ClickHouse
DevOps
CI/CD
AWS
Kubernetes
Amazon EKS
Apply
$26k – $52k per year • Equity 0–0.1% • In office • Full-Time • 3+ years exp • Bengaluru
Python
JavaScript
SQL
Node JS
Databases
DynamoDB
AI/ML
Cursor
LangChain
Claude Code
Embeddings
Prompt Engineering
RAG
OpenAI
Hugging Face
Machine Learning
DevOps
AWS
AWS Lambda
Amazon EC2
IAM
Amazon ECS
Apply
See all jobs
This is one of many
791,964 more open roles from verified company boards, updated every day.