We are seeking a seasoned Senior DevOps Engineer to join our team in Bangalore. This role is designed for a technical leader who specializes in building robust, scalable infrastructure with a "monitoring-first" mindset. You will be responsible for the stability and visibility of our global platforms, ensuring that our Kubernetes ecosystems are not just functional but deeply observable. As a senior DevOps engineer, you will lead the evolution of our Infrastructure-as-a-Service (IaaS) and Kubernetes-as-a-Service (KaaS) offerings. Your primary mission is to architect and maintain a world-class observability stack that provides real-time insights into system health, performance, and distributed traces.
Responsibilities:
- Observability Architecture: Design, implement, and manage a full-stack observability solution using Prometheus, Grafana, Loki, and Tempo (or equivalent LGTM stack).
- Infrastructure Management: Manage and scale production-grade Amazon EKS clusters and AWS resources, specifically optimizing S3 for long-term log storage and data persistence.
- CI/CD Pipeline Engineering: Own and optimize GitLab CI/CD pipelines to automate software delivery, infrastructure updates (IaC), and AMI rotations.
- Performance and Reliability: Lead troubleshooting efforts for complex distributed systems, utilizing distributed tracing to identify bottlenecks across microservices.
- Automation: Replace manual operational tasks with automated workflows using Terraform, Helm, or Python/Go.
Requirements:
- AWS Mastery: 5+ years of experience with AWS core services, with deep expertise in EKS, S3 IAM, and VPC networking.
The Observability Stack:
- Monitoring: Advanced Prometheus configuration (recording rules, alerting rules) and Thanos/Cortex for long-term retention.
- Logging: Centralized logging architecture using Grafana Loki or ELK.
- Tracing: Implementation of distributed tracing using Tempo, Jaeger, or OpenTelemetry.
- Visualization: Expert-level Grafana dashboarding.
- Container Orchestration: Deep understanding of Kubernetes internals, including ingress controllers, service meshes, and cluster autoscaling.
- DevOps Tooling: Proficiency in GitLab CI/CD, Helm charts, and Terraform for Infrastructure as Code.
- Linux/Scripting: Strong Bash and Python skills for system automation.
Preferred Qualifications:
- Experience managing large-scale migrations or "KaaS" platform releases.
- Knowledge of cost optimization on AWS (spot instances, S3 lifecycle policies).
- Experience with GitOps workflows (e. g., ArgoCD or Flux).

