{"id":1227143,"url":"https://alion.io/job/mindsprint-senior-site-reliability-engineer","title":"Senior Site Reliability Engineer","company":{"id":3223551,"name":"Mindsprint","domain":"mindsprint.com","url":"https://alion.io/company/mindsprint-com","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Darwinbox","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":29000,"max_usd":63000,"period":"year","method":"role_seniority_country_cell","sample_n":8},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon Aurora","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon EC2","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon S3","optional":false},{"name":"ArgoCD","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"Azure","optional":false},{"name":"Azure AKS","optional":false},{"name":"Azure DevOps","optional":false},{"name":"Blue-Green Deployment","optional":false},{"name":"Canary Release","optional":false},{"name":"Chaos Engineering","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Commander.js","optional":false},{"name":"Datadog","optional":false},{"name":"Dynatrace","optional":false},{"name":"GitHub Actions","optional":false},{"name":"GitLab CI","optional":false},{"name":"Grafana","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"ISO 27001","optional":false},{"name":"Jenkins","optional":false},{"name":"Karpenter","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"Microsoft Entra ID","optional":false},{"name":"New Relic","optional":false},{"name":"OpenSearch","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"PCI DSS","optional":false},{"name":"Progressive Delivery","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Self-Healing","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"SOC 2","optional":false},{"name":"Terraform","optional":false},{"name":"Amazon Kinesis","optional":true},{"name":"Apache Kafka","optional":true},{"name":"AWS CDK","optional":true},{"name":"DynamoDB","optional":true},{"name":"FinOps","optional":true},{"name":"Istio","optional":true},{"name":"JavaScript","optional":true},{"name":"Linkerd","optional":true},{"name":"Node JS","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Pulumi","optional":true},{"name":"Service Mesh","optional":true}],"status":"live","first_seen_at":"2026-09-25T11:47:55Z","employer_posted_date":null,"last_verified_at":"2026-09-25T11:47:55Z","board_verified":false,"closed_at":null,"days_open":2,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":2},"description":"Role Summary :\n\nWe are hiring a senior Site Reliability Engineer to own the availability, scalability, performance and operability of our production platform. The estate is AWS-first with a growing Azure footprint, and is fully provisioned as code using Terraform and CloudFormation. This is a hands-on engineering role covering all pillars of SRE infrastructure, automation, observability, SLO management, incident response, resilience and security with mentoring responsibility for mid-level engineers and participation in a shared on-call rotation.\n\nKey Responsibilities :\n\n- Cloud Infrastructure (AWS primary, Azure secondary): Architect, build and operate scalable AWS infrastructure (VPC and connectivity, IAM, compute, storage, managed databases, serverless, multi-account governance) and own production Amazon EKS clusters end to end provisioning, upgrades, autoscaling, networking and security. Build and support the secondary Azure footprint and drive cloud cost optimisation.\n\n- Infrastructure as Code: Design and maintain reusable Terraform modules and CloudFormation stacks across multi-account and multi-subscription environments, with remote state, versioning, drift detection and policy-as-code guardrails. Eliminate manual console changes.\n\n- Monitoring & Observability: Own the metrics, logs and tracing stack; build symptom- and SLO-burn-based alerting; instrument services with development teams; and maintain golden dashboards and runbooks that any on-call engineer can use.\n\n- SLIs, SLOs & Error Budgets: Define SLIs for critical user journeys, agree SLOs with stakeholders, operate error-budget policy to balance feature velocity against reliability, and report reliability to leadership through regular service reviews.\n\n- Automation & Toil Reduction: Measure and engineer away repetitive operational work; build tooling in Python, Bash or Go for provisioning, remediation, patching and diagnostics; and implement self-healing and auto-remediation for known failure modes.\n\n- Incident Response & On-Call: Participate in a compensated on-call rotation, act as Incident Commander for high-severity events, run blameless postmortems with corrective actions tracked to closure, and drive down MTTD and MTTR.\n\n- CI/CD & Release Engineering: Build and harden application and infrastructure pipelines; enable blue-green, canary and progressive delivery with automated health gates and fast rollback; and improve DORA metrics including change failure rate.\n\n- Backup, DR & Business Continuity: Own the backup estate across Azure Backup (Recovery Services vaults, retention, immutability, cross-region restore) and AWS Backup; define RPO/RTO per service; and prove recoverability through scheduled restore tests and DR failover drills.\n\n- Capacity Planning & Performance: Forecast demand and plan capacity across compute, storage, network and database tiers; run load and stress testing to validate headroom and autoscaling; tune performance and practise chaos engineering to validate resilience assumptions.\n\n- Security, Compliance & Governance: Enforce least-privilege IAM and centralised secrets management, maintain patch and vulnerability hygiene across hosts and images, and support audit requirements (SOC 2 / ISO 27001 / PCI-DSS) with CIS-aligned hardened baselines.\n\nMust-Have Skills:\n\n- AWS (Primary): EC2, VPC & network design, IAM, S3, RDS/Aurora, ELB/ALB, Route 53, Lambda, CloudWatch, CloudTrail, AWS Backup, Organizations / multi-account governance.\n\n- Kubernetes / EKS: Production EKS ownership cluster provisioning & upgrades, Karpenter / Cluster Autoscaler, HPA, Helm, ingress, IRSA, RBAC, network policies, deep troubleshooting (CNI, CoreDNS, CSI).\n\n- Azure (Secondary): Azure Backup & Recovery Services vaults (mandatory), Azure Site Recovery, VMs, VNets, Entra ID, Storage, AKS, Azure Monitor / Log Analytics.\n\n- IaC: Terraform at expert level (modules, remote state, workspaces, drift detection, CI-driven plan/apply) and strong AWS CloudFormation (nested stacks, StackSets, change sets).\n\n- Observability: Prometheus, Grafana, CloudWatch, Azure Monitor, ELK/OpenSearch, OpenTelemetry; plus one APM (Datadog / New Relic / Dynatrace). SLO and error-budget dashboards.\n\n- SRE Practice: Hands-on SLI/SLO definition, error-budget policy, incident command, blameless postmortems, toil reduction, capacity planning, chaos/DR testing.\n\n- CI/CD: GitHub Actions, GitLab CI, Jenkins or Azure DevOps; ArgoCD or Flux; blue-green and canary deployment with automated rollback.\n\nQualifications:\n\n- Experience: 812 years in IT infrastructure, cloud or DevOps, including a minimum of 5 years in a dedicated SRE / DevOps / Cloud Infrastructure engineering role with production ownership.\n\n- Education: Bachelor's degree in Computer Science, Engineering or equivalent practical experience.\n\n- Certifications (preferred): AWS Solutions Architect / DevOps Engineer Professional; CKA or CKS; Azure AZ-104 or AZ-305; HashiCorp Terraform Associate.\n\n- Soft skills: Strong ownership, calm and structured under production pressure, data-driven on reliability trade-offs, and able to explain risk clearly to non-technical stakeholders.\n\nGood to Have:\n\n- Go for operational tooling or Kubernetes operators; AWS CDK or Pulumi; service mesh (Istio / Linkerd / App Mesh) in production.\n\n- Chaos engineering tooling (AWS FIS, Chaos Mesh, Gremlin); database reliability engineering (Aurora, PostgreSQL, DynamoDB); Kafka / MSK / Kinesis at scale.\n\n- FinOps and cloud cost ownership; data-centre-to-cloud or AWS-to-Azure migration experience; regulated-industry exposure.\nSkills\nDevOps, Site Reliability, AWS, Kubernetes, Terraform, CI/CD Pipeline, Azure, AWS Lambda, CloudFormation, Cloud Infrastructure","description_format":"text","description_chars":5707,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T13:06:44Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":30,"reasons":["seen:1","win:early"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/mindsprint-senior-site-reliability-engineer","json_url":"https://alion.io/job/mindsprint-senior-site-reliability-engineer.json","meta":{"generated_at":"2026-09-28T01:02:56Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":619,"day_limit":5000,"remaining_today":4381,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}