{"id":1285491,"url":"https://alion.io/job/guidewire-site-reliability-engineer","title":"Site Reliability Engineer","company":{"id":6488,"name":"Guidewire","domain":"guidewire.com","url":"https://alion.io/company/guidewire","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":18000,"max_usd":41000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":42},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Go","optional":false},{"name":"Incident Management","optional":false},{"name":"Java","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Python","optional":false},{"name":"Self-Healing","optional":false},{"name":"Spring Boot","optional":false},{"name":"Agile","optional":true},{"name":"Amazon Aurora","optional":true},{"name":"Amazon CloudWatch","optional":true},{"name":"Amazon EKS","optional":true},{"name":"Apache Kafka","optional":true},{"name":"AWS","optional":true},{"name":"Bitbucket","optional":true},{"name":"CI/CD","optional":true},{"name":"Crossplane","optional":true},{"name":"Datadog","optional":true},{"name":"Docker","optional":true},{"name":"FluxCD","optional":true},{"name":"GitHub Actions","optional":true},{"name":"GitOps","optional":true},{"name":"Helm","optional":true},{"name":"IAM","optional":true},{"name":"Jenkins","optional":true},{"name":"Kanban","optional":true},{"name":"Kubernetes","optional":true},{"name":"Linux","optional":true},{"name":"Okta","optional":true},{"name":"OpenTelemetry","optional":true},{"name":"Prometheus","optional":true},{"name":"Scrum","optional":true},{"name":"TeamCity","optional":true},{"name":"Terraform","optional":true},{"name":"Terragrunt","optional":true}],"status":"live","first_seen_at":"2026-08-14T06:20:14Z","employer_posted_date":null,"last_verified_at":"2026-08-14T06:20:14Z","board_verified":false,"closed_at":null,"days_open":46,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":46},"description":"Drive Reliability, Automation & Scale : \n\n- Design, build, and operate highly reliable, scalable infrastructure for a multi-tenant SaaS platform.\n\n- Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.\n\n- Develop internal tools, services, and frameworks to improve efficiency and reduce manual effort.\n\n- Participate in a 24x7 follow-the-sun on-call rotation to support critical production systems.\n\nImprove Platform & Infrastructure : \n\n- Contribute to core platform systems by building features, resolving issues, and enhancing reliability.\n\n- Partner with development teams to ensure systems meet availability, performance, and scalability requirements.\n\n- Proactively identify risks, bottlenecks, and failure modes, and implement solutions before they impact customers.\n\nObservability, Incident Management & Resilience : \n\n- Build and maintain observability systems (metrics, logging, tracing, dashboards).\n\n- Define and track Service Level Objectives (SLOs) and reliability metrics.\n\n- Lead or contribute to incident response, root cause analysis, and blameless postmortems.\n\n- Drive improvements toward self-healing systems and reduced operational toil.\n\nSecurity & Identity : \n\n- Design and support secure access patterns, including SSO, SAML, and OAuth-based authentication systems.\n\n- Ensure platform services meet security and compliance standards.\n\nEnablement & Collaboration : \n\n- Collaborate across engineering teams, providing guidance, feedback, and hands-on contributions.\n\n- Create and maintain documentation, runbooks, and training materials.\n\n- Mentor engineers and promote best practices in reliability engineering and automation.\n\nWho You Are : \n\nInfrastructure Development : \n\n- 8-12 years of hands-on experience in Site Reliability Engineering (SRE), DevOps, Cloud Infrastructure, or a related platform engineering role, with a proven track record of designing, building, and operating highly available, scalable, and reliable production systems.\n\n- Strong programming skills in Python or Go (Java/Spring Boot is a plus).\n\n- Deep experience with AWS and building/operating production systems at scale.\n\n- Hands-on expertise with Kubernetes (EKS), Docker, Helm, CNI, and Ingress networking.\n\n- Strong understanding of Kubernetes primitives and patterns (deployments, services, operators, etc.).\n\n- Experience with Infrastructure as Code (Terraform, Terragrunt, or similar).\n\n- Solid understanding of Linux systems and networking fundamentals.\n\nObservability & Operations : \n\n- Experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch.\n\n- Familiarity with incident management practices and production support in a microservices environment.\n\n- Experience with messaging/streaming systems (e.g., Kafka, SQS) and relational databases (e.g., Aurora, RDS) is a plus.\n\nSecurity & Identity : \n\n- Working knowledge of SSO, SAML, OAuth, and identity providers (Okta is a plus).\n\n- Experience with AWS IAM (roles, policies, IRSA), VPC security groups, and Kubernetes security primitives (RBAC, network policies, pod security standards, secrets management).\n\nDevOps & Delivery : \n\n- Experience with CI/CD and GitOps tools such as GitHub Actions, TeamCity, Jenkins, FluxCD, or Bitbucket.\n\n- Comfortable working in agile environments (Scrum, Kanban).\n\nMindset & Collaboration : \n\n- Strong troubleshooting and problem-solving skills with a proactive, systems-thinking mindset.\n\n- Passion for automation : If you have to do it more than once, automate it.\n\n- Excellent communication skills and ability to work across distributed teams.\n\n- A collaborative team player who can influence, mentor, and lead through technical expertise.\n\n- Demonstrated ability to leverage AI and data-driven insights to improve productivity and outcomes.\n\nPreferred Qualifications : \n\n- Bachelor's degree in Computer Science or related field, or equivalent experience.\n\n- Experience supporting large-scale SaaS platforms.\n\n- AWS or Kubernetes certifications.\n\n- Exposure to modern platform frameworks such as KubeVela (OAM) or Crossplane.\n\n- Contributions to open-source projects.\n\nWhy Guidewire : \n\n- Work on a mission-critical global platform used by leading P&C insurers worldwide.\n\n- Solve complex, real-world infrastructure problems at genuine scale.\n\n- Be part of a collaborative, high-impact engineering culture grounded in integrity, rationality, and collegiality.\n\n- Opportunity to shape the future of a rapidly evolving cloud platform.\n\n- A culture of curiosity and innovation where engineers are empowered to leverage AI and emerging technologies.\nSkills\nGuidewire, Site Reliability, IT Infrastructure, SaaS, DevOps, Cloud Infrastructure, Python, Golang, Kubernetes, AWS, Observability Services","description_format":"text","description_chars":4786,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["InsurTech"],"lifecycle":[{"event":"open","at":"2026-09-26T04:00:00Z"}],"liveness":{"score":10,"band":"cold","label":"Long shot","p_open":0.4,"p_active":0.573,"p_room":0.45,"age_days":45,"expected_fill_days":23,"reasons":["seen:45","win:tail"],"computed_at":"2026-09-29T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/guidewire-site-reliability-engineer","json_url":"https://alion.io/job/guidewire-site-reliability-engineer.json","meta":{"generated_at":"2026-09-30T05:12:05Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3575,"day_limit":5000,"remaining_today":1425,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}