Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 19, 2026. Infosys scores B on the Alion truth index.
Join a high-impact Production Engineering/SRE team where reliability, speed, and customer trust are at the center of everything we do. In this role, you’ll lead the engineering effort to keep large-scale, business-critical platforms stable, performant, and continuously improving-while enabling product teams to ship confidently. You’ll work closely with developers, infrastructure, and security partners to design resilient systems, automate operational toil, and build pragmatic observability that turns signals into action. If you enjoy solving complex production challenges, driving incident excellence, and mentoring teams toward modern reliability practices, this is a place where your leadership will be visible and valued. Expect a collaborative culture, ownership-driven execution, and the opportunity to shape reliability standards across services and environments.
Responsibilities
Key Responsibilities:
Reliability & Production Ownership
- Lead production engineering practices to ensure high availability, scalability, and performance across services and platforms.
- Define and drive SLOs/SLIs, error budgets, capacity planning, and reliability roadmaps aligned to business priorities.
- Partner with engineering teams to design resilient architectures and reduce operational risk through proactive improvements.
Incident Management & Operational Excellence
- Own incident response processes (on-call readiness, triage, escalation, communication) and lead major incident bridges when needed.
- Drive blameless postmortems, root-cause analysis, and corrective/preventive actions to prevent recurrence.
- Establish operational runbooks, playbooks, and production readiness reviews for new releases and changes.
Cloud Operations & Automation
- Lead cloud operations to ensure secure, cost-effective, and reliable environments across regions/accounts/subscriptions.
- Identify toil and implement automation to improve deployment safety, recovery time, and operational efficiency.
- Standardize operational tooling and workflows to improve service health, change success rate, and MTTR.
Leadership & Collaboration
- Mentor engineers and influence cross-functional teams to adopt reliability engineering best practices.
- Provide technical leadership in prioritization, execution planning, and stakeholder communication for reliability initiatives.
Minimum Qualifications:
- BTECH, MTECH, MCA, or MSC in Computer Science, IT, or a related field (or equivalent practical experience).
- 12-14 years of experience in SRE, Production Engineering, or Reliability Engineering roles supporting large-scale systems.
- Strong hands-on experience in cloud operations, incident management, and production support for critical services.
- Proven ability to drive automation initiatives that reduce manual effort and improve system reliability.
- Demonstrated experience leading operational processes such as on-call, postmortems, and production readiness practices.
Technical requirements
SRE, Production engineering,Kubernetes, Terraform, CI/CD pipelines, Observability (Prometheus/Grafana)
Additional responsibilities
Preferred Qualifications:
- Experience designing and implementing SLO/SLI frameworks, error budgets, and reliability KPIs across multiple teams.
- Strong background in observability practices (monitoring, alerting, logging, tracing) and building actionable operational dashboards.
- Expertise in release/change management practices that improve deployment safety and reduce production incidents.
- Experience leading cross-team reliability programs, influencing stakeholders, and driving measurable improvements in uptime and MTTR.
- Track record of mentoring engineers and setting engineering standards for operational excellence and automation at scale.
Preferred skills
Technology->DevOps->Site Reliability Engineering(SRE),Domain->Manufacturing->Production Planning
Education
MCA,MSc,MTech,Bachelor of Engineering,BTech

