Role Summary
We are seeking an experienced Manager, Software Engineering (SRE) to lead a team focused on improving the reliability, scalability, and operational maturity of Sophos cloud services and engineering delivery systems. In this role, you will lead a team distributed across the U.S. and Canada. The team is responsible for production reliability, operational readiness, incident response, escalation management, observability, automation, release reliability, and continuous improvement across cloud-based platforms.
This role requires a leader with a strong background spanning site reliability engineering, platform engineering, DevOps practices, release engineering, infrastructure management, and cloud operations. You should have enough technical depth to guide discussions around AWS, Kubernetes/EKS, CI/CD, Linux, infrastructure as code, observability, cost optimization, security practices, compliance needs, and automation, while primarily focusing on team leadership, execution, stakeholder alignment, prioritization, and improving how engineering teams deliver and operate services at scale.
What You Will Do
- Lead, coach, and develop a team of engineers focused on site reliability, platform operations, release engineering, and automation. Set clear priorities, expectations, and delivery goals aligned to business and engineering needs.
- Support hiring, onboarding, performance management, and career development for team members.
- Foster a culture of ownership, collaboration, continuous improvement, and operational excellence.
- Partner with senior leadership to define and execute the roadmap for SRE and operational improvements.
- Drive improvements in reliability, availability, scalability, deployment confidence, and production readiness.
- Drive operational efficiency and cost optimization practices across cloud platforms, balancing reliability, performance, and responsible cloud spend.
- Help establish best practices for incident response, escalation management, observability, runbooks, post-incident reviews, and toil reduction.
- Promote the use of reliability metrics, operational health indicators, and continuous improvement practices.
- Partner with engineering, architecture, security, release, and operations teams to improve shared tooling and delivery workflows.
- Support improvements across CI/CD pipelines, release automation, deployment reliability, and environment stability. Guide team efforts involving cloud infrastructure, Kubernetes/EKS, Linux systems, Terraform/IaC, Ansible, and automation tooling.
- Encourage responsible use of approved AI tools to improve troubleshooting, documentation, automation, and engineering productivity.
- Ensure effective incident and escalation management practices are in place for timely response, communication, resolution, and follow-up.
- Assess operational situations quickly, prioritize response activities, and balance customer impact, business risk, reliability, security, and delivery needs.
- Drive root cause analysis and long-term corrective actions to reduce repeat incidents.
- Improve observability, alerting, monitoring, and operational visibility across services and platforms.
- Support capacity planning and operational readiness for scalable, high-availability systems.
- Partner with security, AppSec, compliance, and engineering teams to ensure operational practices support secure and reliable service delivery.
- Help manage CI/CD security gates, vulnerability remediation workflows, audit readiness, and compliance-related operational controls.
- Support vendor management for SRE, platform, observability, security, and cloud tooling, including renewals, evaluations, usage reviews, and service escalations.
- Ensure team documentation, runbooks, operational evidence, and process controls are maintained to support audits and internal reviews.
- Lead sprint planning, backlog management, prioritization, and delivery tracking for SRE and platform initiatives.
- Balance competing priorities across reliability work, operational escalations, security needs, technical debt, roadmap delivery, and team capacity.
- Participate in quarterly planning, roadmap development, and cross-team dependency management.
- Work with engineering, product, architecture, and security stakeholders to align reliability and operational priorities with business needs.
- Use Agile/Scrum practices to improve execution, transparency, predictability, and continuous improvement across the team.
- Work closely with engineering teams to understand service needs and ensure SRE priorities align with product and platform goals.
- Lead technical and operational discussions with engineering leaders, product stakeholders, and executive audiences, including Directors, VPs, and C-level staff.
- Present reliability updates, operational risks, roadmap progress, incident trends, and investment needs in a clear, concise, and audience-appropriate way.
- Facilitate meetings, drive decisions, manage follow-ups, and ensure cross-functional conversations result in clear ownership and action.
- Communicate status, risks, operational trends, and improvement opportunities to technical and non-technical stakeholders.
Leadership and Team Management
SRE Strategy and Operational Maturity
Platform, Release, and Automation Enablement
Incident Management and Reliability
Security, Compliance, and Governance
Planning and Delivery
Collaboration and Communication
What You Will Bring
- Extensive experience in site reliability engineering, DevOps practices, platform engineering, infrastructure management, operational efficiency, cloud infrastructure, or production operations.
- Proven experience leading and managing engineers, fostering a culture of collaboration, accountability, ownership, and continuous improvement.
- Strong understanding of reliability engineering, operational excellence, incident management, escalation management, observability, production support, and scalable system operations.
- Strong working knowledge of cloud technologies, especially AWS; Azure experience is a plus.
- Familiarity with operating and supporting containerized platform environments, including Kubernetes/EKS, Docker, Helm, and Linux-based systems.
- Familiarity with infrastructure as code, configuration management, and automation practices using tools such as Terraform and Ansible.
- Understanding of CI/CD pipelines, release automation, deployment practices, and developer tooling such as Jenkins, GitHub Actions, Artifactory, or similar tools.
- Proficiency with scripting and automation concepts using languages such as Python, Bash, or Groovy.
- Experience with monitoring, logging, and observability tools such as ELK stack, Prometheus, Grafana, CloudWatch, or similar platforms.
- Working knowledge of security, application security, vulnerability management, compliance controls, audit readiness, and vendor management in a cloud or SaaS environment.
- Familiarity with Agile/Scrum practices, sprint planning, backlog management, quarterly planning, roadmap development, and cross-team dependency management.
- Excellent problem-solving and analytical skills, with the ability to assess situations quickly, prioritize work, manage escalations, and balance competing technical and business priorities.
- Strong presentation and facilitation skills, with the ability to communicate clearly and confidently with technical teams, senior leaders, executives, vendors, and cross-functional stakeholders.
- Ability to guide technical discussions, evaluate trade-offs, balance reliability, security, compliance, delivery speed, and operational efficiency, and help teams make sound engineering decisions.
- Experience improving operational processes, reducing manual toil, mentoring engineers, managing delivery, and building healthy, accountable teams.
- Experience with Java-based services, build systems, or code review workflows is a plus.
- Comfort using or encouraging AI-assisted engineering tools responsibly within company guidelines.
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. A Master’s degree is a plus.

