We are seeking a highly skilled and experienced Senior DevOps Engineer to join McKesson Speciality Health Technology Products LLC. This role is heavily focused on application support, production stability, and operational excellence in a fast-paced healthcare environment. The ideal candidate will play a critical role in supporting and optimising CI/CD pipelines, cloud infrastructure, and application platforms, ensuring high availability, reliability, and performance across multiple environments. You will work closely with Development, QA, Database, and SRE teams to support the full software development lifecycle (SDLC), production operations, and release management. This role requires strong experience in production support, incident management, and proactive monitoring, along with a clear understanding of the criticality of healthcare production systems.
The core responsibilities for the job include the following:
Application and Production Support:
- Provide L2/L3 support for applications running in Azure AKS Kubernetes clusters.
- Actively monitor production systems using Dynatrace, Azure Monitor, and other observability tools.
- Handle on-call support, respond to incidents, perform troubleshooting, and drive root cause analysis (RCA).
- Ensure system stability, performance, and quick resolution of production issues.
CI/CD and Release Management:
- Design, maintain, and support CI/CD pipelines using GitHub Actions, Jenkins, and Azure DevOps.
- Support release deployments across environments (DEV, QA, PROD).
- Collaborate with teams to improve release cycles, deployment reliability, and rollback strategies.
- Ensure smooth coordination during releases with Dev, QA, and DB teams.
Cloud and Platform Operations:
- Manage and support Azure cloud infrastructure, including AKS, networking, and storage.
- Work with Kubernetes (AKS) for application deployment, scaling, and troubleshooting.
- Support service mesh (Istio) configurations and traffic management.
Automation and Infrastructure:
- Implement automation using scripting (Python, Bash, PowerShell, and Terraform) to improve operational efficiency.
- Support Infrastructure as Code (IaC) practices where applicable.
- Automating repetitive task through Agentic AI
Monitoring, Logging and Security:
- Utilise tools like Dynatrace, Azure Monitor, and logging platforms to ensure observability.
- Integrate and support security and code quality tools such as Wiz CLI and SonarQube.
- Proactively identify potential risks and ensure compliance with enterprise security standards.
Cross-Team Collaboration:
- Work closely with Development, QA, Database, and SRE teams to resolve issues and improve system performance.
- Support troubleshooting across application layers, including APIs, services, and databases.
- Participate in troubleshooting war rooms and critical incident calls.
ServiceNow and Operational Excellence:
- Manage and track incidents, changes, and service requests using ServiceNow.
- Ensure proper documentation, ticket updates, and adherence to SLA/OLA expectations.
- Contribute to runbooks, SOPs, and knowledge base documentation.
Continuous Improvement and Mentorship:
- Mentor junior engineers and improve overall DevOps and operational practices.
- Identify automation and optimisation opportunities to reduce manual effort.
- Drive adoption of DevOps best practices with a focus on reliability and supportability.
Requirements:
- Bachelor's degree in computer science, engineering, or a related field (or equivalent experience).
- 3+ years of experience in DevOps, SRE, or application support roles.
- Strong experience in production support and incident management.
- Strong scripting skills (Python, Bash).
- Experience supporting full SDLC, releases, and multi-environment deployments.
- Ability to troubleshoot complex issues across applications, infrastructure, and databases.
- Strong understanding of production criticality, uptime requirements, and healthcare systems' sensitivity.
- Excellent communication skills and ability to work across multiple teams.
- Willingness to participate in on-call rotation and support critical systems.
Hands-on experience with:
- CI/CD tools: GitHub Actions, Jenkins, Azure DevOps.
- Cloud: Azure (preferred).
- Containers and Orchestration: Docker, Kubernetes (AKS).
- Service Mesh: Istio.
- Monitoring: Dynatrace, Azure Monitor.
- Security/Quality Tools: Wiz CLI, SonarQube.
- ITSM: ServiceNow.

