Salary
≈ $111k – $215k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NextGen Healthcare is an American healthcare technology company headquartered in Atlanta, Georgia, and founded in 1974. The firm develops cloud-based electronic health records, practice management software, and revenue cycle management tools, alongside interoperability platforms like Mirth Connect. It supports more than 150,000 healthcare providers across various medical specialties and operates as a private company under the ownership of Thoma Bravo.
Job Description:
The Senior Cloud Operations Reliability Engineer is responsible for driving operational excellence and strengthening the reliability posture of cloud-based services and supported platforms. This role owns critical reliability initiatives, establishes observability and service health practices, and leads incident response coordination to improve service availability, resiliency, and recovery. The Senior Cloud Operations Reliability Engineer partners with engineering, security, and operations teams to advance reliability practices, support and mature service-level objectives, improve production readiness, and develop reliability-focused automation that reduces operational toil and accelerates incident recovery.- Own service reliability and operational health-establish and maintain SLOs/SLIs, design monitoring and alerting strategies, and drive improvements that enhance service availability and performance across cloud platforms.
- Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root cause analysis, and driving remediation activities with accountability for timeline and resolution quality.
- Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness and resilience; apply Infrastructure as Code practices where appropriate to support recovery, reliability, and operational consistency.
- Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures to ensure rapid problem detection and response.
- Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks; provide recommendations for reliability-focused scaling, performance improvement, capacity planning, and operational readiness of cloud-based services.
- Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability, rollback readiness, and production supportability.
- Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and ensure recovery procedures and documentation reflect current production state and evolving business requirements.
- Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; create and maintain operational documentation, standard operating procedures, and knowledge base materials.
- Support operational adherence to cloud governance, compliance, and security initiatives; including access control, tagging, logging, and audit readiness, and reliability-related documentation.
- Perform other duties that support the overall objective of the position.
Education Required:
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
- Or, any combination of education and experience which would provide the required qualifications for the position.
Experience Required:
- 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated ownership of production systems.
- Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers.
- Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
- Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
- Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
- Experience with Kubernetes operations, containerization, and orchestration platforms.
- Experience with application performance monitoring (APM) and distributed tracing.
- Experience mentoring junior engineers or leading operational improvements initiatives.
License/Certification Required:
- Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer.
- AWS certification: AWS SysOps Administrator or equivalent.
- Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL.
Knowledge, Skills & Abilities:
- Knowledge of: Working knowledge of CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning. Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers; understanding of cloud-native services, networking, security, and compute models.Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, or similar platforms. Security operations, compliance auditing, or audit readiness processes, preferred.
- Skill in: Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.); ability to design effective dashboards, alerts, and health checks. Proficiency in scripting languages (Python, Bash, Go, etc.) to develop automation solutions that reduce manual effort.
- Ability to: Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery. Ability to translate complex technical findings into actionable recommendations; experience influencing cross-functional teams on reliability practices.
The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.
NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
United States
AWS Application Developer lead Engineer
4 hours ago
≈ $25k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Chennai
Java
PowerShell
Python
TypeScript
JavaScript
Java
Spring Boot
Databases
Amazon Redshift
DynamoDB
Frontend
Angular
DevOps
Amazon EC2
Amazon EKS
Ansible
ArgoCD
AWS
AWS Lambda
Azure
Azure DevOps
Bicep
Chef
CI/CD
CloudFormation
Configuration Management
Docker
GCP
GitLab CI
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
Amazon S3
GitLab
IAM
Apply
DevOps Engineer Sênior| Inglês Avançado
4 hours ago
≈ $45k – $114k per year (Estimated) • Remote • Full-Time • São Paulo • Vitoria-Gasteiz • Fortaleza • Rio de Janeiro • Recife
DevOps
AWS
Azure
Azure DevOps
Bicep
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
GitLab
GitLab CI
IAM
Incident Management
Jenkins
Kubernetes
Terraform
Apply
Technical Evangelist Director - Informatica
4 hours ago
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
IT Administrator
4 hours ago
$83k – $113k per year • In office • Full-Time
Node JS
Python
JavaScript
DevOps
AWS
GCP
GitHub
Cybersecurity
Okta
SOC 2
Management
Confluence
Jira
Slack
Apply
DevOps Engineer
4 hours ago
≈ $25k – $62k per year (Estimated) • In office • Full-Time • 8+ years exp • Coimbatore
DevOps
Azure
Azure DevOps
CI/CD
Apply
Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
C#
JavaScript
Python
SQL
C#
.NET
AI/ML
ChatGPT
Claude
Claude Code
Gemini
Prompt Engineering
Frontend
React.js
DevOps
Self-Healing
Apply
Product Manager, NextGen Office (NGO)
13 days ago
≈ $130k – $222k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree • United States
Apply
Deal Desk, Senior Analyst II
15 days ago
≈ $84k – $167k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
Marketing
Salesforce
Apply
Specialist II, Cloud Platform & Operations
20 days ago
≈ $16k – $48k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bengaluru
Python
DevOps
AIOps
AWS
CI/CD
Docker
Dynatrace
GCP
Incident Management
Kubernetes
New Relic
Platform Engineering
Prometheus
SRE
Terraform
IAM
Cybersecurity
Auth0
Okta
Apply
Senior Data Platform Engineer
27 days ago
≈ $123k – $217k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • United States
Python
SQL
Databases
Snowflake
DevOps
AWS
CI/CD
Amazon S3
IAM
Analytics
Power BI
ETL/ELT
Apply
Mech and Robotics Tech
1 hour ago
$60k – $61k per year • In office • Full-Time • 2+ years exp • High School Diploma • United States
Apply
Mech and Robotics Tech
1 hour ago
$62k – $63k per year • In office • Full-Time • 2+ years exp • High School Diploma • United States
Apply
SAP Business Data Cloud (BDC) Consultant
1 hour ago
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
Data Technical Analyst Lead HCV 6205536
1 hour ago
$120k – $140k per year • In office • Full-Time • Charlotte • Raleigh • Dallas • Boston • New York
SQL
Analytics
ETL/ELT
Apply
SAP SLO (DMLT, SNP, SLT) Consultant
1 hour ago
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
This is one of many
368,611 more open roles from verified company boards, updated every day.

