603,101open jobs
31,403companies
86,475added this week
Browse all
Salary
$175k – $200k per year
Location
Remote (United States)
Seniority
Architect · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About Intus Care:

IntusCare is a leader in innovative, data-driven healthcare solutions focused on enabling value-based care organizations. We are building a modern, cloud-native EMR platform purpose-built for PACE & value-based care organizations, designed from the ground up around scalability, interoperability, analytics, and operational efficiency. Data sits at the center of everything we do - from clinical decision support and quality outcomes reporting to payer analytics and population health insights.

Role Overview:

As Director of Site Reliability Engineering, you will own the reliability, availability, and operational excellence of Intus Care's entire product and platform portfolio. You will lead a blended organization that includes a managed SRE services team, an internal QA team, and a growing internal SRE capability - with full accountability for outcomes across all three. You will define the SRE strategy and roadmap, establish SLA and SLO frameworks across all products, design and own Intus Care's incident management process, and build the observability and release gate infrastructure that engineering teams depend on to ship safely and confidently.

This role requires a leader who combines deep SRE and infrastructure expertise with strong operational management skills and the ability to influence engineering culture across multiple teams. You will partner closely with Engineering, Security, and Product leadership, and serve as the primary voice of reliability in cross-functional forums. This is a rare opportunity to build a reliability function from the ground up in a high-stakes healthcare environment where the work directly impacts patient care.

Key Responsibilities:

· Own and execute the SRE strategy and multi-quarter roadmap across reliability, observability, incident management, QA maturity, and release engineering.

· Define, measure, and continuously improve SLAs, SLOs, error budgets, uptime, performance, and operational health metrics across all products and services.

· Lead production reliability for the full platform, including monitoring, alerting, on-call operations, incident response, root cause analysis, and MTTR reduction.

· Establish release readiness standards, deployment safety controls, and quality gates to ensure stable and predictable product releases.

· Manage external SRE vendors and partners, including service delivery, SLA governance, escalations, performance reviews, and compliance expectations.

· Lead QA engineering strategy with a focus on automation, regression prevention, test coverage, and reducing escaped defects in production.

· Partner with Security and Engineering leaders to ensure cloud infrastructure, CI/CD pipelines, and operational tooling meet HIPAA, SOC2, and internal security standards.

· Oversee core platform operations including Azure AKS environments, Kubernetes, GitOps workflows, CI/CD pipelines, GitHub Actions, secrets management, access controls, and audit readiness.

· Drive observability maturity using tools such as Grafana, Prometheus, logging platforms, tracing tools, and automated alerting frameworks.

· Collaborate with Product, Platform, and Engineering teams to embed reliability and quality best practices throughout the software development lifecycle.

· Build, mentor, and scale high-performing SRE and QA teams while fostering a culture of ownership, accountability, learning, and continuous improvement.

· Drive adoption of AI-enabled automation and intelligent tooling to reduce manual toil, improve productivity, and strengthen operational excellence.

Technical Experience

· Strong hands-on experience with cloud infrastructure, preferably Microsoft Azure, including AKS, networking, storage, IAM, and security services.

· Deep expertise in Kubernetes, containerized workloads, and production-scale distributed systems.

· Experience building and managing CI/CD pipelines using GitHub Actions, ArgoCD, Terraform, or similar DevOps tooling.

· Strong background in monitoring, logging, tracing, and observability platforms such as Grafana, Prometheus, Datadog, Splunk, or equivalent.

· Experience with scripting and automation using Python, Bash, PowerShell, or similar languages.

· Strong understanding of release engineering, automated testing frameworks, QA tooling, and shift-left quality practices.

· Experience supporting SaaS applications with uptime, scalability, and security requirements in regulated industries such as healthcare.

· Knowledge of HIPAA, SOC2, vulnerability management, access controls, and infrastructure security best practices.

· Familiarity with databases, APIs, networking, and troubleshooting across modern web application stacks.

· Exposure to AI-powered DevOps / AIOps tooling for incident management, automation, and engineering productivity is a plus.

Requirements

· 12+ years of SRE, infrastructure, or platform engineering experience, with 5+ years of engineering leadership roles.

· Proven track record owning site reliability for complex, multi-tenant SaaS platforms with demanding availability requirements.

· Demonstrated experience defining SLA and SLO frameworks, error budgets, and incident management processes at scale.

· Experience managing vendor relationships for managed infrastructure or SRE services, including SLA governance and performance management.

· Track record leading QA or quality engineering functions, including test automation maturity and release gate ownership.

· Strong communication and cross-functional influence skills - able to represent reliability to both technical and non-technical audiences

Preferred Qualifications

  • Experience in healthcare technology, HIPAA-compliant environments, or other highly regulated SaaS industries.
  • Familiarity with FHIR-native or EMR/EHR platform architectures and their specific reliability requirements.
  • Experience implementing AI-assisted SRE automation including runbook generation, anomaly detection, or incident triage tooling.
  • Background working with Playwright or equivalent test automation frameworks in a QA leadership capacity.
  • Experience building internal SRE capability alongside a managed services provider

Why Join Intus Care?

  • Own and build the SRE function for a modern healthcare EMR platform serving PACE populations - from the ground up.
  • Lead a blended team model combining managed services, internal QA, and internal SRE in a high-growth engineering organization.
  • Work on systems where reliability directly impacts clinical care delivery for vulnerable patient populations.
  • Shape engineering culture in a company that actively embraces AI-assisted software development with Claude Code.
  • Fully remote, collaborative engineering environment with direct access to executive leadership.

Compensation:

The base salary range for this role is $175k- 200k. We expect the ideal candidate to fall near the midpoint of this range, though final compensation will be determined based on experience, skills, and organizational needs. Final compensation will also include a variable component and stock options.

Work location: This is a fully remote role based in the United States.

Sponsorship: This position is not eligible for sponsorship.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
603,101 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$135k – $153k per year • In office • 5+ years exp • Bachelor's Degree
Python
SQL
AI/ML
AI Agents
LLM Guardrails
DevOps
Terraform
CloudFormation
CI/CD
AWS
Management
Agile
Apply
$60k – $109k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
.NET Developer 1 day ago
$26k – $61k per year (Estimated) • In office • 5+ years exp • Bengaluru
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
Frontend
Redux
Webpack
React.js
Vite
Mobile
React Native
Clean Architecture
Dependency Injection
State Management
DevOps
Rest API
Azure DevOps
GitHub Actions
OpenTelemetry
Prometheus
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Management
Agile
Apply
$63k – $160k per year (Estimated) • In office • Internship • Master's Degree
Python
C++
Apply
Design Engineer II 1 day ago
$23k – $49k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Nanjing
Python
C++
SystemVerilog
Chips/EDA
Cadence Palladium
Apply
Tech Lead 2 days ago
$160k – $190k per year • Remote • Full-Time
TypeScript
Apply
$48k – $56k per year • Remote • Contractor
Management
Outlook
Apply
$131k – $253k per year (Estimated) • Remote • Full-Time • 5+ years exp
SQL
AI/ML
Claude
Analytics
Tableau
Looker
Microsoft Excel
Management
Smartsheet
Notion
Airtable
Marketing
HubSpot
Apply
$166k – $287k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
SQL
Databases
PostgreSQL
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Claude Code
dbt
DevOps
Loki
Prometheus
Kubernetes
Grafana
Cybersecurity
HIPAA
Apply
See all jobs
This is one of many
603,101 more open roles from verified company boards, updated every day.