368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$241k – $270k per year
Location
Remote (United States)
Seniority
Staff · 7+ years exp
Overview
Company
Impact
Profile match
Garner Health is a healthcare technology company based in New York City and founded in 2019. The company provides a data-driven platform that identifies high-quality medical providers and offers an employer-sponsored benefit that reimburses employees for out-of-pocket costs when they select top-rated doctors. It operates throughout the United States, serving nearly 800 organizations ranging from small businesses to Fortune 50 companies to improve healthcare outcomes and reduce overall spending.

What you’ll be part of

Garner is on a mission to transform the U.S. healthcare system - and we’re the only proven player doing exactly that. We partner with employers to redesign how healthcare works: applying 550+ proprietary clinical metrics across 80+ specialties to a dataset of 320M+ patients to identify the best-performing doctors, then using compelling incentives to steer members to the care that helps them get healthier, faster.

The result is a rare “win win” - better care and lower costs for both members and employers. In just five years, our work has helped over 2.5 million people access higher-quality care and saved $1B in healthcare costs. We recently raised our Series E and have doubled five years running. If you've ever wanted your work to solve a problem that touches every person in this country, this is the opportunity to do exactly that. You'd be joining a team fundamentally reimagining healthcare in the U.S. - and using AI to scale that impact further and faster than anyone else can.

About the role:

We are seeking an exceptional Staff Site Reliability Engineer to own the reliability strategy for the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. As the most senior reliability voice in the organization, you will set the technical direction for how Garner defines, measures, and upholds production quality, architecting the SLO framework, incident response program, and automation standards that every engineering team builds on. Because our systems directly influence health outcomes for millions of patients, maintaining the highest standards of production quality is imperative. This is an automation-first role: you will use AI tools to continuously convert manual operational work into monitored, hands-free processes, and build the platform that lets every Garner engineer do the same.

Where you will work:

Garner is headquartered in NYC, but this position is available for individuals who are comfortable with remote work and occasional travel to HQ.

What you will do:

  • Own the Reliability Strategy: Architect and own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; design the SLO framework our critical services are measured against and lead the technical decision-making that keeps us ahead of scale
  • Lead the Incident Response Program: Set the standard for how Garner responds to incidents: serve in and level up the on-call rotation, lead response for the most complex escalations, drive deep-dive root cause analysis, and build the review culture that sees corrective actions through to resolution
  • Own Observability: Architect the monitoring, alerting, and observability platform that lets us detect and resolve issues before users feel them, and that lets stakeholders quickly identify the health of every team’s products
  • Translate Ambiguity: Take high-level, ambiguous scaling and reliability requirements and transform them into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
  • Course-Correct Technical Direction: Proactively identify when infrastructure workflows or technical paths are inefficient or fragile and redirect efforts to ensure the highest ROI for the engineering function, paying down impactful tech debt and using AI tools and automation to convert repetitive operational work into hands-free, monitored processes
  • Multiply the Engineering Team: Build and own the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; mentor engineers across the organization and provide high-quality feedback that raises the bar for operational rigor and discipline
  • Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner’s security and HIPAA compliance obligations, and lead rigorous review of infrastructure changes so platform work meets the same standards as our customer-facing products

The ideal candidate has:

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred), with a track record of architecting reliability for systems at scale
  • Experience designing an organization’s reliability practice (SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews) and the judgment to know when to build vs. buy
  • Strong Python or Go skills applied to infrastructure automation (Kubernetes API experience a plus)
  • Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking
  • Mentorship experience and the ability to set technical direction as the senior reliability voice
  • Excellent communication skills-able to make complex reliability concepts land with both technical and non-technical stakeholders
  • Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback

Technologies we use: 

  • AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab

This is a unique opportunity to join a fast-growing company in a transformative role, helping shape the future of healthcare.

Compensation Transparency:

The target salary range for this position is $241,000 - $270,000. Individual compensation for this role will depend on various factors, including qualifications, skills, and applicable laws. In addition to base compensation, this role is eligible to participate in our equity incentive and competitive benefits plans, including but not limited to: flexible PTO, Medical/Dental/Vision plan options, 401(k), Teladoc Health and more.

Fraud and Security Notice: 

Please be aware of recent job scam attempts. Our recruiters use getgarner.com and garnerhealth.com email domains exclusively. If you have been contacted by someone claiming to be a Garner recruiter or a hiring manager from a different domain about a potential job, please report it to law enforcement here and to [email protected].

Equal Employment Opportunity:

Garner Health is proud to be an Equal Employment Opportunity employer and values diversity in the workplace. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics.

Garner Health is committed to providing accommodations for qualified individuals with disabilities in our recruiting process. If you need assistance or an accommodation due to a disability, you may contact us at [email protected]

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$61k – $200k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$145k – $290k per year (Estimated) • In office • 2+ years exp • New York
Python
TypeScript
AI/ML
Claude
Claude Code
dbt
DevOps
AWS
Kubernetes
Cybersecurity
HIPAA
Apply
$211k – $400k per year (Estimated) • Equity • In office • 8+ years exp • New York
Apply
Data Analyst I 7 days ago
$88k – $110k per year • Equity • In office • 1+ year exp • New York
SQL
Databases
Snowflake
AI/ML
AI Agents
dbt
Apply
Product Manager III 14 days ago
$180k – $205k per year • Equity • In office • 3+ years exp • New York
Apply
$149k – $160k per year • Equity • In office • 3+ years exp • New York
SQL
Databases
Snowflake
AI/ML
AI Agents
dbt
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.