371,660open jobs
9,621companies
49,388added this week
Browse all
Salary
$103k – $216k per year (Estimated)
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Transform your pharmacy and healthcare operations with our AI-powered workflow automation platform. Optimize back-office tasks and administrative work, maximizing operational efficiency. Improve productivity and revenue capture for your healthcare organization today.

About Plenful

Plenful is on a mission to transform healthcare operations from the inside out. Fresh off our $50M Series B and backed by Notable Capital, Bessemer Venture Partners, TQ Ventures, Susa/Kivu Ventures, and other leading investors, we’re building the category-defining AI workflow automation platform that healthcare teams rely on to operate smarter, faster, and more efficiently. Our technology empowers healthcare operators across hospital and health systems, pharmacies and payors to eliminate manual work, reduce administrative burden, and improve compliance, all while unlocking critical revenue to fund programs for their in-need patient populations.

Built by healthcare operators for healthcare operators, Plenful is driven by a deep understanding of the challenges facing today’s care teams. We’re passionate about equipping healthcare workers with world-class tools that deliver real, measurable impact, and we’re proud to serve 90+ leading health systems across the country. If you’re excited to help shape the future of healthcare, we’d love to meet you. Apply now to join our growing team.

About the Role

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.

This role is centered on operating real systems at scale - not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.

You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving - from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

What You’ll Do

Reliability Engineering & System Ownership

  • Define and implement SLIs, SLOs, and error budgets across core services.

  • Own production system health: uptime, latency, and availability targets.

  • Improve system resilience through proactive reliability work.

  • Find and mitigate single points of failure across distributed systems.

Production Operations & Incident Response

  • Take part in and improve on-call rotations and incident response.

  • Lead incident triage, mitigation, and resolution in real time.

  • Run blameless postmortems and follow through on action items.

  • Build tooling and automation to cut MTTR (Mean Time to Recovery).

Observability & System Insight

  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.

  • Improve signal quality to cut noise and alert fatigue.

  • Build dashboards and alerts that reflect real system health and user impact.

  • Use observability data to drive performance and reliability improvements.

Performance & Scalability

  • Analyze system performance under load and find bottlenecks.

  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).

  • Partner with engineering teams to improve system efficiency and scaling behavior.

Automation & Reliability Tooling

  • Build automation that eliminates repetitive operational work.

  • Improve deployment safety through reliability checks and safeguards.

  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.

  • Build tools for incident response, debugging, and capacity planning.

Security, Compliance & Operational Maturity

  • Partner with security and compliance to keep systems meeting operational standards.

  • Support audit readiness and reliability-related compliance requirements (Vanta).

  • Integrate monitoring and alerting into security and SIEM workflows.

  • Help mature operational practices across engineering.

You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.

You May Be a Fit If

  • You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.

  • You've operated and debugged distributed systems in production.

  • You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.

  • You've defined and worked with SLOs, SLIs, and error budgets.

  • You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.

  • You can write code or scripts (Python, Bash, etc.) for automation and tooling.

  • You think in systems and reason clearly about failure modes.

  • Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.

Why You'll Love Working Here

  • Mission-Driven, World-Class Team - Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact

  • Opportunities for Growth - Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization

  • Flexible Hybrid Work Environment - We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office

Benefits & Perks

  • Healthcare Coverage - Full medical, dental, and vision insurance for you and participation for your family

  • 401(k) with Company Match - Plenful matches 50% of your first 3% contributed

  • Equity - Every full-time employee shares in our success

  • Unlimited PTO - Take the time you need, when you need it

  • Daily Lunch Stipend - $100/week to cover your midday meals

  • Wellness Stipend - $100/month to support your health and well-being

  • Commuter Benefits - $100/month for SF and NYC-based employees

  • Parental Leave - Paid leave to support growing families

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
371,660 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$130k – $180k per year • Remote • Full-Time • 2+ years exp • Bachelor's Degree
Go
Node JS
SQL
TypeScript
JavaScript
Databases
Amazon Redshift
ClickHouse
Databricks
DynamoDB
ElasticSearch
Google BigQuery
Redis
Snowflake
AI/ML
dbt
DevOps
AWS
CI/CD
CloudFormation
Datadog
Docker
Git
Kubernetes
Terraform
Apply
Full Stack Engineer 9 hours ago
$64k – $150k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree
JavaScript
Python
TypeScript
Python
Django
Django REST Framework
Frontend
React.js
Redux
DevOps
AWS
Azure
CI/CD
GCP
QA
Jest
Pytest
Apply
DevOps Lead (FedRAMP) 9 hours ago
$230k – $260k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree
Python
DevOps
Amazon EKS
AWS
Azure
CI/CD
Configuration Management
GCP
Helm
Kubernetes
Terraform
Cybersecurity
FedRAMP
Apply
$139k – $274k per year (Estimated) • Remote • Full-Time
Go
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Function Calling
Model Context Protocol
DevOps
AWS
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Apply
$167k – $277k per year (Estimated) • Remote • Full-Time
Go
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Function Calling
Model Context Protocol
DevOps
AWS
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Apply
$90k – $197k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco
C#
JavaScript
Node JS
Python
SQL
TypeScript
Databases
PostgreSQL
AI/ML
LLM
Frontend
React.js
DevOps
AWS
CI/CD
Datadog
Docker
Git
GitHub Actions
Grafana
Kubernetes
Rest API
GitHub
QA
Cypress
Playwright
Selenium
Apply
Head of Security 2 months ago
$148k – $316k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • San Francisco
DevOps
AWS
CI/CD
Kubernetes
IAM
Cybersecurity
HIPAA
SOC 2
Threat Modeling
Apply
$136k – $284k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
SQL
Databases
Pinecone
PostgreSQL
Qdrant
Weaviate
AI/ML
BentoML
Braintrust
Embeddings
Kubeflow
Langfuse
LangSmith
LLM
MLFlow
NLP
Prompt Engineering
PyTorch
RAG
Ragas
Scikit-learn
Semantic Search
TensorFlow
vLLM
Weights & Biases
Amazon SageMaker
Anthropic
LLMOps
OpenAI
Semantic Search
AI Agents
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
Rest API
Vector
GitHub
Apply
Senior DevOps Engineer 2 months ago
$104k – $217k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • San Francisco
Databases
PostgreSQL
DevOps
AWS
CI/CD
Terraform
Apply
$196k – $320k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • San Francisco
JavaScript
Python
Ruby
TypeScript
Apply
$100k – $150k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
DevOps
GitHub
Apply
$100k – $150k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
DevOps
GitHub
Apply
$106k – $215k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • San Francisco
Databases
Snowflake
AI/ML
AI Agents
Claude
OpenAI
DevOps
CI/CD
Incident Management
GitHub
Management
Linear
Notion
Marketing
Zendesk
Apply
$120k – $204k per year • In office • Internship • San Francisco
AI/ML
Post-training
RLHF
Apply
$140k – $180k per year • Equity 0.1–1% • In office • Full-Time • 3+ years exp • San Francisco
Python
TypeScript
JavaScript
Frontend
Next.js
React.js
Apply
See all jobs
This is one of many
371,660 more open roles from verified company boards, updated every day.