609,983open jobs
32,963companies
86,576added this week
Browse all
Salary
$104k – $163k per year
Location
Remote (United States)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About the Role

We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.

Key Responsibilities

  • Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently

  • Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior

  • Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems

  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity

  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives

  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure

  • Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments

  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows

  • Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment

Required Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems

  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale

  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)

  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization

  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes

  • Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows

  • Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design

  • Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure

  • Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers

Preferred Qualifications

  • Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)

  • Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management

  • Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically

  • Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms

  • Contribution to open-source observability or infrastructure projects

  • Expertise in network security, application security, or infrastructure hardening

  • Experience with database optimization, query performance tuning, and backup/recovery strategies

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
609,983 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$24k – $48k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Moscow
Python
SQL
PowerShell
Bash
DevOps
Zabbix
Prometheus
Grafana
SLI/SLO/SLA
Management
Jira
ServiceNow
Apply
$29k – $68k per year (Estimated) • Remote • Bachelor's Degree • Moscow
Python
Databases
ElasticSearch
AI/ML
AI Agents
NER
LLM
BERT
DevOps
Docker Compose
HAProxy
Docker
Ubuntu
Nginx
CentOS Stream
Apply
$23k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Pune
Python
Java
SQL
Java
Spring Framework
Maven
DevOps
OpenShift
Git
AWS
Management
Agile
Apply
$31k – $74k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
AI/ML
LangChain
Reinforcement Learning
AI Agents
TensorFlow
PyTorch
RAG
Hallucination
DevOps
Kubernetes
Apply
$13k – $28k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bengaluru
Python
SQL
Python
pySpark
AI/ML
Hadoop
Spark
Analytics
Microsoft Excel
Apply
$77k – $95k per year • Remote • Full-Time • 3+ years exp
AI/ML
Copilot
Claude
Management
Smartsheet
Apply
$77k – $121k per year • Remote • Full-Time • 2+ years exp • Bachelor's Degree
SQL
DevOps
Azure
Apply
Software Engineer II 26 days ago
$93k – $147k per year • Remote • Full-Time • 3+ years exp
C#
C#
.NET
DevOps
Azure
CI/CD
AWS
Docker
Apply
$69k – $100k per year • Remote • Full-Time • 2+ years exp
SQL
Analytics
Tableau
Power BI
Looker
Microsoft Excel
Marketing
Salesforce
Apply
$104k – $163k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
TypeScript
Databases
PostgreSQL
DevOps
Terraform
GitHub Actions
Pulumi
AWS
Docker
AWS Lambda
Amazon S3
Amazon CloudWatch
Amazon EventBridge
Management
Slack
Apply
See all jobs
This is one of many
609,983 more open roles from verified company boards, updated every day.