368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$112k – $218k per year (Estimated)
Location
Remote/Hybrid (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in the United States.

This is an impactful opportunity for an experienced SRE to strengthen the reliability, scalability, security, and operational efficiency of modern cloud platforms.

You will take ownership of production systems across AWS and Kubernetes while establishing measurable reliability standards.

The role combines software engineering, infrastructure automation, observability, incident management, and resilience engineering.

You will work closely with software, platform, security, QA, and product teams to make systems safer and easier to operate at scale.

A major focus will be reducing operational toil through automation, improving deployment reliability, and turning production incidents into lasting improvements.

You will also contribute to responsible cloud cost management while ensuring performance, availability, and security remain strong.

This role is well suited to a technically strong engineer who enjoys solving complex problems and shaping engineering practices for mission-critical systems.

Accountabilities

    • Define, measure, and continuously improve service reliability through SLIs, SLOs, error budgets, availability targets, and capacity planning.
    • Build and maintain comprehensive observability across infrastructure and applications, including metrics, logs, traces, dashboards, and actionable alerting using tools such as CloudWatch, Prometheus, Grafana, New Relic, ELK, or OpenSearch.
    • Participate in production incident response, troubleshooting, escalation, root-cause analysis, blameless post-incident reviews, and corrective-action tracking.
    • Improve operational excellence through production readiness reviews, runbooks, documentation, change management practices, and engineering standards that reduce risk and improve maintainability.
    • Optimize infrastructure for reliability and performance while managing cloud consumption responsibly and supporting FinOps initiatives.
    • Analyze system performance, resource utilization, latency, throughput, and growth trends to identify bottlenecks and implement scalable solutions proactively.
    • Identify repetitive operational tasks and replace them with reliable automation using Python, Bash, Go, CI/CD tooling, and platform APIs.
    • Design and improve CI/CD pipelines that enable safe, repeatable deployments through automated testing, validation, progressive delivery, rollback strategies, and deployment observability.
    • Apply security and compliance controls across cloud environments, including least-privilege IAM, encryption, network security, secrets management, patching, vulnerability management, and auditability.
    • Design, develop, review, and maintain reusable Terraform modules and infrastructure-as-code patterns for AWS environments.
    • Operate and improve Kubernetes platforms, including EKS clusters, workloads, Helm deployments, autoscaling, upgrades, resource management, networking, storage, and workload resilience.
    • Architect, operate, and optimize AWS services including EC2, S3, RDS, EKS, Lambda, VPC, IAM, Route 53, and related cloud services.
    • Design and validate fault-tolerant architectures, backup strategies, recovery procedures, and disaster recovery capabilities, including reliability testing and failure exercises where appropriate.
    • Partner with development teams to improve application operability, instrumentation, deployment patterns, reliability, and production readiness.
    • Contribute to incident response programs, on-call practices, resilience testing, operational standards, and continuous improvement initiatives across engineering teams.
    • Requirements

      • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, cloud infrastructure, or a closely related discipline, with significant production ownership.
      • Strong hands-on experience designing and operating production workloads in AWS, including networking, IAM, compute, storage, databases, DNS, and managed Kubernetes.
      • Advanced Terraform expertise, including reusable modules, remote state, dependency management, environment design, code reviews, and infrastructure lifecycle management.
      • Strong experience with observability and production telemetry, using technologies such as Prometheus, Grafana, CloudWatch, New Relic, ELK, OpenSearch, or comparable platforms.
      • Practical understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, fault tolerance, graceful degradation, and operational toil reduction.
      • Deep Kubernetes knowledge covering EKS, Helm, workload scheduling, networking, storage, autoscaling, upgrades, troubleshooting, and production operations.
      • Strong Linux systems expertise, with the ability to diagnose issues involving CPU, memory, disk, networking, processes, DNS, and application dependencies.
      • Proficiency in Python, Bash, Go, or another general-purpose programming language used for operational tooling and automation.
      • Experience troubleshooting complex production incidents and contributing to incident response, root-cause analysis, postmortems, and corrective actions.
      • Experience designing or operating CI/CD systems such as GitHub Actions, Jenkins, GitLab CI, Argo CD, or comparable technologies.
      • Working knowledge of cloud security practices, including IAM, encryption, secrets management, network segmentation, vulnerability management, and audit controls.
      • Strong written and verbal communication skills, with the ability to collaborate effectively across technical and business teams.
      • Ability to work independently, take ownership of production systems, and make sound technical decisions in complex environments.
      • AWS certifications such as Solutions Architect Professional or DevOps Engineer Professional are desirable.
      • Experience with GitOps, Argo CD or Flux, formal on-call programs, chaos engineering, resilience testing, service meshes, distributed systems, microservices, RDS, DynamoDB, PostgreSQL, multi-account AWS environments, cloud governance, FinOps, or compliance frameworks such as SOC 2, ISO 27001, or PCI DSS is a plus.
      • Benefits

        • Flexible remote or hybrid work options.
        • Comprehensive health and wellness benefits, including medical, dental, and vision coverage with 100% premium coverage for you.
        • Generous paid time off and paid holidays.
        • MyShare Employee Ownership Program.
        • 401(k) plan with up to a 4% employer match, with full vesting from day one.
        • Opportunities to work alongside industry leaders and experienced engineering professionals.
        • Professional growth opportunities in modern cloud, reliability, and energy technology environments.
        • Opportunity to contribute to advanced energy intelligence and infrastructure supporting a more modern, secure, and resilient grid.
        • Occasional travel of 10% or less may be required.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$28k – $66k per year (Estimated) • In office • Internship • 11+ years exp • Bachelor's Degree • Gurgaon
SQL
Databases
Amazon Redshift
Apache Kafka
Databricks
Google BigQuery
MySQL
PostgreSQL
Snowflake
AI/ML
Airflow
Hadoop
Spark
DevOps
AWS
AWS Lambda
CI/CD
Docker
GCP
Kubernetes
Amazon Kinesis
Amazon S3
Analytics
ETL/ELT
Apply
Staff Data Engineer 3 days ago
$213k – $255k per year • Equity • Remote • Full-Time • 5+ years exp
SQL
Databases
Apache Kafka
AI/ML
AI Agents
DevOps
Amazon EKS
CI/CD
Platform Engineering
AWS
Kubernetes
Cybersecurity
GDPR
Design
Webflow
Apply
$184k – $350k per year (Estimated) • Remote/Hybrid • Internship • 15+ years exp • Bachelor's Degree • Arlington
AI/ML
Computer Vision
AI Agents
DevOps
AWS
Platform Engineering
Cybersecurity
GDPR
Apply
$173k – $369k per year (Estimated) • Remote/Hybrid • Internship • 15+ years exp • Bachelor's Degree • San Francisco
AI/ML
Computer Vision
AI Agents
DevOps
AWS
Platform Engineering
Cybersecurity
GDPR
Apply
$148k – $286k per year (Estimated) • Remote/Hybrid • Contractor • 10+ years exp • Arlington
Python
AI/ML
Computer Vision
Embeddings
PyTorch
Time Series Forecasting
DevOps
AWS
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.