368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$25k – $63k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 8+ years exp
Employment
Internship
Overview
Company
Impact
Profile match
This is the jump-off point to learn about working at Roku, including our employees, the Roku culture, our office locations around the world, and student internship opportunities. Roku, jobs, careers, internship, streaming, TV, engineering.

We are seeking a talented and experienced SRE (Site Reliability Engineering) Senior Software Engineer to join our dynamic team. The ideal candidate will have a strong background in SRE practices, cloud infrastructure management, and automation. If you have a consistent track record of architecting and building large-scale systems, enjoy solving intriguing system challenges at internet-scale, are innovative at heart, and have a great balance of skills in learning, organising, building, and making an impact, this role might be a great fit for you!

The candidate will have responsibilities across the following functions:

Design and Infrastructure:

  • Contribute to postmortem culture by facilitating comprehensive, blameless post-incident reviews that identify root causes, contributing factors, and actionable remediation items.
  • Track incident trends to identify systemic issues and prioritise reliability improvements.
  • Implement chaos engineering practices to proactively identify failure modes, validate system resilience, and build confidence in recovery procedures.
  • Conduct game days and disaster recovery exercises.

SRE Process and Principles Implementation:

  • Deploy and evolve SRE practices across the organisation by establishing core SRE principles, frameworks, and methodologies.
  • Define and implement service reliability practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets, to balance innovation velocity with system reliability.
  • Manage Error Budgets as a mechanism for making data-driven decisions about feature velocity vs. reliability.
  • Track, report, and enforce error budget policies, facilitating conversations between engineering and product teams about risk tolerance and release decisions.

Reliability Engineering and Infrastructure:

  • Reduce toil through automation by identifying repetitive operational work and systematically eliminating it through infrastructure-as-code, automation frameworks, and intelligent tooling.
  • Measure and track toil reduction efforts, aiming to keep toil below 50% of team time.
  • Implement capacity planning processes that ensure systems have adequate headroom to meet SLOs during peak traffic, unexpected load spikes, and degraded states.
  • Develop predictive models and automated scaling mechanisms.

Observability, Monitoring and Reporting:

  • Build comprehensive observability systems that provide deep visibility into service health, performance, and user experience. Implement monitoring strategies based on the Four Golden Signals (latency, traffic, errors, saturation) and USE/RED methodologies.
  • Create SRE dashboards and reporting mechanisms that provide real-time visibility into SLO compliance, error budget consumption, and system reliability metrics. Develop executive-level reporting on reliability trends, incident impact, and improvement initiatives.
  • Establish alerting strategies that are actionable, symptom-based, and aligned with SLOs.
  • Reduce alert fatigue by tuning thresholds and eliminating noise while ensuring critical issues trigger appropriate responses.

Collaboration and Leadership:

  • Partner with development teams to implement reliability from the design phase using SRE principles.
  • Conduct design reviews focused on failure modes, scalability, observability, and operational concerns.
  • Guide teams in building services that meet SLO requirements.
  • Collaborate through code reviews and design reviews, ensuring infrastructure-as-code, automation scripts, and reliability improvements follow best practices, are well-documented, and maintain high-quality standards.
  • Manage project priorities using error budgets as a decision-making framework.
  • Leverage agile methodologies while ensuring reliability work gets appropriate prioritisation alongside feature development.

Operational Excellence and Continuous Improvement:

  • Identify and eliminate performance bottlenecks through detailed analysis of metrics, traces, and profiles.
  • Optimise system resources, tune configurations, and implement auto-scaling to ensure SLO compliance during varying load conditions.
  • Drive continuous improvement through SRE feedback loops by analysing SLO violations, incident trends, and toil metrics to identify systemic improvements. Champion the reliability roadmap and advocate for technical debt reduction.
  • Maintain a culture of documentation and knowledge sharing by creating comprehensive runbooks, operational guides, system architecture documentation, and disaster recovery procedures.
  • Ensure operational knowledge is distributed across the team.
  • Track and report on SRE metrics, including SLO compliance rates, error budget consumption, mean time to detection (MTTD), mean time to resolution (MTTR), toil percentage, and reliability improvement velocity.

On-call and reliability:

  • Participate in a 12x7 on-call rotation and be available to work with global teams in the event of critical outages.

Requirements:

  • Preferably 8+ years of experience in DevOps/SRE roles, with demonstrated expertise in implementing SRE principles, SLO/SLI frameworks, and error budget policies in production environments.
  • Deep experience with observability and monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent, including experience building custom dashboards, alerts, and SLO-based monitoring.
  • Strong background in incident management, including experience as an Incident Commander, conducting blameless postmortems, and implementing systematic reliability improvements based on incident learnings.
  • Strong understanding of distributed systems and reliability engineering, including failure modes, fault tolerance patterns, circuit breakers, bulkheads, rate limiting, and graceful degradation strategies.
  • Experience with a number of the following: Kubernetes, Docker, Service Mesh such as Istio, Envoy, Linkerd, Solo & ECS.
  • Experience in cloud-focused software development, preferably in Go, Python, or other object-oriented programming languages.
  • Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation.
  • Experience with CI/CD automation, including GitLab pipelines and other related tools.
  • Strong hands-on experience with cloud platforms such as AWS, GCP or Azure.
  • Proven track record of implementing scalable, high-performance infrastructure solutions in fast-paced, dynamic environments.
  • Demonstrated ability to communicate clearly with both technical and non-technical project stakeholders, with the ability to work effectively in a cross-functional team environment.
  • Self-driven and detail-oriented with the ability to understand complex distributed systems and identify reliability risks proactively.
  • Certifications in relevant technologies, such as Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer, or Certified Information Systems Security Professional (CISSP), are preferred.
  • BS Degree in Computer Science or Equivalent.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
SDET 1 day ago
$12k – $39k per year (Estimated) • In office • 4+ years exp • Gurgaon
Java
DevOps
AWS
Azure
CI/CD
Jenkins
QA
Appium
JMeter
Playwright
Rest-Assured
Selenium
Apply
$18k – $46k per year (Estimated) • Remote • Full-Time • Tula
C#
JavaScript
Node JS
SQL
TypeScript
C#
.NET
Node JS
InversifyJS
Databases
DynamoDB
MySQL
AI/ML
Claude
Copilot
Cursor
OpenAI Codex
Frontend
Angular
React.js
Tailwind CSS
Mobile
Dependency Injection
DevOps
AWS
AWS Lambda
CI/CD
OpenTelemetry
Rest API
Terraform
Amazon CloudWatch
Amazon S3
API Gateway
GitHub
Cybersecurity
HIPAA
Apply
$19k – $53k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Bash
JavaScript
Python
TypeScript
Frontend
Angular
React.js
DevOps
AWS
Azure
Datadog
Docker
GCP
Grafana
Kubernetes
Prometheus
Splunk
IAM
Cybersecurity
Keycloak
Apply
$20k – $45k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Python
pySpark
Databases
Databricks
Microsoft Fabric
AI/ML
Hadoop
Spark
DevOps
AWS
Azure
Analytics
Power BI
Tableau
Apply
$11k – $21k per year (Estimated) • Remote • Saint Petersburg
PowerShell
DevOps
Azure
SLI/SLO/SLA
Windows Server
Management
ServiceNow
Apply
$31k – $60k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
Databases
Apache Kafka
Presto
Trino
AI/ML
Airflow
Hadoop
Spark
DevOps
AWS
GCP
Apply
$25k – $62k per year (Estimated) • In office • 12+ years exp • Bengaluru
Python
DevOps
Amazon EKS
AWS
Azure
Azure AKS
CI/CD
Docker
Envoy
GCP
Google GKE
Grafana
Istio
Kubernetes
Loki
Prometheus
Service Mesh
Terraform
Thanos
Amazon ECS
Apply
$34k – $82k per year (Estimated) • In office • 10+ years exp • Bengaluru
Java
Kotlin
Scala
Databases
Apache Iceberg
Apache Kafka
Delta Lake
Druid
Trino
AI/ML
Airflow
Flink
Hadoop
Spark
DevOps
AWS
GCP
Apply
$25k – $67k per year (Estimated) • In office • 8+ years exp • Bengaluru
Go
Python
DevOps
Amazon EKS
AWS
Azure
Azure AKS
CI/CD
Docker
Envoy
GCP
Google GKE
Grafana
Istio
Kubernetes
Loki
Prometheus
Service Mesh
Terraform
Thanos
Amazon ECS
Apply
Backend Engineer 1 month ago
$46k – $99k per year (Estimated) • In office • 10+ years exp • Bengaluru
JavaScript
Node JS
Python
Scala
SQL
Databases
Aerospike
Apache Kafka
Cassandra
ScyllaDB
Trino
AI/ML
Hadoop
Spark
Frontend
React.js
DevOps
AWS
GCP
Kubernetes
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.