368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$49k – $116k per year (Estimated)
Location
In office
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Stellar Cyber is a company that provides an automation-driven security operations platform seamlessly integrating next-generation SIEM, Network Detection and Response (NDR), and Open Extended Detection and Response (XDR). Their platform utilizes advanced AI to quickly detect and correlate cybersecurity threats across various security tools, offering comprehensive threat intelligence and automated incident response to enhance security operations for enterprises, MSSPs, and MSPs. With a focus on reducing security operation costs and improving threat response times, Stellar Cyber helps organizations protect their entire attack surface, including on-premises, cloud, and IT/OT environments.

Join Stellar Cyber, a fast-growing global leader in cybersecurity trusted by some of the biggest names in the industry. Besides many enterprises and government agencies, nearly 30% of the world’s top MSSPs rely on our platform, and that number is growing every day as more companies recognize the value of next-generation security solutions. We're at the forefront of protecting organizations against sophisticated cyber threats using cutting-edge AI and automation technologies. Our culture is built on diversity, openness, and collaboration, fostering creativity and innovation that drives real impact in the market.

We are seeking a highly skilled Staff Site Reliability Engineer (SRE) to join our team and drive reliability, scalability, and efficiency across our production systems. The ideal candidate will have deep expertise in cloud infrastructure, Kubernetes administration, observability, and incident management, with a proven track record of building and maintaining highly available and resilient platforms. As a senior member of the SRE team, you will not only operate complex distributed systems but also influence architecture, tooling, and best practices to ensure operational excellence.

Please note, as part of our interview process, we may invite candidates for an in-person interview to meet with our team.

Responsibilities:

  • Administer and maintain container orchestration platforms and containerized workloads.
  • Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability.
  • Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
  • Administer and optimize cloud-based environments across multiple providers.
  • Manage and support distributed data platforms and real-time processing systems.
  • Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments.
  • Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability.
  • Automate and orchestrate infrastructure using programming and scripting languages.
  • Perform system administration and networking tasks to support internal and external environments.
  • Collaborate effectively with engineers and stakeholders across different time zones.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Proven success leading large-scale production systems in cloud environments (AWS, GCP, Azure, or OCI).
  • Demonstrated leadership in driving incident response, on-call best practices, and reliability-focused culture.
  • Strong experience with production on-call operations and incident management.
  • Advanced proficiency in Kubernetes administration and troubleshooting.
  • Hands-on experience with observability tools: Prometheus, Grafana, Loki, and Alertmanager.
  • Knowledge in chat-based operations interfaces and/or auto-remediation controllers using AI agentic framework.
  • Understanding of AI agents for Auto-triaging alerts, correlate signals and suggest/root-cause hypotheses
  • Expertise in operating data platforms (Elasticsearch, MongoDB, Spark, Kafka, Redis).
  • Proficiency with public cloud services (AWS, Azure, GCP, or OCI).
  • Strong programming and automation skills in Python and Bash.
  • Deep understanding of Infrastructure as Code (Terraform, Helm).
  • Experience with CI/CD pipelines (GitHub Actions, Bitbucket, ArgoCD).
  • Strong technical background in distributed systems, databases, networking, and Linux administration.
  • Excellent problem-solving, communication, and leadership abilities.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • Certifications in AWS, GCP, Observability, Linux or Kubernetes are a plus..
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$116k – $215k per year (Estimated) • Equity • Remote • Bachelor's Degree
AI/ML
Edge AI
Apply
$100k – $150k per year • Equity • In office • Full-Time • 5+ years exp • Associate's Degree
Databases
Apache Kafka
ElasticSearch
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Kubernetes
IAM
Apply
$33k – $80k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree
Databases
Apache Kafka
ElasticSearch
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Kubernetes
IAM
Apply
$165k – $215k per year • Equity • Remote • Full-Time • 5+ years exp • New York
Bash
Python
Databases
Apache Kafka
ElasticSearch
Redis
AI/ML
Spark
Edge AI
DevOps
Alertmanager
ArgoCD
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
GitOps
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Terraform
GitHub
Apply
$120k – $215k per year • Equity • In office • 5+ years exp
AI/ML
Edge AI
DevOps
IAM
Design
Figma
Management
Confluence
Jira
Miro
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.