368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$56k – $134k per year (Estimated)
Location
In office
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Stellar Cyber is a company that provides an automation-driven security operations platform seamlessly integrating next-generation SIEM, Network Detection and Response (NDR), and Open Extended Detection and Response (XDR). Their platform utilizes advanced AI to quickly detect and correlate cybersecurity threats across various security tools, offering comprehensive threat intelligence and automated incident response to enhance security operations for enterprises, MSSPs, and MSPs. With a focus on reducing security operation costs and improving threat response times, Stellar Cyber helps organizations protect their entire attack surface, including on-premises, cloud, and IT/OT environments.

Join Stellar Cyber, a fast-growing global leader in cybersecurity trusted by some of the biggest names in the industry. Besides many enterprises and government agencies, nearly 30% of the world’s top MSSPs rely on our platform, and that number is growing every day as more companies recognize the value of next-generation security solutions. We're at the forefront of protecting organizations against sophisticated cyber threats using cutting-edge AI and automation technologies. Our culture is built on diversity, openness, and collaboration, fostering creativity and innovation that drives real impact in the market.

We are seeking a highly skilled Staff Site Reliability Engineer (SRE) to join our team and drive reliability, scalability, and efficiency across our production systems. The ideal candidate will have deep expertise in cloud infrastructure, Kubernetes administration, observability, and incident management, with a proven track record of building and maintaining highly available and resilient platforms. As a senior member of the SRE team, you will not only operate complex distributed systems but also influence architecture, tooling, and best practices to ensure operational excellence.

Please note, as part of our interview process, we may invite candidates for an in-person interview to meet with our team.

Responsibilities:

  • Administer and maintain container orchestration platforms and containerized workloads.
  • Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability.
  • Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
  • Administer and optimize cloud-based environments across multiple providers.
  • Manage and support distributed data platforms and real-time processing systems.
  • Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments.
  • Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability.
  • Automate and orchestrate infrastructure using programming and scripting languages.
  • Perform system administration and networking tasks to support internal and external environments.
  • Collaborate effectively with engineers and stakeholders across different time zones.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Proven success leading large-scale production systems in cloud environments (AWS, GCP, Azure, or OCI).
  • Demonstrated leadership in driving incident response, on-call best practices, and reliability-focused culture.
  • Strong experience with production on-call operations and incident management.
  • Advanced proficiency in Kubernetes administration and troubleshooting.
  • Hands-on experience with observability tools: Prometheus, Grafana, Loki, and Alertmanager.
  • Knowledge in chat-based operations interfaces and/or auto-remediation controllers using AI agentic framework.
  • Understanding of AI agents for Auto-triaging alerts, correlate signals and suggest/root-cause hypotheses
  • Expertise in operating data platforms (Elasticsearch, MongoDB, Spark, Kafka, Redis).
  • Proficiency with public cloud services (AWS, Azure, GCP, or OCI).
  • Strong programming and automation skills in Python and Bash.
  • Deep understanding of Infrastructure as Code (Terraform, Helm).
  • Experience with CI/CD pipelines (GitHub Actions, Bitbucket, ArgoCD).
  • Strong technical background in distributed systems, databases, networking, and Linux administration.
  • Excellent problem-solving, communication, and leadership abilities.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • Certifications in AWS, GCP, Observability, Linux or Kubernetes are a plus.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$22k – $57k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
Python
TypeScript
JavaScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
Claude
Claude Code
LLM
Model Context Protocol
Frontend
Angular
Next.js
React.js
DevOps
AWS
AWS Lambda
CI/CD
Docker
GitHub Actions
Terraform
Amazon ECS
Amazon S3
GitHub
IAM
Apply
$125k – $150k per year • In office • Full-Time • 3+ years exp • New York
SQL
Python
AI/ML
AI Agents
DevOps
AWS
Azure
Grafana
Kubernetes
Apply
$220k – $325k per year • Remote/Hybrid • Full-Time • 15+ years exp • New York
DevOps
AWS
Azure
CI/CD
GCP
Kubernetes
Platform Engineering
Cybersecurity
Threat Modeling
Apply
$116k – $215k per year (Estimated) • Equity • Remote • Bachelor's Degree
AI/ML
Edge AI
Apply
$100k – $150k per year • Equity • In office • Full-Time • 5+ years exp • Associate's Degree
Databases
Apache Kafka
ElasticSearch
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Kubernetes
IAM
Apply
$33k – $80k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree
Databases
Apache Kafka
ElasticSearch
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Kubernetes
IAM
Apply
$165k – $215k per year • Equity • Remote • Full-Time • 5+ years exp • New York
Bash
Python
Databases
Apache Kafka
ElasticSearch
Redis
AI/ML
Spark
Edge AI
DevOps
Alertmanager
ArgoCD
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
GitOps
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Terraform
GitHub
Apply
$120k – $215k per year • Equity • In office • 5+ years exp
AI/ML
Edge AI
DevOps
IAM
Design
Figma
Management
Confluence
Jira
Miro
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.