386,863open jobs
10,118companies
50,530added this week
Browse all
Salary
$183k – $275k per year
Location
In office (San Diego)
Seniority
Staff · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Shield AI is an American defence technology company founded in 2015 that builds autonomy software and uncrewed aircraft for military missions. Its Hivemind autonomy stack lets aircraft fly, navigate and complete objectives without GPS, a remote pilot or a communications link, which matters in contested environments where those links are jammed. Headquartered in San Diego, California, the company fields the V-BAT vertical take-off aircraft with United States and allied forces, licenses Hivemind to other manufacturers and is developing the X-BAT jet for high-end missions.

Job Description

    Hivemind is looking for an experienced SRE lead to help drive the establishment of our SRE function.

    As the SRE Lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services. You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and ensure that production systems can be operated and recovered predictably.

    This is a deeply technical, hands-on role. You will investigate complex failures, improve the systems and tooling used to operate our platforms, and turn lessons from incidents into engineering improvements. You will provide technical leadership for reliability engineering, helping teams adopt practices that improve system health without adding unnecessary process.

    You will act as a thought-leader and mentor within the Cloud Engineering and Reliability teams to level up teammates and encourage building with a reliability-first mindset.

What You'll Do

  • Define and implement SLIs, SLOs, and other measures of service reliability
  • Build and improve monitoring, alerting, logging, and tracing for infrastructure and platform services
  • Lead technical response to complex incidents and drive root-cause analysis through resolution
  • Identify recurring failure modes and work with engineering teams to eliminate them
  • Improve system resilience through automation, testing, capacity planning, and failure recovery
  • Develop tooling and automation that reduces manual operational work
  • Partner with product and platform teams to incorporate reliability requirements into system design
  • Establish incident response practices that improve detection, diagnosis, communication, and recovery
  • Mentor product engineers and drive adoption of strong reliability and operational practices
  • Mentor teammates in SRE and Cloud Engineering
  • Define and manage short-and-long term SRE roadmap, distributing work across teammates

Required Qualifications

  • 7+ years of experience in SRE, software engineering, infrastructure engineering, or related fields
  • Experience operating production services with defined availability and reliability requirements
  • Experience implementing SLIs, SLOs, monitoring, alerting, and incident response practices
  • Experience designing and operating infrastructure in AWS or another major cloud environment
  • Experience with infrastructure-as-code and automated infrastructure provisioning
  • Experience supporting containerized applications and distributed systems
  • Experience developing operational tooling or automation using Python, Go, or a similar language
  • Ability to diagnose complex failures across applications, infrastructure, networking, and dependent services
  • Experience leading incident response and root-cause analysis across engineering teams
  • Experience leading and executing on technical vision of a team over multi-quarter timelines

Preferred Qualifications

  • Experience establishing or maturing an SRE function within an engineering organization
  • Experience with Kubernetes and cloud-native observability systems
  • Experience operating systems in regulated or compliance-driven environments
  • Experience with capacity planning, performance analysis, and cloud cost management
  • Background supporting shared infrastructure across multiple products or engineering organizations
  • Experience with building roadmaps in ticketing systems
  • Experience acting as a mentor for other engineers
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
386,863 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Diego
$173k – $277k per year • In office • Full-Time • 13+ years exp • Bachelor's Degree • Austin
Python
DevOps
Akamai
Ansible
AWS
Azure
Bicep
Chef
Cloudflare
Docker
GCP
Kubernetes
Nginx
OpenShift
Puppet
Terraform
Terragrunt
Apply
$25k – $58k per year (Estimated) • In office • Full-Time • Bachelor's Degree • India
Bash
Python
DevOps
Amazon CloudWatch
Amazon EKS
AWS
CI/CD
Docker
Dynatrace
Grafana
Jenkins
Kubernetes
OpenShift
Prometheus
Splunk
SRE
Terraform
Apply
$68k – $188k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • London
JavaScript
Python
SQL
Databases
Apache Kafka
DynamoDB
Kafka
Frontend
React.js
DevOps
Amazon S3
AWS
Kubernetes
Apply
$100k – $230k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise
Python
SQL
AI/ML
Anomaly Detection
Analytics
Power BI
Tableau
Apply
$26k – $67k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Hyderabad • Ahmedabad
C#
C++
Python
SQL
TypeScript
JavaScript
C#
ASP.NET Core
Entity Framework Core
AI/ML
Claude
Copilot
Cursor
Prompt Engineering
Frontend
Angular
RxJS
DevOps
CI/CD
Git
GitHub
Cybersecurity
Qualys Cloud Platform
Apply
$152k – $228k per year • In office • Full-Time • 7+ years exp • San Diego
Go
Python
DevOps
AWS
CI/CD
Kubernetes
Apply
$152k – $228k per year • In office • Full-Time • San Diego
DevOps
AWS
IAM
Kubernetes
Terraform
Apply
$182k – $274k per year • In office • Full-Time • San Mateo
DevOps
AWS
IAM
Kubernetes
Terraform
Apply
$220k – $330k per year • In office • Full-Time • 7+ years exp • San Mateo
Python
DevOps
AWS
Kubernetes
Apply
$182k – $274k per year • In office • Full-Time • 7+ years exp • San Mateo
Go
Python
DevOps
AWS
CI/CD
Kubernetes
Apply
$152k – $228k per year • In office • Full-Time • 7+ years exp • San Diego
Go
Python
DevOps
AWS
CI/CD
Kubernetes
Apply
$152k – $228k per year • In office • Full-Time • San Diego
DevOps
AWS
IAM
Kubernetes
Terraform
Apply
$150k – $262k per year • Equity • In office • Full-Time • 8+ years exp • San Diego
C#
C++
JavaScript
TypeScript
Frontend
Angular
Lit
React.js
Vue.js
Management
ServiceNow
Apply
$82k – $102k per year • In office • Full-Time • Master's Degree • San Diego
DevOps
CI/CD
Shift-Left
Cybersecurity
Shift-Left Security
Apply
$172k – $286k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Diego
Java
Kotlin
Swift
Dart
JavaScript
Databases
Apache Kafka
Kafka
AI/ML
Prompt Engineering
RAG
Frontend
React.js
Mobile
Flutter
React Native
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
See all jobs
This is one of many
386,863 more open roles from verified company boards, updated every day.