664,807open jobs
38,858companies
99,669added this week
Browse all
Salary
$180k – $220k per year
Location
Remote/Hybrid (United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Onebrief is collaboration and AI-powered workflow software designed specifically for military staffs. Onebrief makes the staff as a whole superhuman — faster, smarter, and more efficient — to ultimately support stronger decision-making.

Consequential Work. Dedicated People.

About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Military planning is complex by nature, requiring teams to coordinate information, people, and decisions across systems and locations. Onebrief brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity when the stakes are real.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This role requires regularly working on-site at customer locations in Arlington, VA.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Secret Clearance required; SCI eligibility is a plus.

About The Role

We are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success.

You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation.

In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise.

About You

You care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them.

You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts.

You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality.

What You'll Do

You'll own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.

  • Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it.

  • Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.

  • Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.

  • Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.

What We Look For

  • An active Top Secret clearance

  • 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.

  • Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly.

  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.

Technical expertise

  • Infrastructure as Code: Terraform (or CloudFormation), Ansible.

  • Containers and orchestration: Kubernetes design, deployment, and operations.

  • CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions).

  • Scripting: proficiency with at least one of Python, Go, or Bash.

  • Cloud: Familiarity with AWS or AWS GovCloud.

  • Observability: Grafana stack, ELK stack, or Datadog.

  • Networking fundamentals: core protocols and secure configurations.

Bonus points (nice to have)

  • Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503).

  • GitOps practices and toolchains.

  • Security-minded design for sensitive environments.

  • Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems.

  • Familiarity with on-prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc).

  • Service mesh exposure (Istio, Linkerd).

  • Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD).

  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment.

Notice to Third Party Recruitment Agencies

Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
664,807 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
SRE Sênior 1 day ago
$44k – $107k per year (Estimated) • Remote • Full-Time
Python
Bash
Databases
Apache Kafka
OpenSearch
AI/ML
Copilot
Cursor
Claude
ChatGPT
Edge AI
DevOps
Terraform
Helm
GitHub Actions
Istio
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Service Mesh
FinOps
Incident Management
Apply
In office • Mumbai
JavaScript
TypeScript
Node JS
Frontend
GraphQL
Next.js
React.js
Sass
Mobile
React Native
DevOps
CI/CD
AWS
AWS Lambda
Amazon S3
Apply
$52k – $130k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
JavaScript
Java
Kotlin
TypeScript
SQL
Scala
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
Databricks
Amazon Redshift
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Kubernetes
Incident Management
Apply
$104k per year • In office • Full-Time • 7+ years exp • New York • London
Python
SQL
AI/ML
Pandas
Analytics
Alteryx
Microsoft Excel
Apply
$12k – $27k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Manila
Python
Apply
$180k – $220k per year • In office • TS/SCI • Full-Time • Colorado Springs
Python
JavaScript
Node JS
Bash
Node JS
Commander.js
DevOps
Terraform
Ansible
GitHub Actions
Istio
Loki
VMWare
CloudFormation
Datadog
Linkerd
Prometheus
GitLab CI
CI/CD
GitOps
Jenkins
AWS
Kubernetes
Grafana
Service Mesh
kubectl
Proxmox VE
Hyper-V
Apply
$180k – $220k per year • Remote/Hybrid • Secret • Full-Time • 5+ years exp
Python
JavaScript
TypeScript
Node JS
Bash
Node JS
Commander.js
DevOps
Terraform
Ansible
GitHub Actions
Istio
Loki
VMWare
Datadog
Linkerd
Prometheus
GitLab CI
CI/CD
GitOps
Jenkins
AWS
Kubernetes
Grafana
Service Mesh
kubectl
Proxmox VE
Hyper-V
Apply
$126k – $154k per year • Remote • Full-Time • 2+ years exp
Apply
$170k – $190k per year • In office • TS/SCI • Full-Time • Washington
Apply
$135k – $170k per year • In office • Top Secret • Full-Time • 3+ years exp • Leavenworth
Python
C#
C++
Game Dev
Godot
Apply
See all jobs
This is one of many
664,807 more open roles from verified company boards, updated every day.