781,304open jobs
49,269companies
119,822added this week
Browse all
Salary
$115k – $252k per year
Location
Remote (United States)

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 23, 2026. CACI International scores D on the Alion truth index.

Overview
Company
Impact
Profile match
CACI International is an American technology company founded in 1962 that works almost entirely for United States defence, intelligence and federal civilian agencies. It has moved from being a conventional IT services contractor toward supplying its own technology, with a substantial expeditionary electronic warfare and signals intelligence product line, photonics and optical communications for space, and secure network modernisation. Headquartered in Reston, Virginia, the company employs more than twenty thousand people, the majority holding security clearances, and has grown through a steady programme of acquisitions that add differentiated technology rather than headcount alone.
Job Title: SRE Platform EngineerJob Category: Information TechnologyTime Type: Full timeMinimum Clearance Required to Start: NoneEmployee Type: RegularPercentage of Travel Required: NoneType of Travel: None* * *

The Opportunity:

CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland Security (DHS) Office of the Inspector General (OIG). This role offers a unique opportunity to ensure the reliability, performance, and availability of cutting-edge AI and data analytics platforms that strengthen national security oversight through investigative, audit, and inspection operations.

As an SRE Platform Engineer, you will be the operational guardian of mission-critical systems including OIG Chat-a large language model assistant-enterprise data platforms, and AI-powered applications that federal oversight professionals depend on daily. You will monitor production environments, implement comprehensive alerting and observability, troubleshoot complex incidents, optimize performance and costs, and ensure high availability through proactive capacity planning and reliability engineering. From analyzing performance metrics to identify bottlenecks to supporting disaster recovery and operational readiness, you will apply site reliability engineering principles to maintain service excellence. Working within secure Azure Government environments, you will build operational resilience for a transformational program of national importance. Join us to make a meaningful impact by ensuring mission-critical AI and data capabilities are always available, performant, and reliable.

Responsibilities:

Monitor, maintain, and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms to ensure availability, performance, service health, and adherence to service level objectives (SLOs)

Implement comprehensive observability including alerting, dashboards, health checks, synthetic monitoring, and log analysis using Azure Monitor, Application Insights, Log Analytics, or equivalent tools to enable proactive incident detection

Lead incident response and troubleshooting efforts including root cause analysis, defect resolution, dependency updates, integration validation, and coordination of emergency changes to restore service rapidly

Analyze performance and usage metrics across applications, APIs, AI model endpoints, data pipelines, and infrastructure to identify and remediate bottlenecks, latency issues, resource constraints, and efficiency opportunities

Support capacity planning, resource sizing, autoscaling configuration, and cost optimization for compute, storage, and AI model consumption to balance performance requirements with fiscal responsibility

Implement and maintain backup/restore processes, disaster recovery procedures, high availability architectures, and business continuity capabilities to ensure data protection and operational resilience

Monitor data platform availability including Azure Databricks clusters, data pipelines, storage services, and analytical workloads with alerting for pipeline failures, job errors, and performance degradation

Support deployment and operational readiness for pilot applications and new capabilities including pre-production validation, performance testing, runbook development, and go-live coordination

Provide surge support for complex technical issues, large-scale data collection analysis, analytical environment optimization, and specialized troubleshooting requiring deep platform knowledge

Develop and maintain operational documentation including runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge transfer materials to support sustainable operations

Qualifications:

Required -

Bachelor's degree + 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or related field; equivalencies considered (Master's + 12 years; 21 years with no degree; AA + 17 years)

Must be able to obtain a Active DHS/ EOD Clearance as required.

Extensive experience with Site Reliability Engineering (SRE) principles including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering practices

Proven expertise with Azure cloud services including compute, storage, networking, monitoring, and platform-as-a-service (PaaS) offerings with deep understanding of operational best practices

Strong experience with monitoring and observability tools (Azure Monitor, Application Insights, Grafana, Prometheus, ELK stack) and implementing alerting, dashboards, and log aggregation

Demonstrated ability to troubleshoot complex technical issues across application, platform, and infrastructure layers with strong analytical and problem-solving skills

Desired -

Experience operating AI/ML platforms, large language model services (Azure OpenAI), data analytics platforms (Databricks, Synapse), or high-scale cloud applications in production environments

Hands-on experience with Azure Government or other secure government cloud environments (AWS GovCloud) with understanding of compliance monitoring, security operations, and federal operational requirements

Background in federal government, mission-critical systems, or 24/7 operational environments with experience supporting incident response, change management, and operational excellence programs

-

What You Can Expect:

A culture of integrity.

At CACI, we place character and innovation at the center of everything we do. As a valued team member, you’ll be part of a high-performing group dedicated to our customer’s missions and driven by a higher purpose - to ensure the safety of our nation.

An environment of trust.

CACI values the unique contributions that every employee brings to our company and our customers - every day. You’ll have the autonomy to take the time you need through a unique flexible time off benefit and have access to robust learning resources to make your ambitions a reality.

A focus on continuous growth.

Together, we will advance our nation's most critical missions, build on our lengthy track record of business success, and find opportunities to break new ground - in your career and in our legacy.

Pay Range:

There are a host of factors that can influence final salary including, but not limited to, geographic location, Federal Government contract labor categories and contract wage rates, relevant prior work experience, specific skills and competencies, education, and certifications. Our employees value the flexibility at CACI that allows them to balance quality work and their personal lives. We offer competitive compensation, benefits and learning and development opportunities. Our broad and competitive mix of benefits options is designed to support and protect employees and their families. At CACI, you will receive comprehensive benefits such as; healthcare, wellness, financial, retirement, family support, continuing education, and time off benefits.

Since this position can be worked in more than one location, the range shown is the national average for the position.

The proposed salary range for this position is:

$114,600-$252,100

CACI is an Equal Opportunity Employer.All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, age, national origin, disability, status as a protected veteran, or any other protected characteristic.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
781,304 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Washington
≈ $117k – $216k per year (Estimated) • Remote (United States) • 8+ years exp • Bachelor's Degree
AI/ML
Claude
ChatGPT
DevOps
Terraform
Datadog
Prometheus
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon ECS
Cybersecurity
ISO 27001
SOC 2
NIST 800-53
FedRAMP
Zero Trust
Management
Agile
Waterfall
Apply
$138k – $171k per year • Equity • Remote (United States) • 2+ years exp • Master's Degree
Python
Bash
DevOps
Terraform
Ansible
CI/CD
SRE
Akamai
SLI/SLO/SLA
GitLab
Linux
Management
Confluence
Jira
Agile
Apply
≈ $95k – $198k per year (Estimated) • In office • Wilmington
SQL
Databases
Apache Kafka
DevOps
Splunk
Terraform
Datadog
Dynatrace
Prometheus
CI/CD
Jenkins
AWS
Docker
Kubernetes
Grafana
Amazon EKS
SLI/SLO/SLA
GitLab
Amazon CloudWatch
Management
ServiceNow
ITIL
Apply
$150k – $230k per year • Equity • In office • Full-Time • 6+ years exp • San Mateo
Python
Go
AI/ML
AWS Bedrock
DevOps
Terraform
Ansible
GCP
Helm
GitHub Actions
Istio
CloudFormation
Linkerd
Pulumi
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Service Mesh
Amazon EKS
Google GKE
Azure AKS
Progressive Delivery
IAM
Cybersecurity
SOC 2
HIPAA
FedRAMP
Zero Trust
Sumo Logic
Apply
$120k per year • In office • Full-Time • Park
Java
Databases
DynamoDB
Apache Kafka
AI/ML
AI Agents
DevOps
CloudFormation
Prometheus
AWS
Grafana
AWS Lambda
Amazon EC2
Amazon S3
Amazon CloudWatch
API Gateway
Apply
≈ $46k – $107k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Madrid • Barcelona
DevOps
GCP
Azure
Windows Server
AWS
Linux
Cybersecurity
MITRE ATT&CK
Trend Micro
SIEM
Apply
≈ $21k – $51k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bengaluru • Pune
Python
JavaScript
SQL
AI/ML
Synthetic Data
DevOps
Rest API
Azure DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Shift-Left
Self-Healing
Cybersecurity
Shift-Left Security
Management
Agile
Scrum
QA
Selenium
Cucumber
JMeter
Cypress
Playwright
Gatling
Postman
Rest-Assured
SoapUI
Apply
≈ $35k – $84k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Mumbai
SQL
C#
C#
.NET
Databases
Redis
Neo4j
Oracle
Cassandra
Milvus
Pinecone
Qdrant
Couchbase
AI/ML
Model Context Protocol
Fine-tuning
Multimodal AI
Computer Vision
AI Agents
NLP
RAG
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
Senior AI Engineer 1 hour ago
≈ $38k – $98k per year (Estimated) • Hybrid • Full-Time • Bengaluru
Python
SQL
Databases
Weaviate
Pinecone
FAISS
AI/ML
Copilot
Cursor
LangGraph
LangChain
Claude Code
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
CrewAI
LLM
RAG
OpenAI
Context Engineering
Edge AI
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Azure
CI/CD
Git
AWS
GitHub
Management
Agile
Scrum
Apply
≈ $31k – $87k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • PhD • Bengaluru
Python
Java
SQL
Databases
PostgreSQL
ElasticSearch
Apache Kafka
Trino
AI/ML
AI Agents
Agentforce
DevOps
Terraform
GCP
Azure
AWS
Docker
Kubernetes
Amazon S3
Management
Scrum
Apply
$99k – $207k per year • In office • TS/SCI • 7+ years exp • Bachelor's Degree • King of Prussia
Python
Bash
DevOps
Terraform
Ansible
GCP
Helm
GitHub Actions
AWS CDK
CloudFormation
Packer
Pulumi
GitLab CI
Azure
CI/CD
GitOps
Git
AWS
Kubernetes
Grafana
Amazon EKS
AWS Lambda
Amazon EC2
Linux
Management
Agile
Scrum
Apply
$115k – $252k per year • Remote (United States) • 13+ years exp • Bachelor's Degree • Washington
Python
PowerShell
AI/ML
Edge AI
DevOps
Terraform
Ansible
Azure DevOps
GitHub Actions
CloudFormation
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Configuration Management
Bicep
Azure AKS
Cybersecurity
NIST 800-53
FedRAMP
Management
Agile
Apply
$104k – $218k per year • Hybrid • 7+ years exp • Bachelor's Degree • Ashburn
Python
C
Bash
C
FFmpeg
DevOps
WebRTC
AWS
Ubuntu
Linux
Cybersecurity
Wireshark
Management
Agile
Apply
$75k – $158k per year • In office • Secret • 13+ years exp • Bachelor's Degree • Ashburn
PowerShell
DevOps
Terraform
Ansible
VMWare
Azure
CI/CD
AWS
Configuration Management
Cybersecurity
SonarQube
Management
ITIL
Apply
$87k – $182k per year • In office • Secret • 13+ years exp • Bachelor's Degree • Ashburn
DevOps
GCP
Azure
AWS
SLI/SLO/SLA
Cybersecurity
PKI
Management
ITIL
Service Desk
Apply
≈ $145k – $277k per year (Estimated) • In office • 8+ years exp • Washington
Apply
$192k – $321k per year • Remote (United States) • Full-Time • 10+ years exp • PhD • Chicago • Washington • New York
AI/ML
AI Agents
Agentforce
Marketing
Salesforce
Apply
≈ $121k – $244k per year (Estimated) • Remote (United States) • Contractor • 10+ years exp • Washington
Cybersecurity
Crowdstrike
Apply
$161k – $257k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Washington
Apply
$4k – $14k per year • Equity • Hybrid • 7+ years exp • Bachelor's Degree • Washington
Python
JavaScript
TypeScript
Node JS
AI/ML
Fine-tuning
LLM Guardrails
DevOps
Docker
Kubernetes
Apply
See all jobs
This is one of many
781,304 more open roles from verified company boards, updated every day.