1,088,837open jobs
63,338companies
185,353added this week
Browse all
Salary
≈ $89k – $181k per year (Estimated)
Location
Hybrid (Austin, United States)
Seniority
Middle · 4+ years exp

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Aug 21, 2026. Seekr scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Decision-ready, explainable,sovereign AI. Seekr's trusted AI solutions power critical decisions for government and enterprise sectors where accuracy, transparency, and compliance are paramount.

We're looking for an AI Site Reliability Engineer to ensure the reliability, scalability, and safe operation of Seekr's AI-powered services and supporting infrastructure. You will combine software engineering and site reliability practices with AI/ML operational expertise to improve how models, APIs, data pipelines, and platform services are deployed, monitored, and operated in production. You will partner closely with AI/ML, Platform, Security, and Product teams to build dependable systems that are performant, resilient, and ready to scale. 

The Impact  

You will help define how Seekr operates mission-critical AI systems in production. Your work will make releases safer, incidents less frequent and easier to resolve, and service health more visible and measurable. By establishing meaningful SLOs, strengthening observability, improving resilience, and reducing operational toil, you will directly improve the experience of our customers and engineering teams. 

Duties and Responsibilities  

  • Build safe release processes using CI/CD automation, progressive deployments, automated validation, rollback mechanisms, and deployment health metrics.
  • Define and operate SLIs/SLOs for AI APIs and critical services, covering availability, latency, errors, throughput, model quality, output safety, user-facing correctness, and cost.
  • Develop actionable observability through metrics, logs, traces, dashboards, and SLO-based alerts.
  • Participate in a sustainable on-call rotation; lead incident response, improve runbooks, and facilitate blameless postmortems.
  • Reduce operational toil and improve resilience through infrastructure as code, automation, and disaster-recovery planning.
  • Design automated load, stress, spike, soak, and scalability tests that model realistic AI production traffic.
  • Establish performance baselines, release thresholds, and capacity forecasts for latency, throughput, concurrency, resource utilization, and cost per inference.
  • Validate autoscaling, rate limiting, backpressure, graceful degradation, and recovery from infrastructure and dependency failures.
  • Design and scale GPU-backed model inference on Kubernetes using replicas, autoscaling, batching, caching, and load balancing to meet SLAs/SLOs.
  • Optimize and troubleshoot inference engines and infrastructure across models, Kubernetes, compute, networking, storage, GPU memory, and cost.
  • Support reliable training workloads, including distributed training, scheduling, checkpointing, observability, and failure recovery.
  • Partner with AI/ML, Platform, Security, and Product teams to establish production-readiness standards. 

Qualifications and Skills Required  

  • Bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, production infrastructure engineering, or a similar role.
  • Strong knowledge of distributed systems, cloud infrastructure, networking, containers, and Kubernetes.
  • Experience building and operating CI/CD pipelines, automated validation, and progressive deployment strategies.
  • Proficiency with infrastructure-as-code and configuration-management tools, such as Terraform.
  • Experience with monitoring, logging, distributed tracing, alerting, and incident-management practices, including tools such as Prometheus, Grafana, and OpenTelemetry.
  • Strong scripting and software-development skills in Python, Go, or a similar language.
  • Practical knowledge of SLIs, SLOs, error budgets, capacity planning, production-readiness reviews, and blameless postmortems.
  • Ability to troubleshoot complex systems across application, model-serving, and infrastructure layers.
  • Strong communication and collaboration skills, with a commitment to sustainable and blameless operations.
  • Experience operating machine-learning or generative-AI systems in production is preferred.
  • Familiarity with model serving, inference optimization, LLM gateways, vector databases, GPU infrastructure, ML observability, and model evaluation is preferred.
  • Experience effectively utilizing AI technologies and tools, including large language models, agents, or AI coding assistants, to enhance workflows and operational effectiveness.
  • Hands-on experience with performance-testing tools such as k6, Locust, JMeter, or Gatling.
  • Experience testing distributed systems and interpreting latency percentiles, saturation, throughput, concurrency, and resource-consumption metrics.
  • Familiarity with AI inference benchmarking, GPU profiling, autoscaling, and performance-versus-cost optimization.
  • Experience with inference engines such as vLLM, NVIDIA Triton, TensorRT-LLM, or similar platforms, and scaling GPU workloads on Kubernetes.
  • Working knowledge of GPU infrastructure, memory management, quantization, distributed training, and frameworks such as PyTorch, DeepSpeed, or FSDP. 

About the Company:   

Seekr is a leader in explainable and trustworthy artificial intelligence designed to power mission-critical decisions in enterprises, government, and regulated industries. SeekrFlow™, our end-to-end AI platform, provides secure, auditable AI solutions tailored to sectors where transparency, accuracy, and compliance are paramount. Available across cloud, on-premises, and edge environments, SeekrFlow reduces bias, strengthens data integrity, and simplifies model oversight so organizations can rely on trusted AI decisions in high-stakes settings that impact society’s most sensitive and vital systems. Trusted by leading enterprises and government agencies, we partner with defense, finance, telecom, and critical infrastructure leaders to enable AI solutions that drive real-world results with unmatched transparency and control. We are a team of strategic thinkers and problem-solvers tackling the toughest challenges facing critical infrastructure and global enterprises through best-in-class AI models and customer deployment. Our team operates with unwavering commitment to our core values and mission:

  • We are driven by outcomes-our customers' success is what we strive for every day.
  • We believe trust is earned, which is why we build explainability and transparency into the entire AI lifecycle.
  • We take our responsibility to deliver secure AI seriously.
  • We believe innovation drives progress-we are building the technologies that power the systems our society depends on.

Company Benefits:

  • Meaningful Mission & Impact - Work with a deeply talented, collaborative team solving some of the toughest AI challenges that matter. 
  • Equity Ownership  - RSUs that let you share directly in Seekr’s long-term success and growth.

  • Time Off That Respects Real Life  - Unlimited PTO plus 14 paid company holidays to truly recharge. 
  • Work Your Way  - A flexible hybrid work environment with offices in Reston, VA and Austin, TX.

  • Competitive Total Rewards  - A role-appropriate compensation structure that supports long-term growth, including base salary, bonuses, or commission plans depending on role.

  • 401(k) with Company Match  - Build your future with a retirement plan that includes employer matching. 
  • Comprehensive Health & Wellness  - Medical, dental, vision, and life insurance coverage starting day one-for you and your family.  
  • Parental Leave  - Paid parental leave to support employees as they welcome a new child through birth, adoption, or foster placement.

Important Notice: 

To conform to U.S. Government international trade regulations, applicant must be a U.S. Citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State or U.S. Department of Commerce.

Seekr is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other characteristic protected by applicable law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,088,837 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Austin
≈ $82k – $167k per year (Estimated) • Hybrid • 3+ years exp • High School Diploma
DevOps
Windows
DNS
DHCP
Wi-Fi
Apply
≈ $70k – $150k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Los Angeles
Python
Go
TypeScript
C#
DevOps
Terraform
GitHub Actions
AWS CDK
GitLab CI
CI/CD
Git
AWS
Kubernetes
Amazon EKS
Progressive Delivery
SLI/SLO/SLA
Apply
$93k – $125k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
Python
Bash
DevOps
Splunk
Prometheus
Platform Engineering
Linux
Apply
$135k – $181k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
AI/ML
Computer Vision
DevOps
GCP
Azure
AWS
Linux
Windows
Cybersecurity
Wiz
Management
Agile
Apply
$74k – $116k per year • Equity • In office • Full-Time • 3+ years exp • Associate's Degree • Louisville
PowerShell
DevOps
Windows
Cybersecurity
BeyondTrust
Analytics
Power BI
Management
Power Automate
Microsoft Teams
SharePoint
Microsoft Office
Apply
$180k – $210k per year • Remote (United States) • Full-Time • 7+ years exp • Bachelor's Degree • New York
Python
JavaScript
TypeScript
Node JS
AI/ML
Cursor
Claude Code
AI Agents
LLM
RAG
OpenAI Codex
Frontend
React.js
DevOps
Terraform
GitHub Actions
Azure
CI/CD
AWS
Bicep
Apply
≈ $22k – $44k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru • Mumbai • Thiruvananthapuram
Python
SQL
Scala
Python
pySpark
Databases
Snowflake
Databricks
AI/ML
Spark
DevOps
GCP
Azure DevOps
Azure
AWS
Analytics
Dimensional Modeling
Apply
≈ $18k – $42k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Python
JavaScript
Java
SQL
Node JS
Scala
Databases
Amazon Redshift
DevOps
AWS
AWS Lambda
Amazon S3
IAM
Amazon Kinesis
Analytics
ETL/ELT
AWS Glue
Apply
$151k – $205k per year • Equity • In office • Full-Time • Redmond
Python
Go
Java
Rust
Ruby
C#
C++
MATLAB
DevOps
CI/CD
Linux
Unix
Apply
≈ $11k – $28k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Bengaluru
Python
Java
Databases
Amazon Redshift
Analytics
Tableau
Power BI
MicroStrategy
Apply
≈ $114k – $222k per year (Estimated) • Equity • Hybrid • Secret • 5+ years exp • Reston
Python
Bash
Databases
PostgreSQL
ElasticSearch
OpenSearch
DevOps
Terraform
Ansible
Helm
Cilium
Istio
Loki
Traefik
K3s
Kustomize
Prometheus
Azure
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Linux
TCP/IP
DNS
BGP
Cybersecurity
FedRAMP
CVE
Calico
Apply
≈ $136k – $270k per year (Estimated) • Equity • Hybrid • 8+ years exp • Austin
Python
Rust
C++
AI/ML
Ray Serve
vLLM
Triton Inference Server
Fine-tuning
Quantization
AI Agents
SGLang
TensorRT
TensorRT-LLM
Ray
Edge AI
Multi-Agent Systems
Machine Learning
DevOps
GCP
Helm
OpenTelemetry
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Apply
≈ $117k – $228k per year (Estimated) • Equity • Hybrid • 5+ years exp • Austin
Python
Rust
C++
AI/ML
Ray Serve
vLLM
Triton Inference Server
Fine-tuning
Quantization
AI Agents
SGLang
TensorRT
TensorRT-LLM
Ray
Edge AI
Multi-Agent Systems
Machine Learning
DevOps
GCP
Helm
OpenTelemetry
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Apply
≈ $117k – $228k per year (Estimated) • Equity • Hybrid • 5+ years exp • Austin
Python
Bash
Databases
PostgreSQL
ElasticSearch
OpenSearch
DevOps
Terraform
Ansible
Helm
Cilium
Istio
Loki
Traefik
K3s
Kustomize
Prometheus
Azure
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Linux
TCP/IP
DNS
BGP
Cybersecurity
Calico
Apply
Office Manager 7 days ago
≈ $46k – $89k per year (Estimated) • Equity • Hybrid • 3+ years exp • Austin
Management
Google Workspace
Apply
$187k – $308k per year • Equity • Remote (United States) • Full-Time • 10+ years exp • San Francisco • Austin • Chicago • Seattle
DevOps
gRPC
OpenTelemetry
Datadog
Kustomize
Prometheus
GitLab CI
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Amazon EKS
Linux
DNS
Cybersecurity
PKI
Apply
≈ $117k – $227k per year (Estimated) • Hybrid • Bachelor's Degree • Austin
DevOps
GCP
Kubernetes
Apply
≈ $104k – $236k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Databases
PostgreSQL
Oracle
AI/ML
Replicate
DevOps
Splunk
Dynatrace
AWS
Apply
$175k – $291k per year • Remote (United States) • Full-Time • 12+ years exp • Bachelor's Degree • Austin • Saint Louis
AI/ML
AI Agents
Feature Store
Knowledge Graph
LLM Guardrails
DevOps
GCP
Azure
CI/CD
AWS
Apply
$243k – $364k per year • Equity • Remote (United States) • Full-Time • 15+ years exp • Bachelor's Degree • San Jose • Austin • Seattle
Management
Agile
Apply
See all jobs
This is one of many
1,088,837 more open roles from verified company boards, updated every day.