687,351open jobs
40,172companies
97,615added this week
Browse all
Salary
$86k – $179k per year (Estimated)
Location
Remote/Hybrid (Jersey City, United States)
Seniority
Middle · 4+ years exp
Overview
Company
Impact
Profile match
Exiger transforms supply chain management from a complex challenge into a strategic advantage—driving savings and operational excellence in today’s volatile market. Exiger’s single, intuitive 1Exiger platform provides instant visibility into vast supplier ecosystems through a single pane of glass.

Who We Are:

Exiger transforms supply chains into a strategic advantage, advancing our mission to make the world a safer and more transparent place to succeed. Our AI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging proprietary data and advanced AI to surface risk, automate compliance, and unlock efficiencies and cost savings to strengthen long-term resilience. Trusted by 550+ global customers, including Fortune 500 companies and U.S. government agencies, Exiger is a recognized, award-winning leader in supply chain AI and a FedRAMP® authorized provider to the federal government.

Site Reliability Engineer

Location: U.S. (Hybrid)

This role requires U.S. citizenship and eligibility for a U.S. security clearance.

Role Summary:

Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI-powered platform equips the world's most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data-driven operational decisions.

From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high-stakes environments.

This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job.

You will use your expertise in coding, algorithms, complexity analysis, and large-scale distributed system design to solve the reliability challenges that are unique to operating a mission-critical AI platform in regulated and government environments.

SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

What You'll Do:

  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt.
  • Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight.
  • Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities.
  • Own the reliability of production services from design consulting and launch reviews through steady-state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self-service tooling.
  • Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault-injection testing, running game days that prove the platform degrades gracefully and recovers from failure.
  • Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on-call rotation.
  • Leverage AI-assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively.

What You Need:

  • Bachelor's or Master's degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function.
  • 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work.
  • A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics-driven decisions rather than relying on intuition or anecdote.
  • Experience with chaos engineering or fault-injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience.
  • Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems to root cause.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high-stakes environments.
  • Willingness to travel as needed to support customer engagements.

Nice to Have:

  • 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script.
  • Experience supporting ML or data platforms in production.
  • Familiarity with data warehouses such as Snowflake, Redshift and/or Apache Iceberg.
  • Experience operating in FedRAMP or other regulated or government environments.

Why You'll Love Working at Exiger:

  • High-performance culture rooted in accountability, collaboration, and a shared commitment to excellence.
  • Discretionary Time Off for all employees, with no maximum limits on time off
  • Industry leading health, vision, and dental benefits
  • Competitive compensation package
  • 16 weeks of fully paid parental leave
  • Flexible, hybrid approach to working from home and in the office where applicable
  • Focus on wellness and employee health through stipends and dedicated wellness programming
  • Purposeful career development programs with reimbursement provided for educational certifications

Exiger is named a Leader in the Gartner® Magic Quadrant™ for Supplier Risk Management, twice selected as one of Fast Company's 'Brands That Matter,' and recipient of the Third Party Risk Association's Innovator Award, Exiger's technology has been recognized by leading analyst evaluations and 50+ awards. Learn more at Exiger.com  and follow Exiger on  LinkedIn.

At Exiger, our values define how we work and why we lead. We are mission-inspired, imagination-driven, trust-anchored, and compassion-focused-committed to building technology that makes the world safer, more transparent, and more resilient.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.

Exiger’s hybrid work policy is periodically reviewed and adjusted to align with evolving business needs.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
687,351 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Jersey City
$16k – $46k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Makati
Python
SQL
Python
Flask
FastAPI
Databases
Snowflake
AI/ML
Copilot
Cursor
Claude
Claude Code
Prefect
AI Agents
Anthropic
Devin
GPT-4
DevOps
Rest API
Git
AWS
AWS Fargate
AWS Lambda
Amazon EC2
SLI/SLO/SLA
Amazon S3
Amazon ECS
Amazon EventBridge
AWS Step Functions
Analytics
Alteryx
Management
SharePoint
Apply
$95k – $182k per year (Estimated) • In office • Full-Time • Sydney
C#
Swift
AI/ML
Machine Learning
Frontend
GraphQL
Mobile
SwiftUI
MVI
MVVM
Clean Architecture
Dependency Injection
DevOps
CI/CD
AWS
Apply
$15k – $36k per year (Estimated) • Remote/Hybrid • Full-Time • Manila
JavaScript
PHP
PHP
Magento
Composer
Databases
MySQL
MariaDB
ElasticSearch
OpenSearch
AI/ML
AI Agents
Frontend
Vue.js
GraphQL
Next.js
React.js
Mobile
Dependency Injection
Algolia
DevOps
Rest API
GCP
New Relic
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Configuration Management
Fastly
Linux
Management
Agile
Apply
$37k – $81k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Chennai
Databases
Db2
AI/ML
Anomaly Detection
DevOps
AIOps
Incident Management
Apply
Remote/Hybrid • Full-Time • Melbourne
SQL
DevOps
Terraform
Azure DevOps
Datadog
Azure
CI/CD
AWS
Kubernetes
GitHub
Apply
$114k – $228k per year (Estimated) • Remote/Hybrid • 8+ years exp • Bachelor's Degree • McLean
Python
JavaScript
TypeScript
PowerShell
Bash
AI/ML
AI Agents
LLM Guardrails
DevOps
Rest API
Terraform
Helm
CI/CD
Docker
Kubernetes
Platform Engineering
IAM
Linux
Cybersecurity
Crowdstrike
Qualys Cloud Platform
Microsoft Sentinel
ISO 27001
SOC 2
FedRAMP
Threat Modeling
Kyverno
OWASP
Management
Jira
ServiceNow
Apply
$80k – $166k per year (Estimated) • Remote/Hybrid • Top Secret • 4+ years exp • McLean
Python
SQL
Bash
Perl
AI/ML
AI Agents
DevOps
Incident Management
Linux
Unix
Cybersecurity
FedRAMP
Management
ServiceNow
Service Desk
Apply
$79k – $96k per year • Remote/Hybrid • 2+ years exp • Toronto
Cybersecurity
FedRAMP
Marketing
LinkedIn
Apply
$129k – $256k per year (Estimated) • Remote/Hybrid • 10+ years exp • McLean
Cybersecurity
FedRAMP
Marketing
LinkedIn
Apply
Remote/Hybrid • Full-Time • Huntsville
Cybersecurity
FedRAMP
Apply
$70k – $90k per year • In office • Full-Time • Jersey City
Apply
$140k – $194k per year • In office • Full-Time • 7+ years exp • Plano • Jersey City • Charlotte
Python
SQL
Databases
Databricks
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
FinOps
Analytics
Tableau
Power BI
Management
Agile
Apply
$96k – $200k per year (Estimated) • In office • 4+ years exp • Jersey City
Python
SQL
Databases
PostgreSQL
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
Terraform
GCP
CI/CD
AWS
Bitbucket
Management
Confluence
Jira
Agile
Apply
$166k – $307k per year (Estimated) • In office • 5+ years exp • Jersey City
Apply
$167k – $309k per year (Estimated) • In office • 8+ years exp • Jersey City
AI/ML
PyTorch
Ignite
Apply
See all jobs
This is one of many
687,351 more open roles from verified company boards, updated every day.