368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$117k – $228k per year (Estimated)
Location
Remote/Hybrid (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States.

This is an opportunity to join a critical AI Hardware SRE team responsible for the reliability of next-generation dedicated AI infrastructure.

You will help scale and optimize high-density hardware and software environments across regional data centers.

The role combines automation, observability, infrastructure engineering, networking, and real-time incident response.

You will build Python-based tooling, infrastructure-as-code utilities, telemetry pipelines, and intelligent monitoring solutions.

Your work will directly improve uptime, performance, scalability, and operational efficiency for business-critical systems.

You will collaborate with engineering teams, infrastructure vendors, and field technicians to solve complex reliability challenges.

This role is ideal for an experienced SRE who thrives on ownership, ambiguity, automation, and production-scale infrastructure.

Accountabilities

    • Develop and scale robust Python-based tooling, infrastructure-as-code utilities, and automation frameworks to eliminate operational toil and streamline fleet-wide provisioning.
    • Build automated workflows and API integrations across corporate ticketing systems to accelerate resolution of hardware and network incidents.
    • Apply modern AI and LLM-based development tools to improve technical execution, automate scripting, and evaluate complex infrastructure systems.
    • Work with advanced private cloud and compute technologies to improve availability, latency, scalability, and overall health across high-density hardware environments.
    • Design and implement telemetry pipelines, Prometheus and Grafana dashboards, and AI-driven anomaly detection for bare-metal and virtualized infrastructure.
    • Define operational KPIs, monitoring standards, telemetry baselines, alerting thresholds, and operational readiness criteria for new services and infrastructure deployments.
    • Participate in a 24x7x365 on-call rotation, leading real-time incident response and managing high-severity service disruptions through automated PagerDuty and Slack workflows.
    • Develop detailed technical runbooks, lead incident response bridges, and drive blameless post-mortems that identify systemic improvements and prevent recurring issues.
    • Partner with infrastructure vendors and coordinate on-site field technicians to support hardware reliability, break-fix activities, and uptime objectives.
    • Collaborate across engineering and infrastructure teams to identify reliability gaps, establish best practices, and deliver production-grade solutions to ambiguous technical challenges.
    • Requirements

      • 5+ years of relevant Site Reliability Engineering, infrastructure engineering, systems engineering, or related experience, along with a Bachelor’s degree in Computer Science or a related technical field.
      • Exceptional proficiency in Python and experience developing scalable operational tooling, API integrations, automation frameworks, and infrastructure utilities.
      • Hands-on experience with modern observability technologies such as Prometheus, Grafana, OpenTelemetry, and Loki, as well as familiarity with time-series monitoring and telemetry systems.
      • Strong understanding of advanced networking concepts, including high-bandwidth routing and switching, BGP, and dual-stack IPv4/IPv6 environments.
      • Experience designing and launching new services with clear operational readiness requirements, telemetry baselines, monitoring strategies, and alerting thresholds.
      • Extensive experience creating technical runbooks, leading complex incident response processes, and conducting comprehensive, blameless post-mortems.
      • Strong understanding of distributed infrastructure, high-density compute environments, private cloud technologies, and large-scale content or infrastructure delivery challenges.
      • Ability to leverage AI-assisted development tools and LLM-based approaches to accelerate engineering workflows and solve technical problems effectively.
      • Proven ability to take ownership of ambiguous and complex technical challenges, coordinate cross-functional teams, and drive solutions through to production.
      • Strong communication and collaboration skills, with the ability to work effectively with engineering teams, vendors, and field operations.
      • Willingness to participate in a 24x7x365 on-call rotation and respond effectively to high-severity production incidents.
      • Benefits

        • Comprehensive benefits designed to support employee health, well-being, financial security, and life beyond work.
        • Flexible working options that allow employees to work from home, in an office, or through a combination of both, depending on role and business needs.
        • Opportunity to work on cutting-edge AI hardware, private cloud, distributed infrastructure, and edge technologies.
        • Exposure to large-scale, business-critical systems serving global digital experiences.
        • Collaborative environment with opportunities to work alongside experienced infrastructure, engineering, and technology professionals.
        • Opportunities to develop expertise in SRE, observability, automation, AI-assisted engineering, networking, and high-density compute.
        • Support for professional growth and continued development within a technology-focused environment.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$10k – $34k per year (Estimated) • In office • Tashkent
Java
Kotlin
Java
Hibernate
Maven
Spring Boot
Kotlin
Mockito
Databases
PostgreSQL
DevOps
CI/CD
Docker
Git
Grafana
Kubernetes
Loki
Prometheus
GitLab
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
$17k – $47k per year (Estimated) • Remote/Hybrid • Internship • 5+ years exp • Minsk
Go
JavaScript
Python
AI/ML
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
Prompt Engineering
RAG
Anthropic
Function Calling
OpenAI
Frontend
React.js
DevOps
CI/CD
Docker
Git
Terraform
GitHub
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.