368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$217k – $304k per year
Location
Remote (United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer, Ads based in United States.

This is a senior technical leadership role focused on strengthening the reliability of a large-scale advertising ecosystem.

You’ll shape reliability strategy across critical systems spanning ad serving, auctions, targeting, reporting, measurement, and billing.

The role combines hands-on engineering with architecture, automation, incident leadership, and long-term platform resilience.

You’ll partner with engineering leaders and multiple technical teams to influence roadmaps and improve operational excellence.

Your work will directly support highly available, low-latency systems where reliability and performance have meaningful business impact.

You’ll also mentor engineers and help establish measurable reliability practices across a broad, distributed technology environment.

The position offers significant ownership and the opportunity to define how reliability engineering evolves at organizational scale.

Accountabilities

    • Lead reliability initiatives across critical advertising domains, including ad serving, auctions, targeting, reporting, measurement, attribution, and billing.
    • Partner with engineering leadership to establish and execute roadmaps focused on reliability, scalability, operational excellence, and developer productivity.
    • Design and build scalable platforms, tooling, automation, and infrastructure capabilities that improve system resilience and engineering efficiency.
    • Lead architecture reviews and influence technical decisions for high-traffic, revenue-critical distributed systems.
    • Establish and monitor reliability metrics and SLOs around critical advertiser and platform journeys, using data to identify risks and prioritize improvements.
    • Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses to major production events.
    • Identify systemic reliability risks and implement durable solutions that improve availability, performance, scalability, and operational maturity.
    • Drive automation, observability, incident management, performance optimization, and other practices that strengthen production reliability.
    • Mentor engineers and provide technical leadership across multiple teams, helping raise engineering standards and reliability expertise.
    • Collaborate with Product, Data Science, Infrastructure, and Engineering stakeholders to ensure reliability considerations are embedded into product and infrastructure investments.
    • Requirements

      • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
      • Proven experience evolving and supporting high-traffic, user-facing production environments with demanding availability and performance requirements.
      • Deep expertise in distributed systems, scalability engineering, cloud-native architectures, and highly available system design.
      • Strong software engineering capabilities, ideally with experience in backend programming languages such as Go.
      • Extensive knowledge of observability practices and technologies, including metrics, logging, tracing, alerting, and performance monitoring.
      • Demonstrated experience improving reliability through SLOs, automation, incident management, performance optimization, and systematic operational practices.
      • Strong troubleshooting and problem-solving abilities across complex, modern distributed technology stacks.
      • Excellent cross-functional communication and collaboration skills, with the ability to influence technical direction and align teams around shared reliability goals.
      • Experience supporting advertising technology or other large-scale, revenue-critical platforms is highly desirable.
      • Familiarity with reliability challenges involving ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing is a strong advantage.
      • Experience operating high-QPS, low-latency services where system performance directly affects business outcomes is preferred.
      • Experience establishing reliability programs with measurable business and operational results is a plus.
      • Hands-on experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is beneficial.
      • Familiarity with technologies such as Kafka, ClickHouse, Spark, Flink, or BigQuery is advantageous.
      • Experience partnering with Product, Data Science, and advertising engineering teams, or supporting machine learning inference and recommendation systems at scale, is a plus.
      • Benefits

        • Base salary range of $217,000-$303,900 USD, with final compensation determined by factors such as skills, experience, credentials, and role level.
        • Eligibility for equity in the form of restricted stock units.
        • Comprehensive health benefits, including medical, dental, and vision coverage.
        • 401(k) program with employer matching.
        • Workspace benefits and support for a home office.
        • Personal and professional development funds.
        • Family planning support.
        • Flexible vacation and global days off.
        • 4+ months of paid parental leave.
        • Paid volunteer time off.
        • Flexible-first remote work environment within the United States.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$18k – $39k per year (Estimated) • In office • Full-Time • 3+ years exp • Kolkata
DevOps
Azure
Azure DevOps
Incident Management
Apply
In office • Full-Time • Bachelor's Degree • Bengaluru
DevOps
Incident Management
Apply
In office • Full-Time • Bachelor's Degree • Bengaluru
DevOps
Incident Management
Apply
$14k – $36k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Bengaluru
DevOps
Incident Management
Apply
$19k – $41k per year (Estimated) • In office • Full-Time • 3+ years exp • Gurgaon
DevOps
Incident Management
Apply
$152k – $229k per year • Remote • Full-Time
AI/ML
Human-in-the-Loop
Apply
$162k – $180k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Cybersecurity
HIPAA
Zero Trust
Analytics
ETL/ELT
Apply
$14k – $32k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree
DevOps
Incident Management
Management
ServiceNow
Apply
$29k – $60k per year (Estimated) • Remote • Full-Time • 8+ years exp
DevOps
Azure
Azure DevOps
Apply
$26k – $69k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.