678,285open jobs
39,280companies
100,060added this week
Browse all
Salary
$138k – $278k per year (Estimated)
Location
In office (Jersey City)
Seniority
Staff · 5+ years exp
Overview
Company
Impact
Profile match
JPMorganChase is the largest bank in the United States by assets and one of the most systemically important financial institutions in the world, with a lineage running back through more than a thousand predecessor firms to the 1799 founding of the Bank of the Manhattan Company. It combines a dominant investment bank and markets business with Chase, the largest retail banking franchise in America, plus commercial banking and asset and wealth management. Headquartered in New York, the group is unusual among banks for the scale of its technology spending, running one of the largest engineering organisations of any financial institution and deploying its own internal AI platform across the firm.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Public Cloud team, you will blend hands-on engineering with program leadership to promote platform stability, ensure consistent execution across SRE teams, and partner closely with Engineering and Product to deliver measurable improvements in availability, support outcomes, and cost of failure.

Job Responsibilities

  • Drive consistency across SRE teams: establish and scale a “gold standard” for SLOs, on-call readiness, runbooks, postmortems, action tracking, and operational readiness across AWS/Azure/GCP.
  • Partner with Engineering and Product: embed reliability outcomes into roadmaps and delivery plans; convert incidents, support signals, and error budget trends into prioritized backlog with measurable impact.
  • Operational excellence & BPMs: design and run operating cadences (incident/stability reviews, KPI reviews, planning inputs), standardize intake/prioritization, and ensure closed-loop execution.
  • Risk governance: align reliability operations to risk/control expectations; operationalize recurring remediations (e.g., configuration drift, repeat findings) through centralized automation without sacrificing velocity.
  • Ticket reduction & automation opportunities: identify top drivers of support load and operational toil; build and maintain an automation opportunity pipeline; track adoption and deflection.
  • Platform stability metrics: own cross-platform reporting (SLO attainment, incident trends, MTTR/MTTI, change failure rate, ticket deflection, customer impact/cost of failure).
  • AI for reliability operations: apply AI/LLMs to improve triage, incident summarization, correlation/symptom mapping, and guardrailed automation; measure accuracy, safety, and outcomes.
  • Hands-on leadership: stay close to designs and critical implementations; lead systemic remediation and major incident response improvements.

Required qualifications, skills, and capabilities

  • 5+ years in SRE / production engineering / platform reliability / infrastructure operations at enterprise scale
  • Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms.
  • Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability.
  • Strong experience partnering with Engineering and Product leadership to align priorities and deliver results.
  • Hands-on experience with AWS and/or Azure and automation/IaC fundamentals (e.g., Python/Go/Bash, Terraform, CI/CD).
  • Experience using AI/LLM-enabled approaches in operations (AIOps, AI-assisted troubleshooting, agentic workflows) with appropriate controls and measurement.

Preferred Qualifications

  • Improved SLO attainment and reduced Sev1/Sev2 frequency; fewer repeat incidents.
  • Reduced MTTR/MTTI and lower change failure rate; reduced customer minutes impacted (cost of failure).
  • Measurable ticket reduction/deflection via scaled automation and self-service adoption.
  • Consistent operating model adopted across SRE teams with trusted executive reporting.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
678,285 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Jersey City
$153k – $180k per year • In office • 5+ years exp
Python
Databases
Databricks
AI/ML
LangChain
Spark
MLFlow
SHAP
Fine-tuning
Scikit-learn
Arize Phoenix
LIME
TensorFlow
PyTorch
OpenAI
Hugging Face
Amazon SageMaker
DevOps
GCP
Azure
AWS
Vector
Apply
$152k per year • In office • 5+ years exp • Bachelor's Degree
Python
SQL
Python
pySpark
Databases
Apache Kafka
AI/ML
Hadoop
Spark
DevOps
CI/CD
Docker
Analytics
ETL/ELT
Apply
$107k – $263k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bachelor's Degree • Heredia
DevOps
GCP
Azure
AWS
Cybersecurity
NIST CSF
Zero Trust
Apply
$126k – $228k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
Frontend
Angular
Mobile
MAUI
Xamarin
DevOps
Rest API
Azure
CI/CD
Git
AWS
Apply
Applied ML Scientist 2 months ago
$190k – $250k per year • Remote/Hybrid • 30+ years exp • Master's Degree • New York
Python
C++
MATLAB
Apply
In office • 5+ years exp
Apply
In office • 5+ years exp
Python
Java
Java
Spring Framework
Databases
Cassandra
DynamoDB
DevOps
Ansible
SaltStack
Configuration Management
Management
Agile
Apply
In office • 5+ years exp • Bachelor's Degree
Apply
$152k – $296k per year (Estimated) • In office • 5+ years exp • Jersey City
JavaScript
Java
TypeScript
SQL
COBOL
Java
Spring Boot
COBOL
IBM MQ
Databases
Redis
Apache Kafka
DevOps
Rest API
AWS
Kubernetes
Management
Agile
QA
Swagger
Apply
$162k – $304k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Houston
Visual Basic
Analytics
Microsoft Excel
Apply
Cybersecurity Manager 9 hours ago
$117k – $252k per year (Estimated) • Remote/Hybrid • Contractor • 7+ years exp • Bachelor's Degree • Jersey City
DevOps
CI/CD
Cybersecurity
HIPAA
Apply
$91k – $183k per year (Estimated) • Remote • Contractor • 6+ years exp • Bachelor's Degree • Jersey City
Python
Perl
DevOps
Red Hat
Docker
Kubernetes
Apply
$119k – $237k per year (Estimated) • In office • 3+ years exp • Jersey City
Python
Java
DevOps
CI/CD
AWS
Management
Agile
Apply
$153k – $298k per year (Estimated) • In office • 5+ years exp • Jersey City
Python
Java
DevOps
CI/CD
Management
Agile
Apply
$55k – $107k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Jersey City
AI/ML
Copilot
DevOps
Azure
AWS
Management
ServiceNow
Apply
See all jobs
This is one of many
678,285 more open roles from verified company boards, updated every day.