368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$132k – $268k per year (Estimated)
Location
In office (Plano)
Seniority
Staff
Overview
Company
Impact
Profile match
JPMorgan Chase & Co. is a leading global financial services firm and the largest banking institution in the United States by assets. Headquartered in New York City, the company offers a comprehensive range of financial solutions, including investment banking, asset management, treasury services, and commercial banking. Through its widely recognized consumer division, Chase, it delivers retail banking, credit card, and mortgage services to tens of millions of households across the globe.

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Technology team , you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.

Job Responsibilities

  • Consistently models and champions site reliability culture and practices, documents and shares knowledge within your organization via internal forums and communities of practice
  • Leads initiatives to improve the reliability and stability of your team’s applications and platforms using data-driven analytics to improve service levels, proactively identifying and solving technology-related bottlenecks in areas of expertise
  • Drives collaboration with your team to identify comprehensive service level indicators and the stakeholder partners to establish reasonable service level objectives and error budgets with your customers
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Serves as the main point of contact during major incidents for your application and has the skills to identify and solve the issue quickly to avoid financial loss to the business
  • Offers a high level of technical expertise within one or more technical domains and provides advice and mentorship to other engineers
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

    Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices
  • Fluent in at least one programming language such as: Python, Java/Spring Boot, .Net
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Proficient knowledge and experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection
  • Proficient with continuous integration and continuous delivery practices and tooling
  • Proficient with container and container orchestration
  • Experience with troubleshooting common networking technologies and issues
  • Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, and actively self-educates to evaluate and recommend suitable new technologies

    Preferred qualifications, capabilities, and skills

  • Experience implementing and managing SLOs/SLIs, error budgets, and operational readiness reviews for distributed systems, including leading post-incident analysis and resilience improvements.
  • Hands-on expertise in observability and monitoring tools including Grafana, Dynatrace, Prometheus, Datadog, and Splunk; experience with SLO alerting, white/black box monitoring, and telemetry collection.
  • Advanced proficiency in Python and/or Java for building automation, tooling, and operational workflows.
  • Experience with CI/CD tools and practices including Jenkins, GitLab, and Terraform; strong grasp of DevOps principles and continuous delivery pipelines.
  • Strong incident management experience; effective under pressure with excellent stakeholder communication and the ability to drive root-cause analysis and auto-remediation.
  • Familiarity with containerization and orchestration technologies such as Docker, Kubernetes, and ECS; hands-on experience with infrastructure-as-code is a plus.
  • Exposure to public cloud platforms (AWS or equivalent), including compute, storage, networking, messaging, and infrastructure automation (CloudFormation, Terraform) is desirable.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Plano
$148k – $286k per year (Estimated) • Remote/Hybrid • Contractor • 10+ years exp • Arlington
Python
AI/ML
Computer Vision
Embeddings
PyTorch
Time Series Forecasting
DevOps
AWS
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
$258k – $386k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
Go
Java
Python
Rust
Databases
DynamoDB
MySQL
PostgreSQL
Redis
DevOps
AWS
Incident Management
Kubernetes
Apply
$172k – $301k per year • Equity • In office • Full-Time • 8+ years exp • Minneapolis
C++
Go
Java
Python
AI/ML
AI Agents
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Semantic Search
Anthropic
Human-in-the-Loop
LLM Guardrails
OpenAI
Semantic Search
Structured Outputs
Function Calling
DevOps
Vector
Cybersecurity
Least Privilege
Management
ServiceNow
Apply
$91k – $194k per year (Estimated) • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • Brisbane
Python
SQL
Python
Django
FastAPI
Flask
Databases
Databricks
AI/ML
Function Calling
Hallucination
LLM
NumPy
Scikit-learn
Statsmodels
Streamlit
OpenAI Codex
AI Agents
RAG
DevOps
Azure
CI/CD
Git
Cybersecurity
HIPAA
Analytics
Plotly
QA
Pytest
Apply
$70k – $150k per year (Estimated) • Remote • 3+ years exp
Python
AI/ML
Red Teaming
Cybersecurity
OWASP Top 10
Apply
$157k – $313k per year (Estimated) • In office • Bachelor's Degree • Jersey City
Python
SQL
AI/ML
AI Agents
Analytics
Tableau
Apply
$171k – $307k per year (Estimated) • In office • Jersey City
Apply
$151k – $314k per year (Estimated) • In office • Jersey City
SQL
Databases
Databricks
Apply
$87k – $207k per year (Estimated) • In office • Bachelor's Degree • Jersey City
Python
Analytics
Tableau
Apply
$164k – $332k per year (Estimated) • In office • Jersey City
SQL
Databases
Snowflake
Analytics
Tableau
Apply
$143k – $274k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Antonio • Charlotte • Colorado Springs • Plano • Phoenix
DevOps
IAM
Cybersecurity
CyberArk
Microsoft Entra ID
PCI DSS
Robotics
Path Planning
Management
ServiceNow
Apply
$135k – $217k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Charlotte • Plano • Jersey City • Chicago
DevOps
Azure
Management
Confluence
Jira
Apply
$85k – $163k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Antonio • Plano
Robotics
Path Planning
Management
Jira
Apply
$62k – $126k per year (Estimated) • In office • 3+ years exp • Plano
Python
Cybersecurity
HIPAA
Okta
SOC 2
Management
Confluence
Google Workspace
Jira
n8n
Slack
Zapier
Apply
$111k – $200k per year (Estimated) • In office • Full-Time • 1+ year exp • Plano
JavaScript
Python
C#
C#
.NET
DevOps
Azure
CI/CD
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.