691,154open jobs
40,551companies
97,801added this week
Browse all
Salary
$140k – $281k per year (Estimated)
Location
In office (Jersey City)
Seniority
Staff · 5+ years exp
Overview
Company
Impact
Profile match
JPMorganChase is the largest bank in the United States by assets and one of the most systemically important financial institutions in the world, with a lineage running back through more than a thousand predecessor firms to the 1799 founding of the Bank of the Manhattan Company. It combines a dominant investment bank and markets business with Chase, the largest retail banking franchise in America, plus commercial banking and asset and wealth management. Headquartered in New York, the group is unusual among banks for the scale of its technology spending, running one of the largest engineering organisations of any financial institution and deploying its own internal AI platform across the firm.

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the AI Machine Learning and Data platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.

Job Responsibilities

  • Design and implement solutions to enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast growing demands.
  • Partner with product engineering teams to ensure the AI/ML systems are reliable and high performing.
  • Develop observability, security, automation and fin-ops tools and orchestration.
  • Provide strategic technology leadership by defining and evaluating standards and architecture for reliability, observability and automation frameworks.
  • Build strong cross-functional relationships that foster engagements across the organization and deliver solutions to user problems.
  • Debug and solve issues in a production environment, identify root cause and remediate.
  • Participates in on-call rotations, incident management and escalation workflows.
  • Take full ownership of problems, develop solutions, and acquire new knowledge to complete the task.
  • Mentor and guide junior engineers
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience.
  • Expertise in SRE principles, reliability, scalability and performance of application and infrastructure.
  • Expertise in programming with Python and Infrastructure as Code, tools such as Terraform.
  • Demonstrable experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Experience in architecting distributed systems and cloud-native architecture in AWS.
  • Systematic problem-solving and troubleshooting skills in a complex system. Excellent communication skills and ability to represent and present business and technical concepts to stakeholders.

Preferred qualifications, capabilities, and skills

  • • Prior experience working in AI, ML, or Data engineering.
  • Expertise in container orchestration/Kubernetes.
  • • Prior experience developing Automation frameworks/AI Ops
  • • Prior experience building observability and telemetry tools.
  • • Previous experience as an SRE in a dynamic technology company or startup
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
691,154 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Jersey City
Manager Support 4 hours ago
$37k – $83k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Mexico
DevOps
Azure
AWS
Apply
$70k – $172k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • South San Francisco
C
C
MPI
AI/ML
NCCL
InfiniBand
DevOps
Terraform
Ansible
SLURM
Docker
Kubernetes
HPC
Apply
AI Agent Developer 4 hours ago
$84k – $226k per year (Estimated) • Remote • Tbilisi
Python
Java
C#
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
LLM
Structured Outputs
LLM Guardrails
Tool Use
Game Dev
Unity
Apply
AI Agent Developer 4 hours ago
$21k – $56k per year (Estimated) • Remote • Tashkent
Python
Java
C#
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
LLM
Structured Outputs
LLM Guardrails
Tool Use
Game Dev
Unity
Apply
AI Agent Developer 4 hours ago
$21k – $56k per year (Estimated) • Remote • Astana
Python
Java
C#
AI/ML
Model Context Protocol
Embeddings
Function Calling
AI Agents
LLM
Structured Outputs
LLM Guardrails
Tool Use
Game Dev
Unity
Apply
$21k – $62k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
Go
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Kubernetes
Grafana
Amazon EKS
GitHub
Cybersecurity
SonarQube
Management
Slack
Agile
Apply
$81k – $209k per year (Estimated) • In office • Orlando
Apply
$33k – $86k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
SQL
AI/ML
LangGraph
AutoGen
LangChain
LlamaIndex
Embeddings
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Hallucination
Reranking
Structured Outputs
DevOps
Azure
CI/CD
AWS
Vector
Cybersecurity
Least Privilege
Management
Agile
Apply
$29k – $66k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Grafana
GitHub
Cybersecurity
SonarQube
Management
Slack
Confluence
Jira
Agile
Apply
$70k – $90k per year • In office • Full-Time • Jersey City
Apply
$140k – $194k per year • In office • Full-Time • 7+ years exp • Plano • Jersey City • Charlotte
Python
SQL
Databases
Databricks
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
FinOps
Analytics
Tableau
Power BI
Management
Agile
Apply
$96k – $200k per year (Estimated) • In office • 4+ years exp • Jersey City
Python
SQL
Databases
PostgreSQL
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
Terraform
GCP
CI/CD
AWS
Bitbucket
Management
Confluence
Jira
Agile
Apply
$166k – $307k per year (Estimated) • In office • 5+ years exp • Jersey City
Apply
$167k – $309k per year (Estimated) • In office • 8+ years exp • Jersey City
AI/ML
PyTorch
Ignite
Apply
See all jobs
This is one of many
691,154 more open roles from verified company boards, updated every day.