368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office
Seniority
Senior
Overview
Company
Impact
Profile match
JPMorgan Chase & Co. is a leading global financial services firm and the largest banking institution in the United States by assets. Headquartered in New York City, the company offers a comprehensive range of financial solutions, including investment banking, asset management, treasury services, and commercial banking. Through its widely recognized consumer division, Chase, it delivers retail banking, credit card, and mortgage services to tens of millions of households across the globe.

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.

The Chief Data & Analytics Office (CDAO) at JPMorgan Chase is responsible for accelerating the firm’s data and analytics journey. This includes ensuring the quality, integrity, and security of the company's data, as well as leveraging this data to generate insights and drive decision-making. The CDAO is also responsible for developing and implementing solutions that support the firm’s commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly.

As a Senior Manager - SRE at JPMorgan Chase within the AIML Data Platforms and Chief Data and Analytics Team, you will lead a team of 3-7 SRE engineers and scale the AI/ML Data Platform, deliver advanced technology products focused on data and analytics. You will tackle complex cloud data platform challenges, especially around Data Lake Tools. In this role you will work in an agile environment, collaborating with cross-functional teams.

Job Responsibilities:

  • Leads a team of Site Reliability Engineers and support critical application 24x7
  • Implements Site Reliability Engineering (SRE) best practices to ensure reliability, scalability, and performance of data platforms.
  • Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems.
  • Scale and maintain a managed AWS Databricks platform, and provides engineering and operational support for the platform to Application/Engineering teams.
  • Drives reuse-first adoption of enterprise-authorized AI capabilities within the work environment to improve reliability operations and customer experience outcomes, with human-in-the-loop validation and appropriate handling of sensitive data.
  • Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture.
  • Drives continuous improvement in system observability, alerting, and capacity planning.
  • Collaborates with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Performs platform design, set-up and configuration, workspace administration, resource monitoring, providing engineering support to Data Engineering teams, Data Science/ML, and Application/Integration teams.
  • Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems.
  • Develops secure high-quality production code, and reviews and debugs code written by others.
  • Adds to team culture of diversity, opportunity, and respect.
  • Develops and maintains incident response procedures, including root cause analysis and postmortem documentation.

Required Qualifications, Capabilities, and Skills:

  • Experience with formal training or certification on software engineering concepts.
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management.
  • Ability to manage highly technical team of Site Reliability Engineers.
  • Experience leading teams in the safe use of enterprise-authorized AI capabilities within the work environment for reliability engineering workflows, including validation habits and awareness of data sensitivity.
  • Ability to set and reinforce organization-level practices for reviewing AI-assisted recommendations and escalating uncertain decisions while maintaining resiliency, security, and auditability outcomes.
  • Extensive experience with AWS, Databricks, or Snowflake platform administration and engineering support is a MUST.
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines.
  • Proficient in Python application program development with use of automated unit testing.
  • Experience with Terraform development and understanding of Terraform enterprise.
  • Experience in delivering system design, application development, testing, and operational stability.
  • Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce etc.
  • Excellent troubleshooting, analytical, and communication skills.

Preferred Qualifications, Capabilities, and Skills:

  • Multi Region disaster recovery setup, monitoring and testing.
  • Experience in Data pipelines using Spark.
  • Exposure to AWS & Databricks Platform administration.
  • Knowledge of containerization (Docker, Kubernetes) and orchestration.
  • Familiarity with distributed systems and large-scale data processing.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$148k – $286k per year (Estimated) • Remote/Hybrid • Contractor • 10+ years exp • Arlington
Python
AI/ML
Computer Vision
Embeddings
PyTorch
Time Series Forecasting
DevOps
AWS
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
$258k – $386k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
Go
Java
Python
Rust
Databases
DynamoDB
MySQL
PostgreSQL
Redis
DevOps
AWS
Incident Management
Kubernetes
Apply
$184k – $350k per year (Estimated) • Remote/Hybrid • Internship • 15+ years exp • Bachelor's Degree • Arlington
AI/ML
Computer Vision
AI Agents
DevOps
AWS
Platform Engineering
Cybersecurity
GDPR
Apply
$173k – $369k per year (Estimated) • Remote/Hybrid • Internship • 15+ years exp • Bachelor's Degree • San Francisco
AI/ML
Computer Vision
AI Agents
DevOps
AWS
Platform Engineering
Cybersecurity
GDPR
Apply
$145k – $295k per year (Estimated) • Equity • Remote • Internship • 8+ years exp • San Francisco
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$157k – $313k per year (Estimated) • In office • Bachelor's Degree • Jersey City
Python
SQL
AI/ML
AI Agents
Analytics
Tableau
Apply
$171k – $307k per year (Estimated) • In office • Jersey City
Apply
$151k – $314k per year (Estimated) • In office • Jersey City
SQL
Databases
Databricks
Apply
$87k – $207k per year (Estimated) • In office • Bachelor's Degree • Jersey City
Python
Analytics
Tableau
Apply
$164k – $332k per year (Estimated) • In office • Jersey City
SQL
Databases
Snowflake
Analytics
Tableau
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.