371,660open jobs
9,621companies
49,388added this week
Browse all
Salary
$152k – $309k per year (Estimated)
Location
In office (Plano)
Seniority
Staff
Overview
Company
Impact
Profile match
JPMorgan Chase & Co. is a leading global financial services firm and the largest banking institution in the United States by assets. Headquartered in New York City, the company offers a comprehensive range of financial solutions, including investment banking, asset management, treasury services, and commercial banking. Through its widely recognized consumer division, Chase, it delivers retail banking, credit card, and mortgage services to tens of millions of households across the globe.

Are you passionate about building innovative technology that powers AI and machine learning across a global organization? As part of our team, you’ll help shape the future of model deployment at scale, collaborating with talented engineers and data scientists. You’ll have the opportunity to work on impactful projects, grow your skills, and contribute to a platform that drives real business outcomes. We value creativity, collaboration, and a commitment to excellence.

As a Lead Software Engineer at JPMorgan Chase within Firmwide AI/ML Deployment Platform team, you will work closely with engineers to design, build, and deploy an AI solution that unifies observability data across multi-cloud environments (AWS, Azure, GCP). You will create intelligent systems that correlate cross-platform health metrics, logs, and traces to generate actionable troubleshooting recommendations and automations, enabling self-service remediation, predicting system outages before they occur, and directly reducing ticket volume by identifying and resolving repeated IT incidents. This role bridges the gap between machine learning, multi-cloud infrastructure, and automated IT operations to build predictive, self-healing solutions.

Job responsibilities

  • Build and deploy infrastructure solutions for seamless integration of control plane and user accounts
  • Design pipelines to ingest, aggregate, and correlate telemetry data (metrics, logs, traces) from multi-cloud infrastructures.
  • Architect and implement closed-loop automation playbooks that allow infrastructure to auto-remediate common, repeatable failure modes without human intervention.
  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
  • Build and operationalize LLM/machine learning models for anomaly detection, predictive health monitoring, and forecasting system degradations.
  • Integrate the AI engine with ticket data and map observability insights against ticket trends, cluster repetitive issues, and quantify the platform's impact on reducing Mean Time to Resolution (MTTR).
  • Create user-friendly self-service portals or conversational AI interfaces that allow non-expert teams to diagnose and fix infrastructure issues safely.

Required qualifications, capabilities, and skills

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • AI/ML & Data Science: Strong proficiency in Python, Java alongside experience integrating LLM / ML models. Familiarity with time-series forecasting data analysis pipelines and Natural Language Processing (NLP) for log/ticket clustering is essential. Good understanding of agentic AI concepts (A2A, MCPs, Skills, RAG, etc)
  • Automation & Orchestration: Advanced experience with configuration management tools and automated workflow engines.
  • Integration: Hands-on experience building custom webhooks, APIs, and integrations with ticketing systems like ServiceNow or Jira Service Management.
  • Big Data Pipelines: Competency in managing large-scale, streaming data infrastructure using cloud-native data warehouses (e.g. Snowflake)
  • Cloud & Infrastructure: Expertise in multi-cloud architectures across Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) including on-prem.
  • Observability Frameworks: Experience with enterprise observability stacks such as OpenTelemetry, Prometheus, Dynatrace
  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices

Preferred qualifications, capabilities, and skills

  • Practical experience applying generative AI and agentic workflows to accelerate development (e.g., AI-assisted code and test generation, refactoring, documentation), with strong judgment, governance, and quality control over AI-produced outputs.
  • Experience optimizing performance and reliability of AI-powered user interfaces & proficiency in React framework
  • Knowledge of multi clouds domains
  • Experience working with LLMs
  • Experience in API development and design
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
371,660 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Plano
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
$150k – $180k per year • In office • Full-Time • PhD • New York
Python
AI/ML
Anthropic
Anthropic SDK
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
OpenAI
OpenAI SDK
RAG
DevOps
AWS
Azure
GCP
Apply
$38k – $91k per year (Estimated) • Remote • Full-Time • 8+ years exp • PhD • Guadalajara
PowerShell
Python
Databases
Amazon Aurora
DynamoDB
AI/ML
Amazon SageMaker
AWS Bedrock
AWS Bedrock AgentCore
Ray
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
Datadog
FinOps
GCP
Git
GitLab
GitLab CI
IAM
Jenkins
JFrog Artifactory
Kubernetes
New Relic
Service Mesh
Splunk
Terraform
Cybersecurity
HIPAA
ISO 27001
PCI DSS
SOC 2
Apply
$256k – $320k per year • Remote/Hybrid • Full-Time • Seattle
Go
Python
AI/ML
AI Agents
CrewAI
LangChain
LangGraph
DevOps
AWS
GCP
Kubernetes
Apply
$54k – $128k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Mexico City
C#
JavaScript
Python
TypeScript
C#
ASP.NET Core
Python
FastAPI
Databases
PostgreSQL
AI/ML
AI Agents
Anthropic
Embeddings
Function Calling
LLM
LLM Guardrails
Model Context Protocol
OpenAI
Semantic Search
Semantic Search
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
GitHub
GitHub Actions
Vector
Vercel
Apply
$162k – $328k per year (Estimated) • In office • Seattle
Java
TypeScript
JavaScript
Java
Maven
Spring Boot
Frontend
React Query
Redux
Zustand
React.js
Mobile
State Management
DevOps
AWS
CI/CD
GCP
Git
Jenkins
Apply
$143k – $287k per year (Estimated) • In office • South San Francisco
Apply
$157k – $282k per year (Estimated) • In office • Tampa
Apply
$139k – $282k per year (Estimated) • In office • Bachelor's Degree • Jersey City
Python
DevOps
Ansible
Platform Engineering
Cybersecurity
Wireshark
Apply
$93k – $195k per year (Estimated) • In office • Plano
PowerShell
Python
C#
Java
C#
.NET
Java
Apache Tomcat
AI/ML
AI Agents
Copilot
DevOps
Amazon CloudWatch
Ansible
AWS
Azure
CI/CD
Datadog
Dynatrace
GCP
GitHub
Grafana
Prometheus
Splunk
Terraform
Apply
$115k – $207k per year (Estimated) • In office • Full-Time • Plano • Charlotte
SQL
Databases
CockroachDB
MS SQL
DevOps
Ansible
AWS
Bitbucket
CI/CD
GitHub
Jenkins
Terraform
Cryptography
Vault
Apply
$194k – $368k per year (Estimated) • In office • 10+ years exp • Plano
AI/ML
AI Agents
Human-in-the-Loop
LLM Guardrails
DevOps
Platform Engineering
Apply
$167k – $300k per year (Estimated) • In office • Plano
Databases
Databricks
DevOps
IAM
Apply
$136k – $295k per year (Estimated) • In office • Plano
Python
DevOps
Akamai
Amazon EC2
API Gateway
AWS
AWS Step Functions
CI/CD
Cloudflare
Terraform
Cybersecurity
AWS WAF
Threat Modeling
Apply
$143k – $274k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Antonio • Charlotte • Colorado Springs • Plano • Phoenix
DevOps
IAM
Cybersecurity
CyberArk
Microsoft Entra ID
PCI DSS
Robotics
Path Planning
Management
ServiceNow
Apply
See all jobs
This is one of many
371,660 more open roles from verified company boards, updated every day.