430,068open jobs
14,667companies
59,611added this week
Browse all
Salary
$202k – $424k per year (Estimated)
Location
In office (Palo Alto)
Seniority
Staff · 5+ years exp
Overview
Company
Impact
Profile match
JPMorganChase is the largest bank in the United States by assets and one of the most systemically important financial institutions in the world, with a lineage running back through more than a thousand predecessor firms to the 1799 founding of the Bank of the Manhattan Company. It combines a dominant investment bank and markets business with Chase, the largest retail banking franchise in America, plus commercial banking and asset and wealth management. Headquartered in New York, the group is unusual among banks for the scale of its technology spending, running one of the largest engineering organisations of any financial institution and deploying its own internal AI platform across the firm.

We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.

As a Lead Software Engineer at JPMorganChase within AI/ML Data Platforms, you are an integral part of an agile team that works to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments.

Job Responsibilities

  • Design, build, and maintain end-to-end ML training platform.

  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.

  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.

  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.

  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.

  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)

  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow

  • Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.

  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Required qualifications, capabilities, and skills

  • Formal training or certification on software engineering concepts and 5+ years applied experience.

  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.

  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).

  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).

  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).

  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.

  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).

  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).

  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)

  • Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security.

  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices.

Preferred qualifications, capabilities, and skills

  • Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments.

  • Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management.

  • Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets.

  • Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation.

  • Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts.

  • Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning).

  • Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning.

FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
430,068 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
In office
Python
JavaScript
SQL
Databases
Google BigQuery
BigQuery
DevOps
CI/CD
Docker
GitHub
Apply
Software Engineer II 4 hours ago
$67k – $160k per year (Estimated) • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • Cork
Python
JavaScript
TypeScript
C++
TCL Scripting
Databases
Weaviate
Chroma
Pinecone
AI/ML
Copilot
Claude
Claude Code
Fine-tuning
Prompt Engineering
AI Agents
LLM
RAG
GPT-4
DevOps
GCP
Azure
Git
AWS
Vector
GitHub
Management
Confluence
Apply
$151k – $269k per year • Remote • Top Secret • 11+ years exp • Bachelor's Degree
DevOps
CI/CD
Management
ServiceNow
QA
Selenium
Apply
Software Engineer 4 hours ago
$75k – $80k per year • In office • Bachelor's Degree
Python
SQL
C#
Visual Basic
C#
.NET
DevOps
Azure DevOps
Azure
Apply
$75k – $80k per year • In office • Bachelor's Degree • Eden Prairie
Python
C++
Apply
$47k – $87k per year (Estimated) • In office • 2+ years exp • High School Diploma • Long Beach
Apply
$203k – $377k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • New York
Python
SQL
SAS
Analytics
Tableau
Alteryx
Management
Jira
Apply
In office • 5+ years exp
Python
Java
C#
Java
Spring Boot
C#
.NET
DevOps
Ansible
CI/CD
SRE
Self-Healing
Apply
$77k – $170k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Newark
Databases
Databricks
Apply
$195k – $392k per year (Estimated) • In office • 5+ years exp • New York
Apply
Software Architect 9 hours ago
$218k – $365k per year • In office • Full-Time • 15+ years exp • PhD • San Francisco • Palo Alto • Washington
AI/ML
Copilot
Cursor
Claude
Claude Code
Prompt Engineering
AI Agents
OpenAI Codex
Agentforce
DevOps
GCP
AWS
GitHub
Apply
$165k – $210k per year • Remote/Hybrid • 5+ years exp • Bachelor's Degree • Palo Alto
AI/ML
AI Agents
Design
Figma
Canva
Apply
$183k – $215k per year • In office • Full-Time • Palo Alto
Databases
Redis
Apache Kafka
Frontend
GraphQL
DevOps
gRPC
API Gateway
IoT
MQTT
CoAP
Apply
$68k – $76k per year • In office • Full-Time • High School Diploma • Palo Alto
Cybersecurity
HIPAA
Apply
$229k – $421k per year (Estimated) • Remote/Hybrid • 10+ years exp • Bachelor's Degree • Palo Alto
Python
Java
AI/ML
Airflow
Jupyter Notebook
AI Agents
PyTorch
Ignite
Apply
See all jobs
This is one of many
430,068 more open roles from verified company boards, updated every day.