1,421,583open jobs
83,066companies
211,600added this week
Browse all
Salary
≈ $151k – $300k per year (Estimated)
Location
In office (Palo Alto)
Seniority
Middle · 3+ years exp

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Aug 20, 2026. The Infatuation scores A on the Alion truth index.

Overview
Company
Impact
Profile match
The Infatuation is a restaurant discovery and review publication headquartered in New York, publishing guides, ratings and recommendations on where to eat in cities across the US, the UK and beyond. Founded in 2009, it has been owned by JPMorgan Chase since 2021 and runs editorial teams in cities such as New York and London, with its content also tied to Chase dining benefits. Its own hiring centres on staff writers, editors, video and social producers, audience and partnerships staff, and product and engineering roles for its guides and app.

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level.

As a Software Engineer III at JPMorganChase within the AI/ML data platform team you serve as a seasoned member of an agile team to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments

Job Responsibilities

  • Design, build, and maintain end-to-end ML training platform.

  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.

  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.

  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.

  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.

  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls)

  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow

  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.

  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.

Required qualifications, capabilities, and skills

  • Formal training or certification on software engineering concepts and 3+ years applied experience

  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.

  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).

  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).

  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).

  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.

  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).

  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).

  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)

  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.

  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Preferred qualifications, capabilities, and skills

  • Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments.

  • Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management.

  • Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets.

  • Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation.

  • Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts.

  • Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning).

  • Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning.

FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,421,583 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Palo Alto
AI Software Engineer 2 hours ago
$130k – $136k per year • Remote (United States) • Contractor • 5+ years exp • Winston-Salem
Python
Java
AI/ML
LangChain
LlamaIndex
AI Agents
LLM
RAG
Voice Agents
Tool Use
DevOps
Azure
CI/CD
Cybersecurity
HIPAA
Apply
≈ $113k – $224k per year (Estimated) • In office • Miramar
Python
JavaScript
TypeScript
Node JS
AI/ML
LangGraph
LangChain
Claude Code
Model Context Protocol
Prompt Engineering
AI Agents
AWS Bedrock
CrewAI
AWS Bedrock AgentCore
Human-in-the-Loop
Multi-Agent Systems
Frontend
React.js
DevOps
AWS
GitHub
Management
Confluence
Jira
Apply
Research Scientist V 2 hours ago
$114k – $154k per year • In office • Contractor • 8+ years exp • PhD • Redmond
Python
MATLAB
Apply
≈ $124k – $247k per year (Estimated) • In office • Contractor • Redmond
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Computer Vision
TensorFlow
PyTorch
Machine Learning
Robotics
SLAM
Localization
Apply
Sr. AL/ML Engineer 2 hours ago
$120k – $200k per year • In office • TS/SCI • 8+ years exp • Bachelor's Degree • Annapolis Junction
Python
AI/ML
LangGraph
LangChain
Model Context Protocol
vLLM
AI Agents
LiteLLM
Ollama
LLM
RAG
NVIDIA NIM
NVIDIA NeMo
Context Engineering
LLM Guardrails
Agentic Workflows
Agno
Apply
≈ $21k – $46k per year (Estimated) • In office • 4+ years exp • Hyderabad
Python
JavaScript
TypeScript
Python
Flask
FastAPI
Django
AI/ML
LangGraph
LangChain
LlamaIndex
Fine-tuning
AI Agents
Semantic Kernel
LLM
RAG
Machine Learning
Frontend
Angular
React.js
DevOps
Rest API
Azure
CI/CD
AWS
Apply
≈ $22k – $46k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
Python
Flask
FastAPI
Django
Databases
MySQL
Redis
RabbitMQ
Apache Kafka
Google BigQuery
BigQuery
AI/ML
LangChain
LlamaIndex
Embeddings
Scikit-learn
Prompt Engineering
AI Agents
TensorFlow
Pandas
NumPy
Keras
PyTorch
LLM
Context Engineering
Machine Learning
DevOps
GCP
GitHub Actions
CI/CD
AWS
Apply
≈ $100k – $236k per year (Estimated) • In office • Perth
Python
JavaScript
Node JS
Databases
DynamoDB
Amazon Aurora
AI/ML
LangChain
AI Agents
AWS Bedrock
Copilot Studio
DevOps
CI/CD
AWS
AWS Lambda
IAM
Amazon EventBridge
AWS Step Functions
API Gateway
Analytics
ETL/ELT
Management
Agile
Apply
≈ $15k – $40k per year (Estimated) • In office • Ho Chi Minh City
Python
Go
JavaScript
TypeScript
SQL
Databases
MySQL
PostgreSQL
Redis
pgvector
RabbitMQ
ElasticSearch
Apache Kafka
OpenSearch
AI/ML
Claude
Claude Code
Embeddings
Function Calling
Gemini
LLM
RAG
OpenAI
Structured Outputs
LLM Guardrails
Frontend
Vue.js
Pinia
Vite
Vue Router
Mobile
Dependency Injection
DevOps
Rest API
Helm
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
QA
Playwright
Apply
≈ $18k – $44k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Gurgaon
Python
Java
AI/ML
NLP
DevOps
GCP
Azure
CI/CD
Git
AWS
Management
Agile
Apply
≈ $129k – $267k per year (Estimated) • In office • 5+ years exp • Columbus
AI/ML
Copilot
Cursor
Claude
AI Agents
LLM
Human-in-the-Loop
DevOps
GCP
Azure
Management
Agile
Apply
≈ $129k – $291k per year (Estimated) • In office • Jersey City
Management
Agile
Apply
≈ $113k – $240k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Tampa
Python
Java
AI/ML
Model Context Protocol
Fine-tuning
Embeddings
Prompt Engineering
Function Calling
AI Agents
TensorFlow
PyTorch
RAG
OpenAI
Anthropic
LLM Guardrails
Multi-Agent Systems
DevOps
CI/CD
Cybersecurity
Threat Modeling
Management
Agile
Apply
≈ $136k – $282k per year (Estimated) • In office • 5+ years exp • Master's Degree • Jersey City
Python
JavaScript
Rust
TypeScript
SQL
Databases
PostgreSQL
Redis
Snowflake
Databricks
OpenSearch
AI/ML
Model Context Protocol
Fine-tuning
Embeddings
AI Agents
LLM
RAG
A2A
GraphRAG
Knowledge Graph
LLM Guardrails
Multi-Agent Systems
Machine Learning
Frontend
Svelte
Next.js
React.js
DevOps
Splunk
Azure
CI/CD
AWS
Kubernetes
Bitbucket
Amazon EKS
Cybersecurity
Defense in Depth
Management
Confluence
Agile
Apply
≈ $105k – $210k per year (Estimated) • In office • 3+ years exp • Plano
Python
Databases
Databricks
AI/ML
Spark
Scikit-learn
Pandas
NumPy
LLM
RAG
Feature Store
Machine Learning
DevOps
Terraform
AWS
Docker
Kubernetes
Amazon S3
Amazon ECS
Management
Agile
Apply
≈ $128k – $285k per year (Estimated) • Hybrid • Full-Time • Palo Alto • São Paulo
Python
SQL
Web3
Bitcoin
Ethereum
Staking
Analytics
ETL/ELT
A/B Testing
Apply
Benefits Manager 9 hours ago
≈ $41k – $97k per year (Estimated) • In office • Full-Time • London • Palo Alto • New York • San Francisco • Paris
Apply
Program Scheduler 12 hours ago
$55k – $67k per year • Equity • In office • Full-Time • Bachelor's Degree • Palo Alto
Apply
$99k – $125k per year • Remote (United States) • Full-Time • Palo Alto
Apply
≈ $76k – $177k per year (Estimated) • Equity • Remote (United States) • Full-Time • Palo Alto
DevOps
Incident Management
SLI/SLO/SLA
Management
ITSM
Service Desk
Apply
See all jobs
This is one of many
1,421,583 more open roles from verified company boards, updated every day.