Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 30, 2026. JPMorganChase scores A on the Alion truth index.
As a Lead Software Engineer - Software Reliability at JPMorgan Chase as a part of our product team, you will design, build, and operate scalable, resilient systems using Python and modern reliability practices. You will apply Site Reliability Engineering (SRE) principles to improve availability, performance, and operational excellence, and you will help establish engineering standards that increase system reliability and security. You will collaborate with engineering and product partners to troubleshoot, optimize, and maintain production services while fostering a collaborative and inclusive team culture.
Job Responsibilities
Design and develop scalable and resilient systems using Python to support continuous improvement and apply Site Reliability Engineering (SRE) concepts to enhance system reliability and performance
Execute software solutions, including design, development, and technical troubleshooting and create secure, high-quality production code and maintain algorithms that run synchronously with appropriate systems
Produce or contribute to architecture and design artifacts, ensuring design constraints are met
Gather, analyze, and synthesize data to develop visualizations and reporting for software and system improvement
Identify hidden problems and patterns in data to drive improvements in coding hygiene and system architecture
Implement reliability engineering practices such as monitoring, alerting, and automated recovery
Define and measure Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to track system health
Conduct chaos engineering experiments to test system resiliency and identify weaknesses
Perform performance testing using tools such as JMeter to ensure scalability and stability
Collaborate with product teams to enhance system reliability, scalability, and performance and contribute to software engineering communities of practice and events exploring new and emerging technologies
Foster a team culture of diversity, opportunity, inclusion, and respect and participate in post-incident reviews and drive root cause analysis for system failures
Required qualifications, capabilities, and skills
- Hands-on practical experience in system design, application development, testing, and operational stability and proficient in coding in Python
- Experience developing, debugging, and maintaining code in a large corporate environment with modern programming languages and database querying languages
- Knowledge of the Software Development Life Cycle, AWS cloud exposure, troubleshooting abilities, resiliency, and automation focus
- Understanding of agile methodologies such as CI/CD, application resiliency, and security
- Knowledge of software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, mobile, etc.)
- Experience with AI and full understanding of the SDLC process, and MongoDB
- Familiarity with reliability engineering concepts, including monitoring, alerting, and automated recovery and ability to implement and maintain system health checks and performance metrics
- Commitment to writing maintainable, testable, and high-quality code
- Understanding of SRE principles, including SLIs, SLOs, and error budgets and experience with incident response and root cause analysis
- Experience with chaos engineering practices to test system resiliency and proficiency in performance testing tools such as JMeter
Preferred qualifications, capabilities, and skills
- Familiarity with modern front-end technologies
- Exposure to cloud technologies

