368,941open jobs
9,452companies
47,951added this week
Browse all
Salary
$125k – $168k per year
Location
In office (Charlotte, Jersey City, Plano)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Bank of America is one of the world's leading financial institutions, serving individual consumers, small and middle-market businesses, and large corporations with a full range of banking, investing, and asset management products. Headquartered in Charlotte, North Carolina, it operates an extensive retail banking network alongside robust digital platforms to deliver seamless financial services. Through its wealth management and global markets divisions, including Merrill, the company provides comprehensive investment strategies and corporate advisory services worldwide.

Job Description:

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Leadis responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.

Job Description:

This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week:

40

Pay Transparency details

US - NJ - Jersey City - 101 Hudson St - 101 Hudson (NJ2101)Pay and benefits informationPay range$125,300.00 - $167,900.00 annualized salary, offers to be determined based on experience, education and skill set.Discretionary incentive eligibleThis role is eligible to participate in the annual discretionary plan. Employees are eligible for an annual discretionary award based on their overall individual performance results and behaviors, the performance and contributions of their line of business and/or group; and the overall success of the Company.BenefitsThis role is currently benefits eligible. We provide industry-leading benefits, access to paid time off, resources and support to our employees so they can make a genuine impact and contribute to the sustainable growth of our business and the communities we serve.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Charlotte
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
up to $76k per year (gross) • In office • Full-Time • 5+ years exp • Moscow
Swift
Mobile
Clean Architecture
Fastlane
MVVM
SwiftUI
DevOps
CI/CD
Git
Jenkins
Apply
up to $63k per year (gross) • In office • Full-Time • 5+ years exp • Moscow
SQL
Databases
Apache Kafka
AI/ML
LLM
RAG
DevOps
CI/CD
Git
gRPC
WebSockets
QA
Postman
Swagger
Apply
$16k – $60k per year (Estimated) • In office • Full-Time • PhD • Mumbai
Python
SQL
Python
pySpark
Databases
Presto
Snowflake
AI/ML
Dagster
Prefect
Spark
DevOps
Amazon S3
AWS
CI/CD
Analytics
ETL/ELT
Power BI
Tableau
Apply
$143k – $173k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Wiesbaden
Python
DevOps
Ansible
Terraform
Cybersecurity
Defense in Depth
Wireshark
Zero Trust
Apply
$110k per year • In office • Full-Time • Bachelor's Degree • New York
Python
SQL
Apply
$152k – $185k per year • In office • Full-Time • 7+ years exp • Plano • Jersey City • Charlotte
Databases
GraphDB
Neo4j
Redis
AI/ML
GraphRAG
Knowledge Graph
DevOps
Grafana
Kubernetes
Platform Engineering
Prometheus
Splunk
Apply
$108k – $165k per year • In office • Full-Time • Charlotte • Jersey City • Plano
COBOL
COBOL
IBM MQ
Databases
Db2
IMS
DevOps
Splunk
Apply
$145k – $193k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Washington • Chicago • Denver
Java
Python
AI/ML
Claude
Claude Code
Copilot
DevOps
AWS
Azure
CI/CD
GCP
GitLab CI
Jenkins
GitHub
GitLab
Apply
$79k – $160k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Newark • Plano • Fort Worth • Boston • Charlotte
Python
SQL
Analytics
Tableau
Apply
$143k – $274k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Antonio • Charlotte • Colorado Springs • Plano • Phoenix
DevOps
IAM
Cybersecurity
CyberArk
Microsoft Entra ID
PCI DSS
Robotics
Path Planning
Management
ServiceNow
Apply
$117k – $132k per year • In office • Full-Time • PhD • Charlotte
Python
AI/ML
Anthropic
Anthropic SDK
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
OpenAI
OpenAI SDK
RAG
DevOps
AWS
Azure
GCP
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$122k – $200k per year • In office • Full-Time • 10+ years exp • Charlotte • New York
AI/ML
AI Agents
Context Engineering
Copilot
Hallucination
Human-in-the-Loop
Knowledge Graph
LLM
LLM Guardrails
LLMOps
Prompt Engineering
RAG
DevOps
Azure
CI/CD
GitHub
Kubernetes
Vector
Analytics
A/B Testing
Apply
$135k – $217k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Charlotte • Plano • Jersey City • Chicago
DevOps
Azure
Management
Confluence
Jira
Apply
See all jobs
This is one of many
368,941 more open roles from verified company boards, updated every day.