We are seeking a highly skilled Cloud Platform Systems Engineer/Manager to join our global Cloud Infrastructure and Architecture team, responsible for the stability, reliability, and performance of Clients' mission-critical Enterprise Risk Technology (ERT). These applications under ERT serve as the operational backbone for clients' risk businesses, including compliance risk, operational risk, retail risk, and regulatory pillars.
This is a pivotal role within a high-performing team dedicated to building a world-class DevOps, infrastructure, and architecture group. The candidate will work with the latest technologies, AI-driven automation, and Site Reliability Engineering (SRE) principles to enhance client experience, digitize services, and ensure our platforms can scale for future growth. In this global capacity, candidates will collaborate extensively with Application Development, Infrastructure Setup and Performance Management, Product Rollout using CI/CD Pipelines, Business, and Operations teams to guarantee seamless service delivery.
Responsibilities:
- Containerization and Orchestration: Deep expertise in deploying, managing, and scaling applications using container technologies, including Kubernetes and OpenShift.
- Service Mesh: Strong knowledge of service mesh concepts and hands-on experience with Istio for ingress and egress traffic management.
- Security Proficiency: Solid understanding of security concepts including Single Sign-On (SSO), SAML, Secured SSO, MFA, SSG, COIN authentication, TLS/SSL certificate setup, and authorization protocols such as AD/LDAP.
- WebSphere Application Server: Experience in managing WebSphere environments, including support activities such as JNDI, connection pool, JMS configurations, heap management, and session management.
- Big Data Technologies: Experience supporting data platforms and ingestion pipelines using technologies such as Spark, Hive, and Starburst.
- Cloud Platforms: Experience with cloud migration projects and supporting applications on cloud platforms such as AWS or IBM Cloud.
- Incident and Problem Management: Lead the resolution of critical production incidents, including prioritization, timely escalation, and clear communication to all stakeholders. Conduct comprehensive post-mortems and root cause analyses to ensure permanent solutions and prevent recurrence.
- Technical and Business Support: Serve as a key point of contact for technical and business support for client users. Address daily queries and issues, and proactively drive improvements in stability, efficiency, and risk management.
- Change and Release Management: Support the planning and execution of all system changes, including application releases, infrastructure maintenance, and continuity of business (COB) tests, ensuring production stability is maintained throughout the process.
- System Observability and Monitoring: Maintain and optimize the production monitoring and observability estate. Implement new features and leverage advanced analytics to enhance system visibility, proactive alerting, and predictive fault detection.
- Automation and Efficiency: Identify opportunities to reduce operational toil and mitigate risk by developing and implementing automation tools and processes.
- Collaboration and Stability: Partner closely with development teams to recommend and implement architectural and code improvements that enhance application stability, performance, and recoverability.
Requirements:
- A minimum of 5 years of experience in an infrastructure role managing enterprise-level, business-critical applications.
- Prior direct experience with cloud migration projects and supporting applications on cloud platforms such as AWS or IBM Cloud.
- Strong expertise in the deployment, management, and scalability of applications using container technologies, including Kubernetes and OpenShift.
- Proven expertise in service mesh architecture with hands-on Istio experience for ingress and egress traffic control.
- Comprehensive knowledge of security concepts encompassing Single Sign-On (SSO), SAML-based SSO, MFA, SSG, COIN authentication, TLS/SSL certificate setup, and directory-based authorization protocols such as AD/LDAP.
- Direct experience in administering WebSphere environments, supporting JNDI, connection pool, and JMS configurations, as well as performing heap and session management.
- Experience supporting data platforms and ingestion pipelines leveraging technologies such as Spark, Hive, and Starburst.
- Advanced proficiency in Unix/Linux environments, including complex shell scripting.
- Demonstrated experience with enterprise monitoring tools such as ITRS Geneos and AppDynamics and log aggregation platforms such as Splunk and ELK.
- Hands-on experience with containerization and orchestration platforms, specifically OpenShift and/or Kubernetes.
- Strong database skills, with experience in both relational databases such as Oracle and MSSQL and NoSQL databases such as MongoDB.
- In-depth knowledge of enterprise messaging solutions such as Tibco EMS, IBM MQ, or Kafka.
- Solid understanding of distributed application architecture, including networks, load balancers, storage, and authentication protocols (AD/LDAP).
- Experience working with and troubleshooting REST APIs.
- Exceptional written and verbal communication skills, with the ability to articulate complex technical issues to both technical and non-technical audiences.
Desirable:
- Proficiency in a high-level programming language such as Python or Java for developing automation solutions.
- Experience with modern observability platforms and standards, such as Prometheus and Grafana.
- Familiarity with prompt engineering techniques for interacting with generative AI and large language models (LLMs).
- Bachelor's/University Degree or equivalent professional experience.

