385,264open jobs
10,055companies
49,525added this week
Browse all
Location
In office
Seniority
Junior · 5+ years exp
Overview
Company
Impact
Profile match
American Express is a New York financial services company founded in 1850 as an express freight business that became a payments network and card issuer. Unlike the four-party networks it competes with, it issues most of its own cards and operates its own network, which lets it earn merchant discount revenue as well as interest and annual fees, and supports a premium rewards proposition built on travel and lounge access. Its business spans consumer and small business cards, corporate payments, merchant acquiring and travel services, and it is a component of the Dow Jones Industrial Average.

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.

  • Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
  • Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
  • Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
  • Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
  • Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
  • Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
  • Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
  • Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SRE-owned work items
  • Drive reliability improvements through sprint-based commitments, including automation, operational fixes, and platform enhancements
  • Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of business-critical services
  • Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
  • Support ITIL-based Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and post-implementation validation
  • Monitor network performance and troubleshoot connectivity, latency, and access-related issues impacting platform traffic
  • Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
  • Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
  • Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
  • Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
  • Participate and lead the Development change review and change validation processes
  • Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
  • Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers

Education Qualifications:

  • Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
  • Bachelor’s degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
  • Strong knowledge of operating systems and application runtimes such as Java and .NET
  • Knowledge of distributed systems and service-based architectures from an operations and reliability perspective
  • Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
  • Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
  • Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
  • Knowledge of scripting and automation using languages such as PowerShell and Python
  • Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus

Work Experience:

  • Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
  • Experience supporting production systems in large-scale enterprise environments with a focus on reliability and availability
  • Experience in system administration, infrastructure operations, and network troubleshooting
  • Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
  • Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
  • Knowledge of cloud-based Site Reliability Engineering (SRE) practices with hands-on experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservices-based architectures
  • Experience using enterprise monitoring and alerting platforms such as ELF
  • Exposure to AI-assisted monitoring, automation, or AIOps tools is a plus

    • Proficiency in connecting to and administering servers via SSH (Secure Shell)

  • Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access

Licenses & Certifications

  • Certification in at least one programming language or runtime such as Java, .NET, or Python
  • Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
  • Public cloud certification in AWS or GCP is a plus
  • Certification or training related to AI platforms, analytics platforms, or AIOps is a plus

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
385,264 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • 6+ years exp • Bachelor's Degree
Python
C#
C#
.NET
Databases
Databricks
FAISS
Google BigQuery
Milvus
OpenSearch
Pinecone
Snowflake
Weaviate
AI/ML
AI Agents
AutoGen
AWS Bedrock
Copilot
CrewAI
Function Calling
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
LLM Guardrails
OpenAI
Prompt Engineering
RAG
Semantic Kernel
Vertex AI
DevOps
AWS
Azure
CI/CD
GCP
Git
Vector
Marketing
Salesforce
Apply
$31k – $79k per year (Estimated) • In office • 6+ years exp • Master's Degree • Bengaluru
Java
Python
Scala
AI/ML
AI Agents
Claude
Claude Code
Computer Vision
Cursor
Model Context Protocol
Multimodal AI
NLP
Spark
DevOps
Amazon S3
Analytics
A/B Testing
Apply
$103k – $210k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Pittsburgh • Phoenix • Lakewood • Birmingham
Java
Java
Spring Boot
Databases
Apache Kafka
ElasticSearch
Oracle
DevOps
Dynatrace
Kubernetes
OpenShift
QA
Postman
Apply
$159k – $215k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Springfield
JavaScript
Java
Java
Apache Tomcat
Frontend
JQuery
Less
React.js
DevOps
Amazon S3
AWS
GitHub
Rest API
Management
Jira
Apply
$102k – $138k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • United States
Python
SQL
Databases
Amazon Aurora
Apache Kafka
Databricks
PostgreSQL
Snowflake
DevOps
Amazon Kinesis
AWS
CI/CD
Git
Analytics
ETL/ELT
Apply
$75k – $163k per year (Estimated) • In office • Phoenix
SQL
Databases
Google BigQuery
AI/ML
Flink
DevOps
GCP
IAM
Incident Management
Management
Jira
ServiceNow
Apply
In office • Gurgaon
Python
AI/ML
AI Agents
CUDA
CUDA Toolkit
Embeddings
Multimodal AI
NCCL
PyTorch
Quantization
Reranking
Triton
vLLM
DevOps
Docker
Git
Kubernetes
Apply
$24k – $62k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Bengaluru
JavaScript
PowerShell
Python
C#
C#
.NET
Apply
In office • Bachelor's Degree • Gurgaon
Python
Scala
Python
pySpark
AI/ML
Spark
Apply
$21k – $49k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Google BigQuery
AI/ML
Hallucination
Human-in-the-Loop
LLM
NLP
PyTorch
RAG
Scikit-learn
Spark
TensorFlow
DevOps
AWS
Azure
GCP
Git
Analytics
ETL/ELT
Matplotlib
Seaborn
Tableau
Apply
See all jobs
This is one of many
385,264 more open roles from verified company boards, updated every day.