Salary
≈ $124k – $241k per year (Estimated)
Location
In office (Palo Alto)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
kodypay is on a mission to make in-person payment acceptance easy. today, paying in person presents common problems for businesses, such as high costs, long queues, and limited choice of payment methods. kodypay fully integrates the payment ecosys...
Senior Site Reliability Engineer (Payments Infrastructure)
Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.
Responsibilities
- Participate in a follow-the-sun production on-call rotation as a primary incident responder.
- Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
- Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews.
- Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments.
- Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR.
- Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems.
- Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
- Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.
- Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
- Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.
Leadership & Operational Excellence
- Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact.
- Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure.
- Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments.
- Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
- Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil.
- Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events.
- Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability.
- Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design.
Benefits
- Competitive packages aligned with California market standards
- Lead a dynamic and innovative team in a very rapidly growing company
- Collaborative, inclusive environment where your contributions are recognized and valued
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Palo Alto
Senior Backend Engineer
7 hours ago
$120k – $180k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • London
Python
Python
Celery
Django
FastAPI
Flask
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kong
Kubernetes
Cybersecurity
GDPR
SOC 2
Apply
Director, Software Engineering
7 hours ago
≈ $39k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Pune
C++
Go
Java
Java
Spring Boot
Databases
Apache Kafka
NATS
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Rest API
Apply
Infrastructure/ Security/ DevOps
7 hours ago
$80k – $140k per year • In office • Full-Time • 6+ years exp • London
Python
Databases
MySQL
PostgreSQL
DevOps
AWS
Azure
Azure AKS
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Terraform
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
Full-Stack Engineer
7 hours ago
$80k – $140k per year • In office • Full-Time • 6+ years exp • London
Python
Python
Celery
Django
FastAPI
Databases
Redis
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Apply
≈ $74k – $196k per year (Estimated) • In office • Full-Time • 8+ years exp • Dublin
C#
Java
Node JS
Python
JavaScript
Java
Hibernate
Spring Boot
DevOps
AWS
Azure
GCP
Cybersecurity
CWE
Apply
Senior Product Manager- Payment
12 days ago
≈ $106k – $218k per year (Estimated) • Remote • 5+ years exp • Singapore
Apply
Senior Product Manager- payment
12 days ago
≈ $111k – $175k per year (Estimated) • In office • 5+ years exp • London
Apply
Apply
Apply
Technical Product Owner
1 month ago
In office • Full-Time • 3+ years exp • Hong Kong
AI/ML
AI Agents
Apply
Chief of Staff
7 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
≈ $140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
Executive Director- Applied AI/ML Lead
1 day ago
≈ $139k – $294k per year (Estimated) • In office • 10+ years exp • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
Vice President-Applied AI/ML Lead
1 day ago
≈ $137k – $292k per year (Estimated) • In office • 6+ years exp • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
≈ $150k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
C++
Java
Python
Rust
AI/ML
LLM
Apply
This is one of many
368,530 more open roles from verified company boards, updated every day.

