368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$112k – $218k per year (Estimated)
Location
Remote/Hybrid (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
FreedomPay is an enterprise commerce technology company that provides secure, data-driven payment processing solutions for global businesses. Its fully integrated platform connects point-of-sale systems, payment gateways, fraud protection, and loyalty management across physical and digital environments. Serving industry leaders in hospitality, retail, gaming, and sports venues, the company enables seamless, compliant, and omnichannel transaction management worldwide.

FreedomPay is seeking an experienced Senior Site Reliability Engineer to help ensure the highest possible availability and resiliency of a rapidly growing global payment platform. This full-time salaried position builds on a strong foundation of observability, incident response, and support experience across the development lifecycle - and pushes it forward with AI-driven operations and automation at its core. The right candidate finds real satisfaction in eliminating manual toil, treats every recurring task as an automation opportunity, and is eager to apply modern AI tooling to detect, diagnose, and resolve issues faster than ever before.

About the Role

    You’ll join a team of SREs who work closely with other teams of world-class engineers to tenaciously and creatively solve problems and reduce manual toil wherever possible. We expect AI and automation to be a force multiplier in everything you do - from accelerating root-cause analysis and enriching alerts, to generating runbooks and codifying remediation so that the platform increasingly heals itself.

    Successful candidates are heavily results-driven, bring well-established expertise across both traditional and bleeding-edge technology, and have a strong desire to continuously grow and improve themselves and our platform. This is a global operation spanning multiple regions and time zones, and the role demands the flexibility and commitment that a 24/7 payment platform requires.

    • This position participates in an engineering on-call rotation and provides after-hours support for production issue escalations on a rotational basis.

    • This position is based in the Philadelphia area with a hybrid schedule. Remote arrangements may be considered for exceptional candidates, with occasional travel to Philadelphia required.

Primary Responsibilities:

  • Build and maintain a comprehensive understanding of the platform and custom application stack.
  • Implement, maintain, and continuously improve observability strategies and metrics that ensure complete system health for numerous complex products throughout all stages of the development lifecycle, up to and including production.
  • Continuously identify automation opportunities and follow through to successful implementation, applying AI-assisted tooling to accelerate development and reduce manual effort.
  • Design, build, and maintain automated remediation and self-healing workflows that detect, triage, and resolve common failure modes with minimal human intervention.
  • Leverage AI/ML-driven observability - anomaly detection, alert correlation, and intelligent noise reduction - to surface issues earlier and shorten time to detection.
  • Use AI-assisted analysis to accelerate root-cause investigation, enrich incident context, and generate first-draft postmortems and runbooks for human review.
  • Handle escalations and collaborate effectively with other team members to quickly determine the root cause of any type of service degradation.
  • Implement, maintain, and continuously improve incident response procedures and other operational documentation, automating documentation generation and upkeep wherever practical.
  • Assist with troubleshooting and remediation of failed scheduled jobs and data-related concerns.
  • Champion responsible, secure adoption of AI tooling across the SRE function - sharing patterns, prompts, and automations that raise the productivity of the whole team

AI Enablement & Automation

    AI and automation are central to how this team operates. We are looking for someone who will not only use these tools but help define how the SRE function applies them. In this role you will:

  • Apply AI-assisted development and operations tools - including Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) and the agentic workflows built on them - to write, review, and accelerate automation and infrastructure code.
  • Build and integrate automation that turns repetitive operational work into codified, repeatable, and self-service workflows.
  • Use AIOps and ML-driven observability capabilities within the APM stack for anomaly detection, predictive alerting, and alert correlation.
  • Develop and refine prompts, agents, and integrations that connect monitoring, ticketing, and remediation systems into faster end-to-end response loops.
  • Evaluate emerging AI tooling for reliability and operations use cases, and advocate for adoption where it delivers measurable improvements in toil reduction, MTTR, or availability.
  • Ensure all AI and automation usage adheres to FreedomPay’s security, privacy, and PCI obligations - keeping sensitive data appropriately protected and human review in place for high-impact actions.

Required Background and Experience

  • BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
  • Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
  • Demonstrated history of self-directed learning - someone who independently seeks out knowledge, builds new skills without being told to, and doesn’t wait for formal training to close gaps.
  • Next-level problem-solving abilities and a strong bias toward practical, proven solutions.
  • A track record of identifying and eliminating manual toil through automation.
  • Excellent communication and organizational skills, with a strong sense of ownership and service.

Required Technical Skills

  • Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities; Dynatrace experience strongly preferred, though deep expertise in comparable tools such as Datadog or New Relic where readily transferable.
  • Hands-on experience with AI-assisted development and automation tools - such as Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) - and a demonstrated ability to apply them to real operational and engineering work.
  • Proficiency in scripting and automation - PowerShell and/or Python - to build tooling and remediation workflows.
  • Strong SQL / T-SQL skills.
  • Solid understanding of core networking concepts: DNS, HTTP/HTTPS, load balancing, and TCP/IP routing and switching.
  • Working knowledge of modern technology infrastructure including container orchestration, IaaS/PaaS cloud services, Azure, and VMware.
  • Working knowledge of application development processes.

Preferred Technical Skills and Experience

  • Proven track record of successfully implementing SLI/SLOs and fostering their adoption across an organization.
  • Experience implementing enterprise incident management practices.
  • Experience building AIOps or ML-driven automation into production observability and incident response.
  • Azure Kubernetes Service (AKS) and broader container orchestration experience.
  • Windows Server (IIS) administration.
  • PagerDuty Process Automation (formerly Rundeck) or comparable runbook automation platforms.
  • Comprehensive experience supporting real-time transaction processing applications.
  • PCI policies and best practices.

Additional Experience, a Plus

    • AI/ML model deployment, evaluation, or operations (MLOps).

    • Documentation automation and self-service tooling / service catalog implementation.

    • Experience integrating QA test automation into CI/CD pipelines.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$19k – $28k per year (net) • Remote • Full-Time • Moscow
C#
C++
C++
CMake
DevOps
CI/CD
Git
Management
Jira
Slack
Apply
$14k – $31k per year (Estimated) • Remote/Hybrid • 3+ years exp • Moscow
JavaScript
DevOps
CI/CD
Git
GitLab CI
GitLab
Management
Confluence
Jira
QA
Playwright
Postman
Apply
Backend Engineer 1 day ago
Remote/Hybrid • 1+ year exp
Node JS
TypeScript
JavaScript
Node JS
Express
Nest.JS
Databases
MySQL
PostgreSQL
Redis
Mobile
Firebase
DevOps
CI/CD
GCP
Git
Apply
$132k – $261k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Philadelphia
Design
Figma
Apply
$107k – $237k per year (Estimated) • Remote • PhD
JavaScript
Python
SQL
Databases
Apache Kafka
Snowflake
AI/ML
AI Agents
Streamlit
DevOps
Azure
CI/CD
Cortex
Prometheus
Apply
$78k – $163k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • Philadelphia
PowerShell
Python
SQL
Databases
MS SQL
AI/ML
Anomaly Detection
Claude
Anthropic
DevOps
AIOps
Azure
Datadog
Dynatrace
Incident Management
Kubernetes
New Relic
Splunk
Terraform
VMWare
Apply
$31k – $81k per year (Estimated) • Remote • 2+ years exp • Bachelor's Degree • Dublin
Management
Confluence
Jira
Marketing
Zendesk
Apply
$116k – $225k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Master's Degree
JavaScript
Python
AI/ML
AI Agents
DevOps
AIOps
AWS
Azure
Dynatrace
GCP
Kubernetes
OpenTelemetry
Platform Engineering
SRE
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.