372,576open jobs
9,648companies
49,982added this week
Browse all
Salary
$115k – $224k per year (Estimated)
Location
Remote/Hybrid (Alpharetta, United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
PDI Technologies is an enterprise software company headquartered in Alpharetta, Georgia, and founded in 1983. The company sells back office, fuel pricing, logistics, loyalty, and payments software to convenience retailers, fuel wholesalers, and petroleum distributors. It serves a large share of the North American convenience store market and has expanded internationally through a long series of acquisitions.

PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world.

This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself - reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.

Key Responsibilities

  • Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20-person SRE organization supporting the Paylo product suite.

  • Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.

  • Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.

  • Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.

  • Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.

  • Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.

  • Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.

  • Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.

  • Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.

  • Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.

  • Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.

  • Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.

  • Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.

  • Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.

  • Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.

  • Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.

  • Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.

  • Track and report cost, capacity, and operational KPIs to senior leadership.

Required Qualifications

    • 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role.

    • Proven experience managing managers- you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of ~20 total reports.

    • Strong, hands-on expertise across AWS and Azure- you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so.

    • Strong, hands-on expertise with Kubernetes and Helm- cluster operations, troubleshooting at scale, and chart design/maintenance.

    • Strong, hands-on expertise with Argo CD/Argo Workflowsfor GitOps-based continuous delivery.

    • Strong, hands-on expertise with Infrastructure as Code(Terraform, OpenTofu), including module design and state management.

    • Strong, hands-on expertise with Jenkinsfor CI/CD pipeline design, administration, and automation.

    • Strong, hands-on expertise with Datadog(or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale.

    • Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems.

    • Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility.

    • A strong, visible leadership style - someone who sets clear direction, holds teams accountable, and builds trust across the organization.

    • Applicants must be legally authorized to work in the United States without the need for employer sponsorship, now or in the future. PDI Technologies is unable to offer visa sponsorship for this role.

Preferred Qualifications

    • Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.

    • Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate.

    • Experience with messaging systems (Kafka/SQS/SNS), PagerDuty (or similar), and multi-region/multi-AZ resilience patterns.

    • Prior experience consolidating or standardizing SRE and DevOps practices across multiple product lines or recently-integrated/acquired teams.

    • Experience partnering with product and business stakeholders to translate reliability investments into business outcomes.

What Success Looks Like

    • A stable, well-led SRE organization with clear ownership, career paths, and low regrettable attrition among your managers and their teams.

    • Consistent, Datadog-driven observability and SLOs in place across the organization, with measurable reduction in Sev1/Sev2 incidents and mean time to detect/resolve.

    • Modern, standardized infrastructure practices - GitOps delivery via Argo, IaC via Terraform/OpenTofu, and reliable CI/CD via Jenkins - adopted consistently across teams and clouds.

    • A mature, blameless incident-management culture with strong postmortem follow-through.

    • Strong cross-functional trust with engineering, product, and security/compliance stakeholders.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
372,576 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Alpharetta
$16k – $38k per year (Estimated) • In office • Full-Time • 3+ years exp • India
SQL
Databases
Azure SQL Database
DevOps
AWS
Azure
CI/CD
GitHub
Kubernetes
Terraform
Apply
$24k – $63k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • India
Crystal
Groovy
JavaScript
Perl
Python
Ruby
SQL
TypeScript
Java
Java
Apache Tomcat
Gradle
Hibernate
Maven
Spring Boot
Spring MVC
Databases
Apache Kafka
Db2
Oracle
PostgreSQL
RabbitMQ
AI/ML
Fine-tuning
Frontend
Angular
JQuery
DevOps
Apache HTTP Server
AWS
Azure
CI/CD
Docker
GCP
Jenkins
Kubernetes
Rest API
Cybersecurity
Checkmarx
SonarQube
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
$26k – $63k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Pune
Python
Python
pySpark
AI/ML
Hadoop
Spark
DevOps
AWS
Azure
GCP
Analytics
ETL/ELT
Apply
$177k – $265k per year • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Jersey City • Tampa
Java
Python
Databases
Apache Kafka
Databricks
Snowflake
DevOps
AWS
Azure
GCP
QA
Pact
Apply
$12k – $38k per year (Estimated) • In office • Full-Time • 8+ years exp • Chennai
C#
JavaScript
Python
AI/ML
Embeddings
LLM
Prompt Engineering
RAG
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
GCP
Platform Engineering
QA
Cypress
Playwright
Postman
Rest-Assured
Selenium
TestNG
Apply
Software Engineer II 13 days ago
$19k – $52k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
C#
JavaScript
SQL
C#
ASP.NET Core
AI/ML
Claude
AI Agents
Frontend
React.js
DevOps
Azure
Azure DevOps
Management
Jira
Apply
$49k – $139k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Maidenhead
Python
SQL
AI/ML
Claude
Claude Code
AI Agents
DevOps
AWS
Analytics
A/B Testing
ETL/ELT
Apply
SOFTWARE ENGINEER II 26 days ago
$18k – $50k per year (Estimated) • In office • Full-Time • 2+ years exp • Chennai
C#
JavaScript
SQL
TypeScript
C#
ASP.NET Core
Databases
MS SQL
Frontend
Angular
Mobile
MVC
DevOps
Azure
Azure DevOps
Management
Jira
Apply
Data Engineer II 2 months ago
$20k – $49k per year (Estimated) • Remote/Hybrid • 2+ years exp • Bachelor's Degree • Chennai
Python
SQL
Python
pySpark
Databases
Amazon Redshift
AI/ML
Anomaly Detection
Spark
DevOps
AWS
CI/CD
Git
Terraform
Amazon S3
AWS Step Functions
Analytics
ETL/ELT
Marketing
Salesforce
Apply
$183k – $276k per year • Equity • In office • Full-Time • 8+ years exp • Sunnyvale • Alpharetta
DevOps
AWS
Azure
GCP
Kubernetes
Cybersecurity
Snort
Suricata
Tcpdump
Wireshark
Zero Trust
Apply
API Security Engineer 2 hours ago
$128k – $216k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Alpharetta • Columbus
DevOps
CI/CD
Git
Cybersecurity
ISO 27001
Least Privilege
PCI DSS
Threat Modeling
Apply
$72k – $119k per year • In office • Full-Time • Bachelor's Degree • Boca Raton • Alpharetta • Dayton
Java
DevOps
CI/CD
Git
Apply
$59k – $99k per year • In office • Full-Time • Bachelor's Degree • Alpharetta
C++
JavaScript
Python
SQL
C#
C#
.NET
DevOps
AWS
Azure
Apply
Lead AI Engineer 1 day ago
$106k – $227k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Omaha • Alpharetta
Java
Python
C#
TypeScript
JavaScript
Java
Spring Boot
C#
.NET
AI/ML
Claude
Copilot
LLM
Anthropic
OpenAI
Frontend
Angular
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitLab CI
Jenkins
Kubernetes
OpenShift
Rest API
GitLab
Apply
See all jobs
This is one of many
372,576 more open roles from verified company boards, updated every day.