At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.
Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.
F5 is bringing a better digital world to life by helping organizations create, secure, and run applications that power our lives. Within the Platform Engineering team, this role helps ensure our platform isoperatedsafely, reliably, and with operational excellence.
We’relooking for aPrincipal SREwho leads with kindness,operateswell in a global, follow-the-sun environment, and brings strong execution, documentation, and cross-functional coordination skills.You willbe responsible fordesigning, deploying, andoperatingthe foundational infrastructure that underpins a large-scale, multi-datacenter platform spanning 30+ Points of Presence across the Americas, EMEA, and APAC.
This is a hands-on engineering role. You will own systems from bare metal to application layer -- provisioning physical servers via out-of-band management, building and maintainingProxmox-based hypervisor clusters, managingcontainerizedworkloads, and driving automation across a heterogeneous on-premises and cloud environment. You willoperatewithin a PCI-DSS compliant environment and be expected to contribute to hardening, audit readiness, and security tooling.
You will join an on-call rotation and be expected to respond to and lead incident resolution for production systems across a 24x7 global environment.
What You'll Do
Infrastructure Automation & Configuration Management
Author,maintain, and refactor Ansible playbooks and roles across a large-scale multi-datacenter inventory, covering the full lifecycle from bare-metal provisioning to application deployment
Develop and improve CI/CD pipelines (GitLab CI) for infrastructure automation, including linting, testing, and staged rollout across regions
Manage secrets lifecycle usingHashiCorpVault, includingAppRoleauthentication, secret rotation, and PKI integration
Maintain CMDB/IPAM accuracy inNetBoxas a source of truth for all infrastructure assets
Compute &Virtualization
Deploy and manageProxmoxVE hypervisor clusters on bare-metal HPE hardware, including cluster formation, OVS networking, ZFS storage, and VM replication
Provision and lifecycle-manage virtual machines using cloud-init, QCOW2 images, andProxmoxAPI automation
Manage physical server provisioning end-to-end via HPEiLO(firmware updates, SPP deployment, OS installation via virtual media)
Container & Kubernetes Platforms
Manage self-hosted Kubernetes clusters on-premises, including control plane operations, node provisioning, workload deployment, and upgrade management
OperateDocker-based workloads on infrastructure VMs using compose-driven deployments and container health monitoring
Maintaincontainer image pipelines and registry infrastructure (Azure Container Registryor AWS ECR)
Cloud Platforms
Engineer andmaintaininfrastructure on AWS and Azure, integrating cloud resources with on-premises systems (DNS, monitoring, identity, networking)
Apply cloud cost awareness, security best practices, andIaCprinciples (IAM, security groups, networking, storage) across AWS and Azure environments
Networking & Core Services
Operateand troubleshoot core distributed services including authoritative DNS (BIND9), recursive DNS (Unbound), load balancing (HAProxy), and high-availability VIPs (Keepalived/VRRP)
Maintaindirectory services (OpenLDAPmaster-replica topology) and AAA infrastructure (FreeRADIUS) used for SSH, VPN, and network device authentication
Manage OVS-based network configurations, VLAN topologies, and bonded NIC arrangements across hypervisor fleets
Observability & Security
Maintainand extend monitoring infrastructure (Prometheus,Observium) across a global fleet including SNMP polling, metrics collection, and alerting
Managecentralisedlog aggregation pipelines (Fluentbit) and ensure log delivery integrity across DCs
Operateruntime security tooling (Falco) and file integrity monitoring (AIDE) in production environments
Support PCI-DSS compliance activities including CIS hardening, audit logging (auditd), and participation in control reviews
Reliability & Incident Response
Participatein a 24x7 on-call rotation, responding to and leading production incident resolution
Conduct blameless post-mortems and drive remediation of root causes through automation and system improvements
Define and track SLOs/SLIs for critical infrastructure services
Identifyand address single points of failure; design and implement HA improvements
What We're Looking For
Required
7+ years of experience in a Site Reliability Engineering, DevOps, or Infrastructure Engineering role in a production environment
Strong Linux systems administration skills (RHEL/CentOS preferred) includingsystemd, networking, storage, kernel tuning, and package management
Proficiencywith Ansible(or similar tool)for large-scale configuration management, including role design, inventory management, and CI/CD integration
Hands-on experience with at least one hypervisor platform, preferablyProxmoxVE, Harvester (Kubevirt)or similar (VMware vSphere, KVM)
Production experienceoperatingon-premiseKubernetes clusters (rke2, k3s,etc)
Practical AWSorAzure experience includingcompute, networking (VPC/VNet, security groups, DNS), IAM, and managed services
Solid understanding of networking fundamentals: VLANs, bonding/LAG, routing, BGP concepts, DNS, load balancing, andfirewallrule management
Experience with secrets management platforms (HashiCorpVault or equivalent)
Familiarity with PCI-DSS requirements as they apply to infrastructure -- hardening standards (CIS benchmarks), audit logging, access control
Experience writing andmaintainingCI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)
Demonstrable on-call experience and comfort leading incident response in a global production environment
Preferred
Experience managing bare-metal server fleets including out-of-band management tools (HPEiLO, IPMI, or equivalent)
Experience with CMDB/IPAM tooling
Working knowledge of LDAP directory services and RADIUS authentication
Exposure to network monitoring tooling (Observium,LibreNMS, or similar SNMP-based platforms)
Scriptingproficiencyin Python or Bash for tooling and automation tasks
Experienceoperatingin colocation / carrier-neutral datacenterenvironments (Equinix, Interxion, or similar)
What You'll Need to Succeed
Comfortoperatingautonomously in a globally distributed, remote-first team across multiple time zones
A bias toward automation: ifyou'vedone something manually twice, you should be automating it
Strong written communication skills -- you will document systems, write runbooks, and produce post-mortems
The ability to context-switch between strategic work (architecture improvements, tooling development) and urgent operational issues (incident response)
Willingness toparticipatein a 24x7 on-call rotation with a fair and well-supported rotation
Nice to Have
Experience withOSTree/ atomic update workflows for OS lifecycle management
Familiarity with DDoS mitigation platforms (Coreroor similar)
Exposure to BGP route reflector concepts and network-functionvirtualization
Experience with Pulp or other on-premises RPM/package repository management systems
Contributions to open-source infrastructure tooling
The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.
Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or@myworkday.com).
Equal Employment Opportunity
It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting [email protected].

