This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Infrastructure Engineer based in United States.
As a Staff Infrastructure Engineer, you will serve as a senior technical leader responsible for building and evolving highly reliable cloud infrastructure at global scale.
You will shape architecture, reliability standards, and platform strategy across mission-critical, high-throughput distributed systems.
The role combines hands-on engineering with broad technical influence, giving you ownership of complex infrastructure challenges from design through operation.
You will modernize infrastructure through automation, Infrastructure-as-Code, GitOps, and self-healing systems while reducing operational toil.
Working closely with Product, Security, and Core Engineering teams, you will create scalable platform capabilities that accelerate development across the organization.
You will also strengthen observability, incident management, security, and resiliency practices across a multi-region cloud environment.
This is an opportunity to make a significant technical impact within a remote-first engineering culture focused on ownership, reliability, and continuous improvement.
Accountabilities
- Architect, build, and optimize highly available, multi-region cloud infrastructure capable of supporting mission-critical, high-throughput applications with strong reliability and low latency.
- Drive the transition from manually configured environments toward standardized, globally scalable infrastructure managed entirely through code.
- Design and implement automated, self-healing infrastructure using Infrastructure-as-Code, GitOps practices, modern CI/CD pipelines, and other automation technologies.
- Establish and continuously improve platform reliability, resiliency, availability, and operational standards across assigned infrastructure domains.
- Build comprehensive observability and telemetry capabilities using technologies such as Prometheus, Grafana, and OpenTelemetry to identify bottlenecks and reliability risks proactively.
- Own incident management activities, lead blameless post-mortems, identify systemic improvements, and continuously strengthen the reliability baseline.
- Develop reusable platform capabilities and "Golden Paths" that reduce friction for engineering teams and accelerate safe, consistent software delivery.
- Partner with Product, Security, and Core Engineering teams at the architectural and design level to enable platform adoption and improve engineering practices.
- Mentor engineers and influence technical standards and practices across the broader engineering organization.
- Architect and operate high-performance, low-latency private connectivity to improve the experience and reliability of external customers.
- Apply security as a core engineering consideration across cloud infrastructure, networking, API gateways, load balancing, access controls, and network isolation.
- Take end-to-end ownership of complex distributed-systems challenges, balancing technical excellence, operational reliability, and business impact.
- 8+ years of experience owning outcomes within complex, large-scale distributed systems or other mission-critical technical environments.
- Advanced AWS expertise, including hands-on experience designing and operating production cloud infrastructure.
- Strong Infrastructure-as-Code experience, particularly with Terraform and reproducible infrastructure environments.
- Extensive hands-on experience with Kubernetes, particularly EKS, as well as Docker and containerized workloads.
- Strong knowledge of GitOps workflows and tools such as Flux, Argo, and GitHub Actions.
- Strong programming and automation capabilities using Python, Go, Bash, or comparable languages.
- Deep experience implementing and operating observability solutions using Prometheus, Grafana, OpenTelemetry, or similar technologies at scale.
- Solid understanding of cloud security, API gateways, load balancing, network isolation, and secure infrastructure design.
- Demonstrated ability to lead complex technical initiatives independently, make sound architectural decisions, and influence engineering practices beyond your immediate team.
- Strong communication, collaboration, and mentoring skills, with the ability to work effectively across engineering and business functions.
- Experience with tokenization, payment processing, cryptography, security products, or other highly regulated or security-sensitive technology environments is a plus.
- Bachelor’s degree in a relevant discipline is preferred.
- Experience with distributed data streaming platforms such as Kafka or Amazon MSK is a plus.
- Experience with database performance tuning and query optimization is beneficial.
- Familiarity with Java and the Spring Framework is advantageous.
- Comfortable working in a fast-paced environment where ownership, curiosity, continuous learning, and creative problem-solving are highly valued.
- Must be based in an eligible U.S. hiring location; certain roles may also be available in select Canadian provinces.
- Competitive annual salary ranging from $145,000 to $260,000.
- 401(k) with a 4% company match and immediate vesting for U.S. employees.
- Early equity opportunities.
- Medical, dental, and vision insurance.
- HSA and FSA options.
- Life and disability insurance.
- Pet insurance.
- Unlimited paid time off.
- Paid national holidays.
- Global parental leave program.
- Internet stipend.
- Learning and professional development stipend.
- Office setup stipend for U.S. employees.
- Flexible working hours.
- Remote and hybrid work opportunities.
- Opportunities for internal promotion and career growth.
- Lunch-and-learn sessions, team events, and company summits.
- Weekly office lunches and employee discount programs.

