This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Platform Engineer based in the United States.
The Staff Platform Engineer will lead the design and delivery of scalable cloud infrastructure, CI/CD systems, and platform tooling across complex engineering environments.
You will architect cloud-native solutions that prioritize reliability, resilience, security, observability, and cost efficiency.
The role combines hands-on engineering with growing ownership of platform strategy and DevSecOps practices.
You will work closely with engineering, QA, and product teams throughout the software development lifecycle to enable fast and dependable delivery.
A key focus will be building and operating infrastructure for modern AI/ML workloads, including model serving, GPU orchestration, LLM gateways, and vector storage.
You will also contribute to platform standards, incident response, technical decision-making, and the mentorship of other engineers.
This is an opportunity to join a highly experienced technical environment where ownership, craftsmanship, adaptability, and production impact are valued.
Accountabilities:
- Lead the end-to-end design and implementation of CI/CD pipelines, improving quality, reliability, and efficiency across engineering delivery workflows.
- Architect and implement cloud-native infrastructure designed for scalability, resilience, performance, and cost efficiency.
- Design, deploy, and manage Kubernetes clusters and containerized workloads at scale.
- Implement and maintain infrastructure as code across multiple environments.
- Drive observability, performance optimization, monitoring, and alerting across production systems.
- Embed DevSecOps practices into engineering workflows, including security scanning, secrets management, and access control.
- Lead production incident response, root cause analysis, remediation, and post-incident reviews.
- Design and operate AI/ML platform infrastructure, including model serving and deployment, GPU workload orchestration, LLM gateways and observability, vector store infrastructure, and CI/CD for AI/ML systems.
- Apply an AI-forward approach to daily engineering work by leveraging modern AI assistants to improve productivity, quality, and delivery speed.
- Collaborate with engineering, QA, and product teams across the full software development lifecycle to align infrastructure with product and delivery objectives.
- Communicate infrastructure decisions, technical tradeoffs, risks, and recommendations clearly across technical and cross-functional teams.
- Participate in architecture and design reviews, sprint ceremonies, release planning, and other engineering activities.
- Lead platform initiatives end-to-end while progressively taking greater ownership of infrastructure strategy.
- Establish and contribute to platform standards, engineering best practices, and reusable approaches that improve reliability and consistency.
- Mentor junior engineers, share technical knowledge, and support the development of platform engineering capabilities across the team.
- Design, build, and maintain reusable AWS CDK constructs for Amazon EKS environments that enable efficient provisioning and management of Kubernetes infrastructure.
- 5-7 years of professional experience in DevOps or platform engineering, with demonstrated growth in infrastructure ownership and technical leadership.
- Strong scripting and programming capabilities in languages such as Python, Go, and Bash.
- Hands-on experience with at least one major cloud platform and strong understanding of its services and capabilities.
- Strong Kubernetes and container orchestration expertise.
- Extensive experience with infrastructure as code.
- Proven ability to design, implement, and own CI/CD pipelines.
- Experience with observability, monitoring, alerting, and production performance management.
- Strong understanding of cloud networking, security, IAM, and related infrastructure controls.
- Familiarity with microservices and distributed systems architecture.
- Experience designing or operating AI/ML platform infrastructure, including model serving, deployment, GPU workloads, LLM gateways, observability, vector stores, and AI/ML CI/CD.
- Demonstrable day-to-day experience using AI-forward tools such as Claude, Cursor, or similar technologies.
- Strong problem-solving skills and sound judgment when navigating ambiguous or complex technical challenges.
- Deep hands-on Amazon EKS experience is highly valued, including designing, building, and maintaining reusable AWS CDK constructs.
- Experience with service mesh, multi-cloud or hybrid environments, or relevant cloud certifications is a plus.
- Strong ownership mindset, with the ability to identify problems, take action, and follow initiatives through to completion.
- Adaptability and curiosity, with a willingness to continuously learn as cloud and AI technologies evolve.
- Clear and direct communication style, with the confidence to provide constructive feedback and engage in difficult technical discussions.
- Strong attention to engineering craftsmanship and a commitment to delivering high-quality production systems.
- Ability to work effectively in fast-moving environments where timelines, budgets, and resources may be constrained.
- Comfortable collaborating with highly experienced engineering teams and contributing to a culture of continuous improvement.
- Salary range of $109,000-$150,000 USD, with compensation determined based on factors including qualifications, experience, skills, seniority, location, performance, travel requirements, and organizational needs.
- Paid time off.
- Medical, dental, and vision insurance for eligible employees.
- 401(k) for eligible employees.
- Fully remote position open to candidates based anywhere in the United States.
- Opportunity to work on production-ready AI systems and modern cloud infrastructure.
- Collaborative environment with highly experienced engineering professionals.
- Opportunities to contribute to platform strategy, technical standards, and engineering best practices.
- Mentorship and professional growth opportunities.
- Equal employment opportunity regardless of legally protected characteristics.
Requirements:
Benefits:

