This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Platform Engineer, Cloud Infrastructure based in United States.
This is a senior, hands-on platform engineering opportunity focused on building and operating resilient cloud-native infrastructure at scale.
You’ll own production Kubernetes platforms and the underlying systems that engineering teams rely on to deliver reliable products.
The role combines deep infrastructure engineering with software development in Go, Python, or Java.
You’ll shape platform architecture across networking, service mesh, observability, infrastructure as code, and production operations.
You’ll also lead complex, multi-quarter initiatives, including cloud migrations, reliability improvements, and platform modernization.
Success means creating scalable capabilities, improving developer velocity, and ensuring systems remain secure, observable, and dependable.
Working fully remotely, you’ll collaborate with highly technical distributed teams while mentoring engineers and influencing technical direction.
Accountabilities
- Design, build, and operate production Kubernetes clusters, including networking, workload isolation, resource management, scheduling, and multi-region architectures.
- Work deeply with Kubernetes internals, including CNI networking, NetworkPolicy enforcement, cluster behavior under load, resource quotas, and custom controllers or operators.
- Design and operate service mesh capabilities, including mTLS, workload identity, service-account authentication and authorization, and traffic management.
- Optimize containerized workloads for performance, cost efficiency, scalability, and effective resource utilization.
- Develop and maintain production-grade services, Kubernetes controllers, middleware, and platform components using Go, Python, or Java.
- Build HTTP, REST, and gRPC interfaces used by internal engineering teams, while maintaining strong unit and integration test coverage.
- Lead platform-level incident response, troubleshoot distributed systems using logs, metrics, traces, and profiling, and produce postmortems that drive lasting improvements.
- Define and implement SLOs, actionable alerting, dashboards, metrics, and distributed tracing to strengthen platform reliability.
- Own infrastructure as code at scale, creating reusable modules and evolving them as platform requirements change.
- Build and improve CI/CD and GitOps workflows that enable engineering teams to release software safely, frequently, and reliably.
- Plan and execute cloud migration initiatives, including moving production workloads between providers or environments while minimizing disruption and maintaining reliability.
- Partner with product, security, and infrastructure teams to gather requirements, evaluate trade-offs, conduct design reviews, and establish technical direction.
- Mentor engineers and contribute to higher standards of engineering, reliability, testing, and operational excellence.
- 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including substantial experience operating production distributed systems.
- Demonstrated experience building and operating production Kubernetes platforms, rather than simply deploying applications onto existing clusters.
- Strong production programming experience in Go, Python, or Java, with the ability to work effectively in complex existing codebases.
- Experience taking ambiguous technical problems from initial design through production implementation and ongoing operation.
- Proven experience planning and executing cloud migrations involving production workloads across providers or environments.
- Degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- Strong understanding of Kubernetes networking, CNI, NetworkPolicy, resource management, scheduling, and cluster behavior.
- Hands-on experience with service mesh technologies such as Istio, Envoy, Linkerd, or equivalent, including mTLS and workload identity.
- Strong Linux fundamentals, including cgroups and resource management.
- Significant infrastructure-as-code experience using Terraform or an equivalent technology.
- Production experience with at least one major cloud platform such as AWS, GCP, or Azure; multi-cloud experience is highly valuable.
- Experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, and PromQL or comparable query languages.
- Production experience with relational databases, particularly PostgreSQL or managed PostgreSQL-compatible services, including an understanding of replication and failover.
- Strong Docker and container tooling experience across the software delivery lifecycle.
- Advanced debugging, troubleshooting, and performance-profiling capabilities.
- Experience with Kubernetes controllers, operators, API-server extensions, or Go testing frameworks such as Ginkgo and Gomega is an asset.
- Additional Python expertise and the ability to read Java are valuable.
- Experience with identity and access technologies such as SSO, Keycloak, OIDC, SAML, Vault, or cloud-based secrets management is preferred.
- Experience designing high-availability and disaster-recovery architectures across regions or cloud providers is an advantage.
- Familiarity with monorepos, Bazel, compliance frameworks such as SOC 2 or GDPR, and reliability considerations for LLM-backed systems is beneficial.
- Experience with Alibaba Cloud and large-scale data technologies such as Apache Spark or Apache Flink is highly preferred.
- Strong analytical and problem-solving skills, excellent written and verbal communication, and the ability to produce clear design documents and postmortems.
- Comfortable working independently in a distributed, fast-moving, highly technical environment while collaborating effectively across teams.
- Fully remote position open to candidates in United States.
- Senior-level opportunity with significant technical ownership and influence over cloud infrastructure.
- Opportunity to work with Kubernetes, cloud platforms, service mesh, observability, infrastructure as code, and modern software engineering practices.
- Exposure to complex, multi-quarter initiatives including cloud migrations and large-scale reliability improvements.
- Collaborative distributed environment with opportunities to mentor engineers and shape engineering standards.
- Flexible remote collaboration across an international technical team.
- Opportunity to contribute to innovative cloud-native infrastructure and emerging technologies, including AI-enabled systems.

