We are looking for a seasoned Platform Engineer to join our core infrastructure team and drive the next generation of our large-scale distributed systems. You will own mission-critical components that underpin products serving hundreds of millions of users worldwide, writing high-performance, production-grade code in C++, Go, and/or Rust. This is a high-impact, high-autonomy role. You will collaborate with world-class engineers across infrastructure, reliability, security, and product engineering to define platform primitives, establish engineering best practices, and mentor the next generation of systems engineers.
The core responsibilities for the job include the following:
Systems Design and Core Engineering:
- Design, build, and maintain high-performance platform services and infrastructure components in C++20 Go, and/or Rust that operate at petabyte scale and handle millions of QPS.
- Architect low-latency, fault-tolerant distributed systems including service meshes, RPC frameworks, data pipelines, and storage engines.
- Own the full lifecycle of platform components from design docs and RFC proposals through implementation, testing, rollout, and on-call.
- Identify and eliminate performance bottlenecks through profiling, benchmarking, and first-principles analysis of CPU, memory, I/O, and network behavior.
Platform Reliability and Observability:
- Define and drive SLOs/SLAs for platform services; lead blameless post-mortems and implement systemic fixes.
- Build and extend internal observability frameworks: distributed tracing, metrics pipelines, structured logging, and alerting.
- Drive chaos engineering and fault-injection initiatives to proactively surface reliability gaps at enterprise scale.
Developer Experience and Ecosystem:
- Build internal tooling, SDKs, and abstractions that raise the productivity of hundreds of engineers across the organization.
- Establish and enforce build system standards (Bazel / CMake), dependency management, and cross-language interop (C++/Go/Rust FFI).
- Champion security-by-default practices: memory safety, sandboxing, supply-chain integrity, and secure-coding guidelines.
Leadership and Collaboration:
- Author technically rigorous design documents and lead design reviews with senior and principal engineers.
- Mentor and code-review junior and mid-level engineers; raise the engineering bar across the platform org.
Platform Reliability and Observability:
- Define and drive SLOs/SLAs for platform services; lead blameless post-mortems and implement systemic fixes.
- Build and extend internal observability frameworks: distributed tracing, metrics pipelines, structured logging, and alerting.
- Drive chaos engineering and fault-injection initiatives to proactively surface reliability gaps at enterprise scale.
Developer Experience and Ecosystem:
- Build internal tooling, SDKs, and abstractions that raise the productivity of hundreds of engineers across the organization.
- Establish and enforce build system standards (Bazel / CMake), dependency management, and cross-language interop (C++/Go/Rust FFI).
- Champion security-by-default practices: memory safety, sandboxing, supply-chain integrity, and secure-coding guidelines.
Leadership and Collaboration:
- Author technically rigorous design documents and lead design reviews with senior and principal engineers.
- Mentor and code-review junior and mid-level engineers; raise the engineering bar across the platform org.

