About the Role
This is a backend-architecture-heavy Platform Engineer role sitting within an engineering team of roughly 15 people, including researchers, AI specialists, and serial startup builders. You will own the reliability, scale, performance, and developer experience of core infrastructure and systems, with direct impact on how fast, efficient, and cost-effective the platform is to build on and run.
What You'll Do
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Write clean, maintainable production code to automate systems, improve backend services, and create internal developer tooling.
What We're Looking For
2 to 4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost.
AI-specific infrastructure experience is required, not just general full-stack or DevOps background.
Deep hands-on experience with AWS and containerized systems: Terraform, Kubernetes/EKS, Docker, EC2, networking, load balancers, and secrets management.
Track record building or operating CI/CD, release automation, observability, alerting, and incident response systems.
Strong backend engineering judgment: service architecture, APIs, databases, async systems, queues, and production failure modes.
Experience designing for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
Demonstrated focus on reducing cloud spend through architecture improvements, autoscaling, workload placement, caching, or cleanup systems.
Ability to write production code and apply software engineering judgment across infrastructure, backend systems, and developer workflows.
Strong ownership mindset and ability to learn quickly in a fast-moving environment.
Compensation & Benefits
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore. Candidates based in the US must be located in San Francisco. Candidates in Europe may be considered as fully remote independent contractors.

