{"id":1262468,"url":"https://alion.io/job/cumulus-labs-founding-engineer-ml-platforms-engineer","title":"Founding Engineer — ML Platforms Engineer","company":{"id":3779705,"name":"Cumulus Labs","domain":"cumuluslabs.io","url":"https://alion.io/company/cumulus-labs","size_band":"11-50","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Work at a Startup","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":150000,"max":300000,"currency":"USD","period":"year","gross":null,"usd_annual":300000},"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"Kubernetes","optional":false},{"name":"Claude Code","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-25T20:57:57Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-30T02:40:11Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"About the role\nCumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're looking for an ML Platforms Engineer to help build and run the orchestration layer underneath our inference and agent products, the system that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization.\nWe care more about how you think than which languages are on your resume. Our stack today includes Go, Kubernetes, and Terraform, but we're looking for someone who can walk into any part of a production system, understand it, and make it better, not someone who only knows one toolchain.\nWhat you'll do\nBuild and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime\n\nDesign and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant-facing inference gateway\n\nOwn observability for the fleet: metrics, logs, and traces at scale\n\nDebug hard, systems-level problems across the stack, from scheduling logic down to GPU memory and networking\n\nMake real architectural decisions, not just implement someone else's design\n\nShip fast, own your systems end to end, and work directly with the founder\n\nWhat we're looking for\nExcellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem\n\nReal production experience, ideally with systems that had to stay up and scale under load\n\nStrong design instincts: you can reason about tradeoffs, not just follow a framework's conventions\n\nFast learner who can go deep in unfamiliar territory; specific experience with Go or Kubernetes is a plus, not a requirement\n\nComfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor\n\nYou want to work in person, in a small team, solving problems nobody has solved before\n\nWhy Cumulus\nWe're small, early, and building the systems layer for the next generation of AI infrastructure. You'll own real infrastructure from day one, not tickets in a backlog.\n1. Intro call with a founder (30 min) - background, mutual fit\n2. Technical deep dive (30-45 min) - a real problem from our orchestrator, discussed live\n3. Take-home or paired session on a scoped systems problem\n4. References\nWe move fast: most candidates hear back within a few days at each stage, and we aim to get from first contact to offer in under two weeks.","description_format":"text","description_chars":2472,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence"],"lifecycle":[{"event":"open","at":"2026-09-25T20:57:57Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":3,"expected_fill_days":22,"reasons":["conf:0","win:early"],"computed_at":"2026-09-29T05:45:00Z"},"pay":{"stated_usd_annual":300000,"is_top_pay":true},"html_url":"https://alion.io/job/cumulus-labs-founding-engineer-ml-platforms-engineer","json_url":"https://alion.io/job/cumulus-labs-founding-engineer-ml-platforms-engineer.json","meta":{"generated_at":"2026-09-30T05:39:05Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4054,"day_limit":5000,"remaining_today":946,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}