You will own the entire path from camera to customer: real-time video ingestion from factory floors with unreliable networks, GPU inference at scale, stream recovery, autoscaling, cost engineering, and the reliability of a system customers watch live on a dashboard all day. Architecture is not handed to you; you are the architecture. The uptime bar is absolute. Unlike most software, where an hour of downtime just means users resume an hour later, here every minute down means real people on the shop floor are blind, and that minute of the factory's day is unrecorded and unrecoverable. Downtime is a permanent loss for a customer running a live operation, and you'll build like that's true, because it is. It cares very little about specific keywords or whether you've touched its exact stack; it cares whether you've taken a hard real-time system and scaled it an order of magnitude, under production pressure, as the person responsible.
Responsibilities:
- The entire path from camera to customer: real-time video ingestion over unreliable factory networks, GPU inference at scale, stream recovery, autoscaling, and cost engineering.
- The reliability of a live system customers watch all day, held to an absolute uptime bar.
- End-to-end production ownership: design, build, deploy, on-call, and the consequences.
- Company-wide infrastructure, working directly with the founders
Requirements:
- Must-Have (at least the core signal):
- Scaled a real-time or high-throughput system (video, voice, streaming, trading, telemetry; the domain matters less than the physics) ~10x under production pressure, as the owner.
- Deep systems fluency: concurrency and backpressure, GPU utilization, Kubernetes autoscaling, designing for partial failure, and major cost reduction with the reasoning behind it.

