A well-resourced AI research and infrastructure team building models that keep learning after deployment, along with the inference systems that serve them at frontier-level speed and cost. Small, technically deep team; distributed by default with regular in-person collaboration.
Requirements
Strong first-principles reasoning and deep technical expertise (e.g. ML systems, compilers, mathematics, physics, competitive programming, security, or HPC)
Hands-on experience with GPU-level performance engineering (Triton/CUDA) and/or model training and post-training systems
Creative, high-agency builder who has mastered something hard, with a track record of tackling genuinely novel technical problems
Comfortable moving fluidly between open-ended research and shipping production code
Responsibilities
Design and optimize inference systems that serve continuously adapting models with frontier-level performance and economics
Develop hardware-informed quantization approaches and write/optimize Triton and CUDA kernels
Convert agent traces into synthetic reinforcement learning environments for continued model training
Identify and resolve performance bottlenecks across compute, memory, and communication
Efficiently deploy LoRAs, learned memory, specialized weights, and other adaptation techniques in production
Implement policy updates derived from live model interactions and post-deployment feedback loops

