{"id":1230908,"url":"https://alion.io/job/archetype-ai-machine-learning-systems-engineer","title":"Machine Learning Systems Engineer","company":{"id":1814726,"name":"Archetype AI","domain":"archetypeai.io","url":"https://alion.io/company/archetypeai","size_band":"11-50","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Kula","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Mateo, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":166000,"max_usd":302000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":637},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"Linux","optional":false},{"name":"Multimodal AI","optional":false},{"name":"Physical AI","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"Quantization","optional":false},{"name":"Rust","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"CUTLASS","optional":true},{"name":"KV Cache","optional":true},{"name":"SGLang","optional":true},{"name":"TensorRT","optional":true},{"name":"TensorRT-LLM","optional":true},{"name":"Time Series Forecasting","optional":true},{"name":"Triton","optional":true},{"name":"vLLM","optional":true}],"status":"live","first_seen_at":"2026-09-24T14:38:56Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-25T22:29:54Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About Job\nAt Archetype AI, we’re building the world’s first physical AI platform to bring artificial intelligence into the real world. Our foundation model, Newton, understands the physical world through objective sensor data and generates real-time insights into complex physical behaviors, from industrial machinery and systems to wearable devices and smart environments.\nFormed by a high-caliber team from Google and backed by one of Silicon Valley’s most renowned venture funds, Archetype AI is in a Series A phase and rapidly advancing its technology for the next big leap. This is a unique opportunity to join an exciting, fast-growing AI team based in the heart of Silicon Valley.\nAbout the Role\nYou will own the serving path for Newton and related multimodal models. Much of our inference stack is Rust-native: model nodes in our agent runtime, built on Rust ML stacks (candle, Burn) with custom GPU kernels, plus the routing layer that streams real-time inference to GPU nodes. You will drive GPU utilization, numerical precision, and low-latency serving from the kernel up.\nWhat You'll Own\nBuild and own model nodes in our Rust inference runtime: loading, warmup, batching, streaming, GPU memory pools.\n\nOptimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, parity against research.\n\nOwn inference routing and serving: streaming path API to GPU node, request batching, SLO-backed latency and cost.\n\nProductionize research checkpoints: export, compilation, quantization, parity evals, and rollout.\n\nBuild the observability inference needs: latency histograms, GPU metrics, OOM signatures, replayable traces.\n\nKey Qualifications\n6+ years software engineering, several of them in ML systems, inference, or high-performance GPU computing.\n\nHas owned a production serving path end to end, not only benchmarked models.\n\nExpert in PyTorch with models shipped to production; knows what will be slow before the profiler runs.\n\nStrong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight or equivalent profiling.\n\nRust or C++ alongside Python, Linux performance, production ops; ready to work in Rust daily.\n\nWorks with researchers: can translate an architecture change into serving work.\n\nNice to Have\nRust ML stacks: candle, Burn, or comparable GPU compute in Rust.\n\nCustom kernels and compiler stacks: Triton, CUTLASS, TorchInductor, TensorRT.\n\nQuantization and mixed precision in production with a numerical-correctness suite.\n\nMultimodal, video, embedding or time-series serving, not only decoder-only chat LLMs.\n\nHigh-performance serving stacks (vLLM, SGLang, TensorRT-LLM): continuous batching, paged KV cache.","description_format":"text","description_chars":2663,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Foundation Models","Industrial AI"],"lifecycle":[{"event":"open","at":"2026-09-25T14:38:56Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":31,"reasons":["conf:3","win:early"],"computed_at":"2026-09-26T02:23:09Z"},"pay":null,"html_url":"https://alion.io/job/archetype-ai-machine-learning-systems-engineer","json_url":"https://alion.io/job/archetype-ai-machine-learning-systems-engineer.json","meta":{"generated_at":"2026-09-26T02:23:09Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2433,"day_limit":5000,"remaining_today":2567,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}