{"id":1202093,"url":"https://alion.io/job/maker-maker-ai-researcher-efficient-inference","title":"RESEARCHER, EFFICIENT INFERENCE","company":{"id":2683108,"name":"Maker Maker AI","domain":"makermaker.ai","url":"https://alion.io/company/maker-maker-ai","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"B","score":75,"open_postings":5,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-27T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":184000,"max_usd":333000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":637},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Knowledge Distillation","optional":false},{"name":"Mixture of Experts","optional":false},{"name":"Model Distillation","optional":false},{"name":"Post-training","optional":false},{"name":"PyTorch","optional":false},{"name":"Quantization","optional":false},{"name":"Speculative Decoding","optional":false},{"name":"Machine Learning","optional":true},{"name":"Multi-Agent Systems","optional":true}],"status":"live","first_seen_at":"2026-05-18T20:46:06Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-27T22:24:41Z","board_verified":true,"closed_at":null,"days_open":132,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":131},"description":"ABOUT THE COMPANY\nWe're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site\nABOUT THE ROLE\nYou'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run.\nThis is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).\nWHAT YOU'LL DO\nResearch and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic\n\nDesign and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding\n\nInvestigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning\n\nRun controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks\n\nCo-design methods with the inference engineering team: push results to where they actually run, not stop at the paper\n\nRead deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack\n\nPublish when the work warrants it; share findings internally\n\nPartner with model and training researchers so efficiency choices align with model architecture and post-training decisions\n\nWHAT WE'RE LOOKING FOR\nStrong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent\n\n5+ years of hands-on research experience\n\nDeep familiarity with both training and inference performance characteristics\n\nFluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving-framework level when methods require it\n\nTrack record of moving efficiency research from prototype to production\n\nStrong statistical expertise: you'd notice a flawed comparison before someone else points it out\n\nStrong written communication\n\nPublished research at NeurIPS, ICML, ICLR, MLSys, or comparable venues\n\nNICE TO HAVE\nPhD in ML, systems, or related field\n\nOpen-source contributions to quantization, speculative-decoding, or efficient-inference libraries\n\nExperience with hardware-aware optimization and accelerator-specific tooling\n\nBackground in numerical methods, low-precision arithmetic, or\n\napproximate computation\n\nTHIS ROLE IS PROBABLY NOT FOR YOU IF\nYou want to focus on pretraining large models from scratch (that's a different role)\n\nYou prefer abstract algorithmic research without hands-on implementation\n\nYou want a fixed benchmark with stable targets (our targets shift with what our models actually need to do)","description_format":"text","description_chars":3161,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"phd","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-24T21:28:19Z"}],"liveness":{"score":7,"band":"cold","label":"Long shot","p_open":1,"p_active":0.263,"p_room":0.28,"age_days":131,"expected_fill_days":24,"reasons":["conf:0","win:tail","crowd:"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/maker-maker-ai-researcher-efficient-inference","json_url":"https://alion.io/job/maker-maker-ai-researcher-efficient-inference.json","meta":{"generated_at":"2026-09-28T02:53:57Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1757,"day_limit":5000,"remaining_today":3243,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}