{"id":2243371,"url":"https://alion.io/job/nebius-senior-machine-learning-engineer-llm-inference-optimization-2","title":"Senior Machine Learning Engineer, LLM Inference Optimization","company":{"id":6364,"name":"Nebius","domain":"nebius.com","url":"https://alion.io/company/nebius","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Palo Alto, United States","San Francisco, United States"],"countries":["US"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":{"min":195200,"max":262200,"currency":"USD","period":"year","gross":null,"usd_annual":262200},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Knowledge Distillation","optional":false},{"name":"KServe","optional":false},{"name":"KV Cache","optional":false},{"name":"LLM","optional":false},{"name":"Machine Learning","optional":false},{"name":"Mixture of Experts","optional":false},{"name":"Model Distillation","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Quantization","optional":false},{"name":"Ray","optional":false},{"name":"Ray Serve","optional":false},{"name":"SGLang","optional":false},{"name":"Speculative Decoding","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"TPOT","optional":false},{"name":"Triton","optional":false},{"name":"Triton Inference Server","optional":false},{"name":"vLLM","optional":false},{"name":"VLM","optional":false},{"name":"AI Agents","optional":true},{"name":"AWQ","optional":true},{"name":"Function Calling","optional":true},{"name":"GPTQ","optional":true},{"name":"InfiniBand","optional":true},{"name":"NCCL","optional":true},{"name":"NVLink","optional":true},{"name":"Post-training","optional":true},{"name":"Structured Outputs","optional":true},{"name":"TensorRT","optional":true},{"name":"Tool Use","optional":true}],"status":"live","first_seen_at":"2026-07-22T22:54:04Z","employer_posted_date":"2026-10-08","last_verified_at":"2026-10-11T16:04:46Z","board_verified":true,"closed_at":null,"days_open":80,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":80},"description":"About Nebius:\nNebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.\nBuilt by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.\nListed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.\nThe role \nAs a Senior Machine Learning Engineer on the Applied AI team at Nebius Token Factory, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, quantization and model compression, speculative decoding, KV-cache optimization, inference engines, serving architecture, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization, and cost per token while preserving model quality and reliability.\nYou will work with kernel and platform engineers to investigate serving problems, compare configurations, and resolve performance and quality regressions. You will evaluate model-and engine-level optimizations together with distributed inference designs, including request routing, scheduling, prefill/decode coordination, and multi-node GPU execution. Your improvements will be validated through reproducible benchmarks and safe production rollouts under real-world workloads. \nYour responsibilities: \nOwn optimization projects for specific model families, customer endpoints, or serving backends.\n\nCompare inference engines and recommend practical serving configurations suited to individual workloads.\n\nDiagnose model-quality and performance regressions during production rollouts.\n\nImprove LLM and VLMendpoint latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.\n\nDeploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIADynamo.\n\nDevelop production model-compression workflows covering quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.\n\nImplement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. For prefill/decode disaggregation (PDD), evaluate KV-cache transfer, worker placement, and capacity balancing to determine which workloads benefit from the architecture.\n\nDesign and improve LLM request routers and scheduling policies that balance worker utilization queuing, request characteristics, and KV-cache locality while meeting latency and reliability targets \nScale dense and mixture-of-experts inference across multiple GPU nodes. Select and tune tensor, pipeline, data, and expert parallelism - including wide expert parallelism (WideEP) - with attention to hardware topology, expert load balance, and communication overhead. \nBuild reproducible benchmark harnesses measuring for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Validate architecture choices under representative traffic and consistent GPU budgets.\n\nCollaborate with GPU kernel engineers and platform engineers to trace bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.\n\nProduce design documents, performance reports, rollout plans, and customer-facing technical explanations.\n\nMust-haves: \nStrong engineering skills in Python and PyTorch.\n\nHands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference.\n\nPractical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIADynamo, Ray Serve, KServe, or an equivalent internal system.\n\nA strong understanding of transformer inference bottlenecks involving KVcache, attention, memory bandwidth, batching, parallelism, and long-context serving.\n\nAbility to quantify trade-offs among latency, throughput, quality, utilization, and cost.\n\nExperience designing or optimizing distributed inference system architecture, with hands-on work in one or more areas such as request routing, distributed scheduling, PDD, multi-node inference, or expert parallelism.\nStrong communication skills and collaboration skills across research, kernel, infrastructure, product, and customer teams.\n\nNice-to-haves: \nExperience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant.\n\nExperience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or related inference acceleration methods.\n\nExperience supporting agentic workloads involving including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.\n\nFamiliarity with CUDAor Triton; the role does not require kernel engineering to be the candidate's primary specialization.\n\nContributions to open-source projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, or KServe.\n\nImplementation experience with cache-aware request routing, PDD, or WideEP, including diagnosing communication bottlenecks and load imbalance.\nFamiliarity with GPU communication libraries and interconnects, such as NCCL, NVLink, Infiniband, or RoCE.\nKey employee benefits in the US:\nHealth insurance: 100% company-paid medical, dental, and vision coverage for employees and families.\n\n401(k) plan: Up to 4% company match with immediate vesting.\n\nParental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.\n\nRemote work reimbursement: Up to $85/month for mobile and internet.\n\nDisability & life insurance: Company-paid short-term, long-term and life insurance coverage.\n\nPay Transparency\nWe offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.\nBase Compensation Range\n$195,200—$262,200 USD\nBenefits & Perks:\nCompetitive compensation\nCareer growth and learning opportunities\nFlexibility and ownership\nCollaborative and innovative culture\nOpportunity to work on impactful AI projects\nInternational environment and talented teams\nWhat's it like to work at Nebius:\nFast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI \nEqual Opportunity Statement:\nNebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.\nApplicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. \nIf you need accommodations during the application process, please let us know.","description_format":"text","description_chars":7603,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Health insurance","Insurance coverage","Life insurance","Parental leave"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","Education","Data Centers & Colocation"],"lifecycle":[{"event":"open","at":"2026-10-11T01:05:42Z"}],"visa":[],"liveness":{"score":12,"band":"cold","label":"Long shot","p_open":1,"p_active":0.426,"p_room":0.28,"age_days":80,"expected_fill_days":28,"reasons":["conf:3","win:tail","crowd:brand"],"computed_at":"2026-10-11T19:19:04Z"},"pay":{"stated_usd_annual":262200,"is_top_pay":true},"html_url":"https://alion.io/job/nebius-senior-machine-learning-engineer-llm-inference-optimization-2","json_url":"https://alion.io/job/nebius-senior-machine-learning-engineer-llm-inference-optimization-2.json","meta":{"generated_at":"2026-10-11T19:19:04Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":6056,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":6364},"rest":"https://alion.io/mcp/rest/get_company?id=6364"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fnebius-senior-machine-learning-engineer-llm-inference-optimization-2"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fnebius-senior-machine-learning-engineer-llm-inference-optimization-2"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fnebius-senior-machine-learning-engineer-llm-inference-optimization-2"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/nebius-senior-machine-learning-engineer-llm-inference-optimization-2\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fnebius-senior-machine-learning-engineer-llm-inference-optimization-2"}]}