1,456,462open jobs
86,522companies
226,511added this week
Browse all
Salary
$195k – $262k per year
Location
Remote (United States)also open in Austria, Belgium, Bulgaria, Croatia, Czech Republic, Denmark +25
Seniority
Senior

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Jul 22, 2026.

Overview
Company
Impact
Profile match
Nebius is an international technology company that specializes in building full-stack artificial intelligence infrastructure and high-performance cloud GPU platforms for AI model development. Headquartered in Amsterdam, Netherlands, the firm operates energy-efficient data centers across Europe and North America to provide scalable compute, storage, and software tools for machine learning workloads.

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role  

As a Senior Machine Learning Engineer on the Applied AI team at Nebius Token Factory, you will own model and endpoint optimization from model artifacts through production deployment. Your work will span model internals, quantization and model compression, speculative decoding, KV-cache optimization, inference engines, serving architecture, and benchmarking to improve latency, throughput, memory efficiency, GPU utilization, and cost per token while preserving model quality and reliability.

You will work with kernel and platform engineers to investigate serving problems, compare configurations, and resolve performance and quality regressions. You will evaluate model-and engine-level optimizations together with distributed inference designs, including request routing, scheduling, prefill/decode coordination, and multi-node GPU execution. Your improvements will be validated through reproducible benchmarks and safe production rollouts under real-world workloads. 

Your responsibilities:  

  • Own optimization projects for specific model families, customer endpoints, or serving backends.

  • Compare inference engines and recommend practical serving configurations suited to individual workloads.

  • Diagnose model-quality and performance regressions during production rollouts.

  • Improve LLM and VLMendpoint latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.

  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIADynamo.

  • Develop production model-compression workflows covering quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.

  • Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. For prefill/decode disaggregation (PDD), evaluate KV-cache transfer, worker placement, and capacity balancing to determine which workloads benefit from the architecture.

  • Design and improve LLM request routers and scheduling policies that balance worker utilization queuing, request characteristics, and KV-cache locality while meeting latency and reliability targets 
  • Scale dense and mixture-of-experts inference across multiple GPU nodes. Select and tune tensor, pipeline, data, and expert parallelism - including wide expert parallelism (WideEP) - with attention to hardware topology, expert load balance, and communication overhead. 
  • Build reproducible benchmark harnesses measuring for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. Validate architecture choices under representative traffic and consistent GPU budgets.

  • Collaborate with GPU kernel engineers and platform engineers to trace bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.

  • Produce design documents, performance reports, rollout plans, and customer-facing technical explanations.

Must-haves:  

  • Strong engineering skills in Python and PyTorch.

  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference.

  • Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIADynamo, Ray Serve, KServe, or an equivalent internal system.

  • A strong understanding of transformer inference bottlenecks involving KVcache, attention, memory bandwidth, batching, parallelism, and long-context serving.

  • Ability to quantify trade-offs among latency, throughput, quality, utilization, and cost.

  • Experience designing or optimizing distributed inference system architecture, with hands-on work in one or more areas such as request routing, distributed scheduling, PDD, multi-node inference, or expert parallelism.
  • Strong communication skills and collaboration skills across research, kernel, infrastructure, product, and customer teams.

Nice-to-haves:  

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant.

  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or related inference acceleration methods.

  • Experience supporting agentic workloads involving including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.

  • Familiarity with CUDAor Triton; the role does not require kernel engineering to be the candidate's primary specialization.

  • Contributions to open-source projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, or KServe.

  • Implementation experience with cache-aware request routing, PDD, or WideEP, including diagnosing communication bottlenecks and load imbalance.
  • Familiarity with GPU communication libraries and interconnects, such as NCCL, NVLink, Infiniband, or RoCE.

Key employee benefits in the US:

  • Health insurance:  100% company-paid medical, dental, and vision coverage for employees and families.

  • 401(k) plan:  Up to 4% company match with immediate vesting.

  • Parental leave:  20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.

  • Remote work reimbursement:  Up to $85/month for mobile and internet.

  • Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.

Pay Transparency

We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.

Base Compensation Range

$195,200—$262,200 USD

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,456,462 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Palo Alto
$106k – $205k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Richardson • Santa Clara • Austin
AI/ML
Quantization
MLIR
Apache TVM
IREE
Apply
$219k – $250k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Jose • McLean • Cambridge • New York
AI/ML
Fine-tuning
RLHF
NLP
Transfer Learning
PyTorch
LLM
Tokenization
Self-Supervised Learning
Hugging Face
SFT
Pre-training
Machine Learning
DevOps
AWS
Apply
$350k – $400k per year • In office • Full-Time • 5+ years exp • Master's Degree • McLean • San Francisco • Cambridge • San Jose • New York
AI/ML
DeepSpeed
DGL
Fine-tuning
RLHF
Quantization
Multimodal AI
NLP
Transfer Learning
PyTorch
LLM
Tokenization
Self-Supervised Learning
Time Series Forecasting
Hugging Face
SFT
Pre-training
NVIDIA NeMo
Recommender Systems
Machine Learning
DevOps
AWS
Apply
$197k – $225k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • New York • McLean
Python
Go
Java
C++
Scala
C++
TensorFlow C++
PyTorch C++
AI/ML
Spark
Reinforcement Learning
Scikit-learn
Transformers
TensorFlow
Pandas
NumPy
PyTorch
Ray
Explainable AI
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Management
Agile
Apply
$230k – $262k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • New York • San Francisco • McLean • Cambridge • San Jose
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
RAG
CUDA
Hugging Face
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$200k – $300k per year • Hybrid • Full-Time • 2+ years exp • Palo Alto
Python
TypeScript
Python
Django
AI/ML
AI Agents
LLM
Braintrust
Machine Learning
Apply
$230k – $270k per year • Hybrid • Full-Time • 7+ years exp • Palo Alto
AI/ML
Quantization
Multimodal AI
Time Series Forecasting
Amazon SageMaker
AWS Trainium
DevOps
CI/CD
AWS
Docker
Kubernetes
Apply
$230k – $280k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Palo Alto
Python
AI/ML
AI Agents
Mistral
LLM
OpenAI
Anthropic
LLM Guardrails
Agentic Workflows
Apply
≈ $265k – $543k per year (Estimated) • In office • 5+ years exp • High School Diploma • Palo Alto
AI/ML
AI Agents
DevOps
CI/CD
Apply
$80k – $110k per year • In office • Full-Time • 4+ years exp • Palo Alto
AI/ML
Claude
ChatGPT
Perplexity
Apply
See all jobs
This is one of many
1,456,462 more open roles from verified company boards, updated every day.