1,454,406open jobs
86,438companies
225,454added this week
Browse all
Salary
$195k – $262k per year
Location
Remote (United States)
Seniority
Senior

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Jul 22, 2026.

Overview
Company
Impact
Profile match
Nebius is an international technology company that specializes in building full-stack artificial intelligence infrastructure and high-performance cloud GPU platforms for AI model development. Headquartered in Amsterdam, Netherlands, the firm operates energy-efficient data centers across Europe and North America to provide scalable compute, storage, and software tools for machine learning workloads.

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role  

Nebius Token Factory is looking for scientists who can turn frontier LLM and VLM inference bottlenecks into research questions, conduct rigorous experiments, and translate their findings into production capabilities. You will own focused research and optimization projects, write strong code, and produce prototypes, publications, and technical artifacts that engineers can build on.

Your research will span quantization, distillation, speculative decoding, KV-cache optimization, and model/runtime co-optimization, alongside inference engines and distributed inference architectures. You will investigate how these approaches interact, evaluate their trade-offs across model quality, latency, throughput, memory footprint, and cost per token, and work with engineering teams to bring promising results into production.

Your responsibilities:  

  • Lead research projects in efficient LLM and VLM inference, from hypotheses and experiments through ablations, prototypes, and production handoff.
  • Develop and evaluate methods spanning quantization, quantization-aware training (QAT), distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.
  • Build prototypes using PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks, and collaborate with MLEs and platform engineers to turn them into production components.
  • Research LLM request routing and scheduling strategies, including cache-aware load balancing and prefill-decode disaggregation (PDD). Evaluate how queueing, KV-cache transfer, worker placement, and prefill/decode capacity allocation affect latency, throughput, and serving cost.
  • Investigate multi-node inference for dense and mixture-of-experts models, including wide expert parallelism (WideEP). Study expert placement, load imbalance, and computation/communication trade-offs, and use the findings to guide model and system architecture choices.
  • Develop rigorous evaluations covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token. Compare serving architectures under consistent workload conditions and GPU budgets.
  • Work with MLE, GPU kernel, backend infrastructure, product, and customer teams to select research priorities with measurable production impact.
  • Share results through internal reports, technical blogs, papers, and open-source artifacts, and mentor engineers and scientists on experimental design, scientific rigor, and model/system trade-offs.

Must-haves:  

  • A PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied mathematics, or a closely related discipline.
  • A strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas.

  • Strong Python and PyTorch implementation skills, with the ability to turn ideas into experiments and working prototypes.

  • Deep knowledge of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving trade-offs.

  • Strong experimental design skills covering ablations, baselines, metrics, statistical reasoning, and failure analysis.

  • Excellent written and verbal communication.

Nice-to-haves:  

  • First-author publications at venues such as NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, or HPCA.

  • Experience deploying ML models or inference optimizations in production.

  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.

  • Experience applying post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation to inference quality or efficiency.

  • Open-source research artifacts, widely used benchmarks, technical blogs, or invited talks demonstrating contributions to efficient AI systems.

  • Research or implementation experience in one or more areas of distributed inference system architecture, such as LLM request routing, PDD, multi-node serving, or MoE expert parallelism, including WideEP.

Key employee benefits in the US:

  • Health insurance:  100% company-paid medical, dental, and vision coverage for employees and families.

  • 401(k) plan:  Up to 4% company match with immediate vesting.

  • Parental leave:  20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.

  • Remote work reimbursement:  Up to $85/month for mobile and internet.

  • Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.

Pay Transparency

We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.

Base Compensation Range

$195,200—$262,200 USD

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,454,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Palo Alto
$129k – $216k per year • Equity • In office • 4+ years exp • Bachelor's Degree • Santa Rosa
Python
JavaScript
TypeScript
C#
AI/ML
Copilot
Embeddings
AI Agents
TensorFlow
PyTorch
LLM
RAG
Frontend
GraphQL
Angular
React.js
DevOps
Rest API
VMWare
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Bitbucket
GitHub
Cybersecurity
Least Privilege
Management
Confluence
Jira
Apply
$129k – $216k per year • Equity • In office • Master's Degree • Calabasas
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Reinforcement Learning
Scikit-learn
TensorFlow
PyTorch
Machine Learning
Chips/EDA
Cadence Virtuoso
Management
Agile
Apply
≈ $132k – $251k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Colorado Springs
Python
C#
C++
AI/ML
LangGraph
AutoGen
LangChain
Prompt Engineering
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Tool Use
DevOps
CI/CD
Git
Bitbucket
Management
Confluence
Jira
Agile
Apply
AI Software Architect 10 days ago
≈ $182k – $375k per year (Estimated) • Remote (Romania) • Full-Time • 7+ years exp • PhD
Databases
Weaviate
Neo4j
Databricks
Milvus
Pinecone
Qdrant
AI/ML
AutoGen
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
OpenAI
LLMOps
GraphRAG
Knowledge Graph
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GitHub Actions
Azure
CI/CD
Apply
≈ $45k – $114k per year (Estimated) • Remote (Romania) • Full-Time • Romania
C++
C++
PyTorch C++
AI/ML
Qwen
OpenCV
YOLO
Fine-tuning
Embeddings
Multimodal AI
Knowledge Distillation
Computer Vision
ONNX
OpenVINO
PyTorch
CLIP
Data Augmentation
OCR
ONNX Runtime
Model Distillation
Apply
$200k – $300k per year • Hybrid • Full-Time • 2+ years exp • Palo Alto
Python
TypeScript
Python
Django
AI/ML
AI Agents
LLM
Braintrust
Machine Learning
Apply
$230k – $270k per year • Hybrid • Full-Time • 7+ years exp • Palo Alto
AI/ML
Quantization
Multimodal AI
Time Series Forecasting
Amazon SageMaker
AWS Trainium
DevOps
CI/CD
AWS
Docker
Kubernetes
Apply
$230k – $280k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Palo Alto
Python
AI/ML
AI Agents
Mistral
LLM
OpenAI
Anthropic
LLM Guardrails
Agentic Workflows
Apply
≈ $265k – $543k per year (Estimated) • In office • 5+ years exp • High School Diploma • Palo Alto
AI/ML
AI Agents
DevOps
CI/CD
Apply
$80k – $110k per year • In office • Full-Time • 4+ years exp • Palo Alto
AI/ML
Claude
ChatGPT
Perplexity
Apply
See all jobs
This is one of many
1,454,406 more open roles from verified company boards, updated every day.