716,098open jobs
42,633companies
101,517added this week
Browse all
Salary
$32k – $86k per year (Estimated)
Location
Remote (EAEU)

Confirmed on the employer's own hiring board on Sep 22, 2026. First seen by Alion on Sep 23, 2026. Wildberries scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Wildberries is Russia's largest e-commerce platform and online retailer, offering a vast array of consumer goods ranging from apparel and electronics to home essentials. The company focuses on facilitating digital trade for millions of third-party merchants while managing an extensive logistics network of fulfillment centers and local pickup points. Headquartered in Moscow, Russia, it expands commercial access and fast delivery services across multiple international markets.

Направление работы:

ML Platform - платформенная команда в отделе Trust & Safety Wildberries. Мы отвечаем за ML-инфраструктуру модерации контента и карточек товаров. Ежедневно через наши системы проходят десятки миллионов карточек - это сотни миллионов предсказаний по 100+ ML-моделям. Модели крутятся в Nvidia Triton Inference Server на GPU-кластере в нескольких дата-центрах; пиковая нагрузка на инференс - порядка 400-500K RPS, десятки инстансов Triton.

Исторически платформа выросла из модерации, сейчас становится самостоятельным юнитом и расширяется на все направления T&S. В отделе работают 50+ DS, и единой платформы для них пока нет. Мы это меняем - строим общую мультитенантную ML-платформу для всего отдела.

Ищем ML Infrastructure Engineer (MLOps) на инфраструктурный слой платформы: GPU-кластер, разделение ресурсов между командами, ML-тулинг, среда обучения, стандартизация пайплайнов.

Стань частью команды!

Вам предстоит:

  • Отвечать за GPU-кластер целиком: драйверы, GPU Operator, DCGM-мониторинг, утилизация, планирование ёмкости;

  • Выстроить стратегию разделения GPU между командами в гетерогенной среде, квоты и приоритеты;

  • Поднять единый Kubeflow и ClearML как стандарт для DS-команд отдела;

  • Оптимизировать inference-инфраструктуру - автоскейлинг GPU-нод, bin-packing, утилизация;

  • Строить инфраструктуру Embedding Store (pgvector, FAISS, qdrant);

  • Общаться с DS-командами, понимать их потребности и переводить в инфраструктурные решения;

  • Раскатать платформу на остальные команды Trust & Safety с полноценной мультитенантностью.

Формат работы - гибридный или удаленный по договоренности с руководителем.

Вы нам подходите, если:

  • Имеете глубокое понимание Kubernetes: операторы, scheduling, resource management, GPU в K8s (device plugin, GPU Operator);

  • Имеете практический опыт с NVIDIA GPU: драйверы и CUDA-стек, DCGM, MIG или time-slicing;

  • Имеете опыт развёртывания и поддержки MLOps-платформ для DS-команд (ClearML, MLflow, Kubeflow, Airflow или аналоги);

  • Владеете Linux и bare-metal, IaC (Ansible или Terraform), Helm;

  • Имеете опыт разработки на Python или Go: операторы, экспортеры, внутренний тулинг;

  • Умеете взаимодействовать с DS-командами и переводить потребности в технические решения;

Будет плюсом:

  • Опыт с Triton Inference Server, vLLM, KServe или аналогами;

  • Опыт разделения и виртуализации GPU в Kubernetes для multi-tenant окружений;

  • Понимание векторных БД и их оптимизации (pgvector, FAISS, Milvus);

  • Сторадж под данные обучения: S3, Ceph, NFS;

  • Multi-node training: NCCL, InfiniBand, RDMA.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
716,098 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Moscow
MLOps Engineer (MLE) 5 hours ago
$32k – $86k per year (Estimated) • Remote • Moscow
Python
Python
FastAPI
Asyncio
Databases
PostgreSQL
ClickHouse
pgvector
Qdrant
Apache Kafka
OpenSearch
AI/ML
vLLM
MLFlow
Triton Inference Server
ONNX
TensorRT
Kubeflow
PyTorch
ClearML
KServe
Triton
TorchServe
ONNX Runtime
DevOps
gRPC
Terraform
Ansible
Helm
GitLab CI
Kubernetes
Error Budget
SLI/SLO/SLA
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Brașov
Python
DevOps
CI/CD
GitLab
Linux
Windows
Management
Agile
Apply
Geophysicist 1 hour ago
$32k – $89k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Eleusis
Python
Management
Microsoft Office
Apply
$44k – $112k per year (Estimated) • In office • Full-Time • 6+ years exp • Bengaluru
Python
JavaScript
SQL
Node JS
AI/ML
LangGraph
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
NLP
LLM
RAG
Hallucination
OpenAI
Anthropic
Structured Outputs
LLM Evaluation
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Rest API
GCP
GitHub Actions
CI/CD
AWS
Docker
Apply
MDR Services Engineer 5 hours ago
$29k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Athens
Python
PowerShell
Bash
DevOps
Prometheus
Azure
AWS
Cortex
Linux
Windows
Cybersecurity
Crowdstrike
Microsoft Sentinel
Microsoft Defender
IBM QRadar
SIEM
Apply
MLOps Engineer (MLE) 5 hours ago
$32k – $86k per year (Estimated) • Remote • Moscow
Python
Python
FastAPI
Asyncio
Databases
PostgreSQL
ClickHouse
pgvector
Qdrant
Apache Kafka
OpenSearch
AI/ML
vLLM
MLFlow
Triton Inference Server
ONNX
TensorRT
Kubeflow
PyTorch
ClearML
KServe
Triton
TorchServe
ONNX Runtime
DevOps
gRPC
Terraform
Ansible
Helm
GitLab CI
Kubernetes
Error Budget
SLI/SLO/SLA
Apply
$27k – $63k per year (Estimated) • In office • 1+ year exp • Moscow
DevOps
Linux
Unix
TCP/IP
Apply
$36k – $88k per year (Estimated) • Remote • Moscow
AI/ML
Claude Code
NLP
VLM
Transformers
LLM
Apply
$50k – $101k per year (Estimated) • In office • 4+ years exp • Moscow
Python
SQL
Databases
Redis
ClickHouse
Apache Iceberg
Cassandra
Presto
Apache Kafka
Apache Hudi
Trino
Redpanda
AI/ML
Spark
Airflow
dbt
Flink
Feast
Machine Learning
DevOps
SLI/SLO/SLA
Amazon S3
Apply
$36k – $88k per year (Estimated) • Remote • Moscow
AI/ML
Claude Code
NLP
VLM
Transformers
LLM
Apply
$38k – $70k per year (Estimated) • Remote • 6+ years exp • Moscow
Databases
RabbitMQ
Apache Kafka
DevOps
Rest API
SOAP
Management
Confluence
Jira
Draw.io
UML
BPMN
QA
Swagger
Postman
Apply
$21k – $25k per year • In office • 1+ year exp • Moscow
Apply
$21k – $46k per year (Estimated) • Remote/Hybrid • 3+ years exp • Moscow
DevOps
Windows
Apply
$14k – $28k per year (Estimated) • Remote • Part-Time • Moscow
AI/ML
ChatGPT
Apply
$35k – $74k per year (Estimated) • In office • 3+ years exp • Moscow
Python
Python
FastAPI
AI/ML
MLFlow
Fine-tuning
Prompt Engineering
AI Agents
Mistral
Kubeflow
TensorFlow
PyTorch
LLM
RAG
Apply
See all jobs
This is one of many
716,098 more open roles from verified company boards, updated every day.