514,690open jobs
17,495companies
76,717added this week
Browse all
Salary
$160k – $240k per year
Location
In office
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Headquartered in New York City, New York, Bloomberg is a global leader in financial technology, data, and business media. The company is best known for its proprietary Bloomberg Terminal, which provides financial professionals with real-time market analytics, trading tools, and execution capabilities. Through its multi-platform news division, it delivers economic reporting, research, and analysis across television, digital, print, and audio outlets worldwide.

Our team:

Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference Platform provides production-grade managed infrastructure for hosting, deploying, and serving all machine learning models, both predictive and cutting-edge generative models. We abstract away infrastructure complexity, empowering engineering teams to focus on creating intelligent applications with guaranteed scalability, performance, and governance. Our platform is built on the open-source KServe project, and the CNCS AI Inference team is a primary contributor to its development.

We'll trust you to:

  • Design and build scalable infrastructure for both online and offline inference workloads.
  • Lead integration of high-performance inference runtimes and serving frameworks, including TensorRT, vLLM, ONNX, and Triton.
  • Drive architecture and technical decisions across Bloomberg’s inference platform, balancing latency, throughput, reliability, and cost.
  • Partner across engineering teams to improve model deployment, observability, and production performance.
  • Mentor junior engineers on system design, debugging, and performance optimization.

You'll need to have:

  • 5+ years of professional software engineering experience.
  • Experience designing, building, and operating production distributed systems.
  • Strong systems intuition and a track record of debugging and optimizing performance-critical services.
  • Ability to own problems end-to-end and quickly ramp up in unfamiliar technical areas.
  • 4+ years of demonstrated experience working with an object-oriented programming language.
  • A degree in Computer Science, Electrical Engineering, or equivalent practical experience.

We'd love to see:

  • Experience deploying and operating machine learning systems at scale.
  • Experience with inference optimization techniques such as batching, caching, request scheduling, or memory-aware serving.
  • Familiarity with PyTorch and GPU software stacks such as CUDA and NCCL.
  • Exposure to high-performance interconnects and distributed computing technologies such as NVLink, InfiniBand, or MPI.
  • Experience with Kubernetes and cloud-native infrastructure.
  • Experience with load balancing, request routing, or traffic management systems.

Representative projects:

  • Autoscaling a heterogeneous compute fleet to match supply and demand aross diverse inference workloads.
  • Building production-grade deployment pipelines to safely roll out new models to millions of users.
  • Developing new inference capabilities such as structured sampling, prompt caching, and advanced serving optimizations.
  • Analyzing observability data from real production workloads to improve latency, throughput, and resource efficiency.

Salary Range = 160,000 - 240,000 USD Annual + Benefits + Bonus

The referenced salary range is based on the Company's good faith belief at the time of posting. Actual compensation may vary based on factors such as geographic location, work experience, market conditions, education/training and skill level.

We offer one of the most comprehensive and generous benefits plans available and offer a range of total rewards that may include merit increases, incentive compensation (exempt roles only), paid holidays, paid time off, medical, dental, vision, short and long term disability benefits, 401(k) +match, life insurance, and various wellness programs, among others. The Company does not provide benefits directly to contingent workers/contractors and interns.

Discover what makes Bloomberg unique - watch our podcast series for an inside look at our culture, values, and the people behind our success.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
514,690 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$120k – $317k per year (Estimated) • Remote/Hybrid • Internship • Bachelor's Degree • San Francisco
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
TensorRT
PyTorch
LLM
Tokenization
CUDA
Edge AI
ExecuTorch
ONNX Runtime
Speculative Decoding
KV Cache
Vision-Language-Action
DevOps
eBPF
Apply
$98k – $201k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree • Courbevoie
AI/ML
Qwen
DeepSeek
vLLM
CUDA Toolkit
Fine-tuning
AI Agents
SGLang
TensorRT-LLM
Mistral
RAG
MiniMax
CUDA
Triton
Hugging Face
NVIDIA NeMo
LLM Evaluation
LLM Guardrails
EU AI Act
DevOps
Kubernetes
Cybersecurity
Zero Trust
Apply
$111k – $177k per year • Equity • Remote • Full-Time • 17+ years exp • Bachelor's Degree • United States
Python
Java
Python
Dask
Java
Hazelcast
Databases
Redis
Cassandra
Milvus
Pinecone
Presto
Qdrant
ScyllaDB
RabbitMQ
Apache Kafka
Apache Pulsar
Trino
Apache Ignite
AI/ML
Spark
vLLM
SGLang
TGI
Flink
LLM
RAG
Ray
Triton
InfiniBand
DevOps
GCP
OpenShift
Rancher
VMWare
Azure
GitOps
ArgoCD
AWS
Kubernetes
Amazon EKS
Google GKE
Azure AKS
Game Dev
Coherence
Apply
$161k – $241k per year • In office • Secret • Full-Time • Colorado Springs
Python
Java
C++
MATLAB
MATLAB
Signal Processing Toolbox
AI/ML
CUDA Toolkit
CUDA
Apply
$179k – $205k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • McLean • Plano
Python
Java
Scala
Python
Dask
AI/ML
Spark
Scikit-learn
TensorFlow
PyTorch
Explainable AI
DevOps
GCP
Azure
CI/CD
AWS
Management
Agile
Apply
$110k – $210k per year • Remote/Hybrid • 4+ years exp • Bachelor's Degree
AI/ML
LLM
Time Series Forecasting
Management
Agile
Scrum
Apply
$155k – $285k per year • In office • 4+ years exp • Bachelor's Degree
SAS
Apply
Apply
Apply
$160k – $240k per year • In office • 4+ years exp
Python
Java
DevOps
Terraform
Ansible
eBPF
IAM
Cybersecurity
CyberArk
Least Privilege
Teleport
Apply
See all jobs
This is one of many
514,690 more open roles from verified company boards, updated every day.