997,024open jobs
59,481companies
165,643added this week
Browse all
Salary
$200k – $250k per year
Location
In office (Chicago)

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026. DRW scores B on the Alion truth index.

Overview
Company
Impact
Profile match

DRW

DRW is a proprietary trading firm active across futures, options, fixed income and digital assets. It builds its own low-latency technology and quantitative research infrastructure. The firm also invests in real estate, venture and cryptocurrency ventures.

DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk.

Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets.

We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters-it's how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.

About the Role  

We're looking for an AI Inference Platform Engineer to build, operate,and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use. 

You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy.  You'll own the serving platform end-to-end: onboardingnewly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet. 

What You'll Do  

  • Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes. 
  • Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems. 
  • Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies. 
  • Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models. 
  • Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing. 
  • Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads. 
  • Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement. 
  • Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments. 
  • Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements. 
  • Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads. 

What We're Looking For  

The Tech  

  • Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions. 
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang. 
  • Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention. 
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching. 
  • Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes. 
  • Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode. 
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads. 
  • Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana. 
  • Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers. 
  • Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness. 

The Intangibles  

  • You take a measurement-driven approach to performance optimization. 
  • You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries. 
  • You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems. 
  • You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes. 
  • You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance. 
  • You communicate clearly and can explain complex performance tradeoffs across engineering teams. 

The annual base salary range for this position is $200,000 to $250,000 depending on the candidate’s experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus. In addition, DRW offers a comprehensive suite of employee benefits including group medical, pharmacy, dental and vision insurance, 401k (with discretionary employer match), short and long-term disability, life and AD&D insurance, health savings accounts, and flexible spending accounts.

For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at https://drw.com/privacy-notice.

California residents, please review the California Privacy Notice for information about certain legal rights at https://drw.com/california-privacy-notice.

[#LI-VD1]  

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
997,024 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Chicago
$181k – $323k per year • Equity • In office • Full-Time • 4+ years exp • PhD • San Francisco
AI/ML
Reinforcement Learning
Multimodal AI
Computer Vision
VLM
Synthetic Data
World Models
Apply
LLM Platform Engineer 15 days ago
$172k – $329k per year • Equity • In office • Full-Time • 4+ years exp • San Francisco
Python
AI/ML
LoRA
Fine-tuning
Embeddings
Multimodal AI
Knowledge Distillation
Computer Vision
AI Agents
PEFT
Transformers
LLM
RAG
Reranking
Hybrid Search
Text-to-Speech
Model Distillation
DevOps
CI/CD
AWS
Apply
≈ $70k – $162k per year (Estimated) • In office • Full-Time • 1+ year exp • PhD • Prairie View
Python
AI/ML
Scikit-learn
Computer Vision
TensorFlow
Keras
PyTorch
Ray
Machine Learning
Analytics
Seaborn
Matplotlib
Plotly
Apply
$150k – $200k per year • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree • United States
Python
AI/ML
Fine-tuning
JAX
Multimodal AI
Computer Vision
AI Agents
TensorFlow
PyTorch
Machine Learning
DevOps
GCP
Apply
≈ $139k – $263k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • United States
Python
SQL
C++
AI/ML
Machine Learning
Apply
≈ $16k – $43k per year (Estimated) • In office • 5+ years exp • Gurgaon
Python
JavaScript
TypeScript
Node JS
Databases
MySQL
PostgreSQL
Redis
RabbitMQ
Apache Kafka
AI/ML
LLM
RAG
Agentic Workflows
DevOps
Terraform
OpenTelemetry
CloudFormation
Datadog
Prometheus
Pulumi
Azure
CI/CD
AWS
Kubernetes
Grafana
Bicep
IAM
Apply
≈ $59k – $151k per year (Estimated) • In office • Bachelor's Degree • Singapore
DevOps
Istio
OpenTelemetry
Linkerd
PagerDuty
Prometheus
Grafana
Chaos Engineering
Self-Healing
Incident Management
Error Budget
SLI/SLO/SLA
Apply
In office • 1+ year exp • Astana
Python
JavaScript
PHP
TypeScript
Python
SQLAlchemy
FastAPI
Asyncio
Pydantic
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
LangGraph
LangChain
Claude Code
LoRA
Model Context Protocol
vLLM
Fine-tuning
Function Calling
AI Agents
Langfuse
PEFT
CrewAI
LLM
RAG
Google ADK
OpenAI
OpenAI Codex
Frontend
Socket.IO
DevOps
WebSockets
Git
Docker
Linux
QA
Pytest
k6
Locust
Apply
Team Lead DevOps 3 hours ago
≈ $22k – $47k per year (Estimated) • Remote (likely EAEU) • 5+ years exp • Saint Petersburg
Databases
PostgreSQL
Redis
ClickHouse
Apache Kafka
DevOps
Ansible
Loki
VMWare
Prometheus
GitLab CI
CI/CD
Kubernetes
Grafana
Sealed Secrets
SLI/SLO/SLA
Linux
Apply
≈ $40k – $92k per year (Estimated) • In office • 3+ years exp • Moscow
C#
C#
.NET
Databases
PostgreSQL
Apache Kafka
DevOps
Rest API
OpenShift
OpenTelemetry
WebSockets
CI/CD
Kubernetes
Graylog
Management
UML
BPMN
Apply
$200k – $250k per year • In office • Full-Time • Chicago
Python
AI/ML
Prompt Engineering
Human-in-the-Loop
DevOps
Platform Engineering
API Gateway
Management
Agile
Apply
$175k – $225k per year • In office • Bachelor's Degree • Chicago
Python
C++
Apply
$175k – $225k per year • In office • 2+ years exp • Bachelor's Degree • Greenwich
Python
Rust
C++
Zig
DevOps
Linux
Apply
Research Engineer 2 months ago
$175k – $225k per year • In office • 2+ years exp • Bachelor's Degree • New York
Python
Rust
C++
Zig
DevOps
Linux
Apply
≈ $120k – $265k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Montreal
AI/ML
CUDA Toolkit
Quantization
ONNX
TensorRT
OpenCL
TensorFlow
PyTorch
CUDA
Feature Store
Machine Learning
Apply
$78k – $108k per year • Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Chicago
Cybersecurity
HIPAA
Management
Outlook
Apply
≈ $109k – $210k per year (Estimated) • Remote (United States) • Full-Time • 3+ years exp • Bachelor's Degree • Louisville • Charlotte • Tampa • Fort Lauderdale • Washington
Cybersecurity
HIPAA
Apply
≈ $37k – $78k per year (Estimated) • Hybrid • 2+ years exp • Bachelor's Degree • Chicago
Analytics
Microsoft Excel
Management
Microsoft Office
Apply
$117k – $175k per year • Equity • In office • Full-Time • 5+ years exp • Chicago • Victoria • Toronto • Calgary • Montreal
Java
C#
C#
.NET
Apply
Data Scientist 1 day ago
$105k – $124k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Minneapolis • Atlanta • Chicago • Charlotte
Python
SQL
SAS
Databases
Databricks
AI/ML
Machine Learning
DevOps
Azure
Linux
Apply
See all jobs
This is one of many
997,024 more open roles from verified company boards, updated every day.