1,443,753open jobs
85,355companies
221,329added this week
Browse all
Salary
≈ $152k – $336k per year (Estimated)
Location
In office (New York)

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on May 20, 2024. Jane Street scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Jane Street (Jane Street Capital) is a quantitative trading firm and liquidity provider specializing in automated market making, algorithmic execution, and high-frequency trading. Headquartered in New York, New York, the firm operates major regional hubs in London, Hong Kong, Singapore, and Amsterdam. Its core activities involve deploying proprietary quantitative research, machine learning models, and low-latency infrastructure to facilitate continuous trading across exchange-traded funds (ETFs), equities, options, fixed income, futures, and derivative products across more than 200 trading venues globally.

We are looking for an engineer with experience in low-level systems programming and optimization to join our growing ML team. 

Machine learning is a critical pillar of Jane Street's global business. Our ever-evolving trading environment serves as a unique, rapid-feedback platform for ML experimentation, allowing us to incorporate new ideas with relatively little friction.

Your part here is optimizing the performance of our models - both training and inference. We care about efficient large-scale training, low-latency inference in real-time systems, and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking, and host- and GPU-level considerations. Zooming in, we also want to ensure our platform makes sense even at the lowest level - is all that throughput actually goodput? Does loading that vector from the L2 cache really take that long?

If you’ve never thought about a career in finance, you’re in good company. Many of us were in the same position before working here. If you have a curious mind and a passion for solving interesting problems, we have a feeling you’ll fit right in. 

There’s no fixed set of skills, but here are some of the things we’re looking for:

  • An understanding of modern ML techniques and toolsets
  • The experience and systems knowledge required to debug a training run’s performance end to end
  • Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores, and the memory hierarchy
  • Debugging and optimization experience using tools like CUDA GDB, NSight Systems, NSight Compute
  • Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS
  • Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization, and asynchronous memory loads
  • Background in Infiniband, RoCE, GPUDirect, PXN, rail optimization, and NVLink, and how to use these networking technologies to link up GPU clusters
  • An understanding of the collective algorithms supporting distributed GPU training in NCCL or MPI
  • An inventive approach and the willingness to ask hard questions about whether we're taking the right approaches and using the right tools

If you're a recruiting agency and want to partner with us, please reach out to  [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,443,753 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
New York
$195k – $262k per year • Remote (United States) • Palo Alto
Python
AI/ML
DeepSpeed
CUDA Toolkit
RLHF
Function Calling
TRL
Transformers
PyTorch
LLM
Ray
Synthetic Data
CUDA
Triton
DPO
SFT
PPO
GRPO
Post-training
Megatron-LM
FSDP
NCCL
InfiniBand
Agentic Workflows
Tool Use
RLAIF
Reward Modeling
Machine Learning
DevOps
SLURM
Kubernetes
Apply
$195k – $262k per year • Remote (United States) • Palo Alto • San Francisco
Python
AI/ML
Ray Serve
vLLM
Triton Inference Server
Quantization
Knowledge Distillation
Function Calling
AI Agents
VLM
SGLang
AWQ
GPTQ
TensorRT
TensorRT-LLM
PyTorch
LLM
Ray
Mixture of Experts
KServe
TPOT
Triton
Post-training
NCCL
InfiniBand
NVLink
Structured Outputs
Speculative Decoding
KV Cache
Tool Use
Model Distillation
Machine Learning
Apply
$180k – $225k per year • Remote (United States) • United States
AI/ML
vLLM
TensorRT
Apply
$195k – $262k per year • Remote (United States) • PhD • Palo Alto • San Francisco
Python
AI/ML
vLLM
CUDA Toolkit
RLHF
Quantization
Knowledge Distillation
VLM
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Mixture of Experts
Synthetic Data
CUDA
Triton
DPO
SFT
Post-training
Speculative Decoding
KV Cache
Model Distillation
RLAIF
Machine Learning
Apply
≈ $144k – $255k per year (Estimated) • Remote (United States) • 5+ years exp • United States
AI/ML
NCCL
NVLink
DevOps
Linux
Apply
$210k – $275k per year • In office • Full-Time • New York
Apply
$40k – $42k per year • In office • Full-Time • High School Diploma • New York
Apply
Diesel Mechanic 1 day ago
$60k – $80k per year • Equity • In office • Full-Time • 5+ years exp • New York
Apply
$120k – $145k per year • In office • New York
Apply
up to $72k per year • In office • Full-Time • 2+ years exp • New York
Management
Microsoft Office
Apply
See all jobs
This is one of many
1,443,753 more open roles from verified company boards, updated every day.