429,799open jobs
14,642companies
59,404added this week
Browse all
Salary
$236k – $330k per year
Location
In office (Bellevue)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Snowflake is an American cloud data company founded in 2012 by former Oracle database architects and headquartered in Bozeman, Montana. Its platform separates storage from compute so organisations can scale query power independently of data volume, run workloads across Amazon Web Services, Azure and Google Cloud, and share live datasets with partners without copying them. Originally a cloud data warehouse, it has expanded into data engineering, application development with Snowpark and Native Apps, and AI features through Cortex, and it listed on the New York Stock Exchange in 2020 in the largest software offering to that point.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset - who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.

Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.

Our work spans the full inference stack-from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.

Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization-rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development.

Recent innovations from Snowflake AI Research include Arctic Inference, our open-source inference system, and technologies such as Shift Parallelism, which dynamically adapts parallelism to workload characteristics; SwiftKV, which reduces redundant prefill computation; Arctic Speculator and SuffixDecoding for fast speculative decoding; Jacobi Forcing for causal parallel decoding; and Semi-Persistence for fast model swapping and dynamic multi-model serving.

This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI.

Responsibilities

  • Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.

  • Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.

  • Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.

  • Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.

  • Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.

  • Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.

  • Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.

  • Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.

  • Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.

  • Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.

  • Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.

  • Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.

  • Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.

  • 5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.

  • Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.

  • Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.

  • Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.

  • Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.

  • Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.

  • Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.

  • Demonstrated ability to operate as an independent problem identifier and solver -recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.

  • Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.

  • Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.

  • Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
429,799 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
Remote
Databases
Snowflake
AI/ML
OpenAI
Anthropic
DevOps
Datadog
Docker
Cloudflare
Cybersecurity
Censys
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara • Austin • Hillsboro • Boulder • Redmond
AI/ML
CUDA Toolkit
CUDA
DevOps
HPC
Apply
$224k – $357k per year • In office • Full-Time • 12+ years exp • PhD • Santa Clara
Python
C++
AI/ML
CUDA Toolkit
Reinforcement Learning
JAX
CUDA
Game Dev
NVIDIA PhysX
Robotics
Isaac Sim
MuJoCo
Isaac Lab
Reinforcement Learning
Apply
$152k – $296k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Montreal
Python
SQL
Databases
Snowflake
AI/ML
Cursor
Claude
Spark
Claude Code
Anomaly Detection
Time Series Forecasting
Red Teaming
Apply
Risk Modeling Intern 4 hours ago
$36k – $55k per year (Estimated) • Remote/Hybrid • Internship • Raleigh
Python
SAS
Databases
Snowflake
DevOps
AWS
Analytics
Microsoft Excel
Apply
$143k – $188k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • New York
Databases
Snowflake
AI/ML
AI Agents
Web3
Rollup
Marketing
Salesforce
Marketo
Apply
$160k – $230k per year • Remote/Hybrid • Full-Time • 4+ years exp • Menlo Park
Python
Go
JavaScript
Java
TypeScript
Databases
Snowflake
AI/ML
AI Agents
Frontend
React.js
Apply
$85k – $202k per year (Estimated) • In office • Full-Time • Oslo
Python
SQL
Databases
Snowflake
AI/ML
LangChain
AI Agents
Accelerate
PyTorch
LLM
RAG
Hugging Face
Apply
$166k – $239k per year • Remote/Hybrid • Full-Time • 7+ years exp • Menlo Park
Databases
Snowflake
AI/ML
AI Agents
Apply
$236k – $339k per year • Remote/Hybrid • Full-Time • Menlo Park • Bellevue
Databases
Snowflake
AI/ML
AI Agents
Apply
$100k – $130k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bellevue
Apply
$92k – $120k per year • Remote • Full-Time • 7+ years exp • Bellevue
Apply
$42k per year • In office • Bellevue
Apply
$77k – $225k per year (Estimated) • In office • Confidential • Internship • PhD • Bellevue
Apply
$200k – $333k per year (Estimated) • Equity • In office • Bellevue
Python
Go
DevOps
gRPC
Terraform
Istio
Envoy
etcd
Kubernetes
Service Mesh
Apply
See all jobs
This is one of many
429,799 more open roles from verified company boards, updated every day.