Updated Aug 26, 2026

Flash Attention Jobs - Remote & On-site

Flash Attention jobs curated from vetted startups and product companies, updated daily. Many listings are remote-friendly and provide salary transparency so you can compare roles. Browse and apply today.

Open positions
14
Companies
10
Salary range
$275K – $500K
Average salary
$353K
AI/ML
Data Science
Backend
Frontend
Mobile
DevOps
Web3
Games
Hardware
Robotics
Security
QA
Executive
Networking
Product
Design
Analytics
Support
Enterprise Apps
Quantum

Tenstorrent

tenstorrent.com
Tenstorrent is a Canadian semiconductor company founded in 2016 and led by veteran chip architect Jim Keller. It designs artificial intelligence accelerators and RISC-V processors, and sells intellectual property licences as well as hardware, with an open-source software stack. Investors and licensees include major automotive and technology groups seeking an alternative to closed accelerator ecosystems.
tenstorrent.com • HQ: Toronto, Canada • Embedded Systems • Hardware • Machine Learning • Est. 2016
HQ: Toronto, Canada • Embedded Systems • Hardware • Machine Learning • Est. 2016
Top 25% payVerified live · 23 min ago 3 months ago

Machine Learning Research Engineer (LLMs & AI Systems)

$100k – $500k per year • Middle • 4+ years expIn office (Boston, United States) • PhD • Full-Time • Middle
Python
AI/ML
Flash Attention
LLM
PyTorch
Quantization
Edge AI
Apply
Open 338 daysVerified live · 23 min ago 12 months ago

Machine Learning Engineer, AI Models

In office • Full-Time
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
CUDA
CUDA Toolkit
Flash Attention
PyTorch
Quantization
TensorFlow
Edge AI
Apply
Report

Anthropic

anthropic.com
Anthropic is an American artificial intelligence safety and research company founded in 2021 by former OpenAI researchers, among them the siblings Dario and Daniela Amodei. It develops the Claude family of large language models and ships them through a consumer assistant, an enterprise developer platform and the Claude Code agentic coding tool, alongside open standards such as the Model Context Protocol. Incorporated as a public benefit corporation and headquartered in San Francisco, the company concentrates on interpretability, alignment and reliability research and counts Google and Amazon among its largest investors.
anthropic.com • HQ: San Francisco, United States • Code Intelligence • AI Agents • Machine Learning • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2021
HQ: San Francisco, United States • Code Intelligence • AI Agents • Machine Learning • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2021
Open 342 daysTop 25% pay 12 months ago

Performance Engineer, GPU

$280k per year • In office (San Francisco, United States) • Visa sponsorship • Bachelor's Degree
Claude
CUDA Toolkit
Flash Attention
JAX
Multimodal AI
PyTorch
Quantization
CUDA
Triton
Anthropic
NCCL
NVLink
Apply
Report

Akrys.ai

akrys.ai
Akrys AI is a healthtech company building a clinical intelligence platform to convert doctor-patient conversations into structured records and clinical notes, offering AI-assisted support for healthcare workflows.
akrys.ai • Bengaluru • Health Care • Digital Health • Medical AI • Est. 2025
Bengaluru • Health Care • Digital Health • Medical AI • Est. 2025
21 day ago

Lead AI Engineer

$37k – $90k per year (Estimated) • Lead • 8+ years expIn office (Bengaluru, India) • Staff
Python
Bash
C++
Rust
SQL
C++
PyTorch C++
TensorFlow C++
Python
Dask
AI/ML
Fine-tuning
Hallucination
Knowledge Distillation
LoRA
QLoRA
RLHF
DPO
PPO
Pre-training
SFT
Speech Recognition
Accelerate
DeepSpeed
Encoder-Decoder
Flash Attention
Gemma
JAX
Kaldi
Llama
LLM
Mistral
MLFlow
NER
NLP
NumPy
ONNX
PEFT
Prompt Engineering
PyTorch
Qwen
RAG
Ray
Scikit-learn
SciPy
Spark
Synthetic Data
TensorFlow
TensorRT
Transformers
Triton Inference Server
TRL
vLLM
Whisper
xFormers
FSDP
Hugging Face
NCCL
NVIDIA NeMo
TorchServe
AI Agents
TGI
Weights & Biases
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
SLURM
Vector
Cybersecurity
GDPR
HIPAA
Apply
Report

Scale AI

scale.com
Scale AI is an American company founded in San Francisco in 2016 that supplies the training data, evaluation and tooling behind large artificial intelligence systems. It began with human-labelled annotation for autonomous driving and computer vision, then expanded into reinforcement learning from human feedback, expert data generation, model evaluation and full-stack deployment platforms for enterprises and governments. In 2025 Meta acquired a large minority stake and hired co-founder Alexandr Wang, after which the company continued under new leadership serving defence, public sector and commercial customers.
scale.com • HQ: San Francisco, United States • Defense AI • Data & Analytics • Data Collection • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2016
HQ: San Francisco, United States • Defense AI • Data & Analytics • Data Collection • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2016
Verified live · 2 hours ago 11 months ago

Tech Lead Manager- MLRE, ML Systems

$242k – $517k per year (Estimated) • Equity • Lead • In office (San Francisco, United States) • Full-Time • Staff
CUDA
CUDA Toolkit
Flash Attention
LLM
PyTorch
RLHF
GRPO
Post-training
PPO
Multimodal AI
Function Calling
Apply
Verified live · 2 hours ago 1 year ago

ML Research Engineer, ML Systems

$232k – $508k per year (Estimated) • Equity • In office (San Francisco, United States) • Full-Time
CUDA Toolkit
Flash Attention
LLM
PyTorch
CUDA
Multimodal AI
RLHF
Post-training
Function Calling
Apply
$171k – $400k per year (Estimated) • Equity • Junior • 1+ year expIn office (San Francisco, United States) • Master's Degree • Full-Time • Junior
CUDA Toolkit
Flash Attention
LLM
PyTorch
RLHF
GRPO
Post-training
PPO
AI Agents
Apply
Report

Modal

modal.com
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.
modal.com • HQ: New York, United States • MLOps • LLM & Generative AI • Machine Learning • Information Technology • Hardware • Cloud Computing • AI Infrastructure • Artificial Intelligence • Est. 2021
HQ: New York, United States • MLOps • LLM & Generative AI • Machine Learning • Information Technology • Hardware • Cloud Computing • AI Infrastructure • Artificial Intelligence • Est. 2021
Top 25% payVerified live · 11 hours ago 2 months ago

Member of Technical Staff - Research, Inference

$150k – $350k per year • Staff+ • In office (New York, San Francisco, United States) • Full-Time • Staff
Flash Attention
LLM
Multimodal AI
Quantization
SGLang
Analytics
Seaborn
Matplotlib
Apply
Report

Tether

tether.to
Tether is a digital asset company founded in 2014 that issues USDT, the largest stablecoin in circulation and the most heavily traded asset in crypto markets. Each token is intended to hold a one-to-one peg with the United States dollar and is backed by a reserve portfolio dominated by short-dated Treasury bills, with regular attestation reports published on the reserve composition. Headquartered in San Salvador, El Salvador, the group has expanded beyond stablecoin issuance into gold-backed tokens, Bitcoin mining, energy, peer-to-peer communications and local artificial intelligence infrastructure.
tether.to • HQ: San Salvador, El Salvador • Crypto Payments • Data Centers • Artificial Intelligence • Blockchain & Crypto • Cryptocurrencies • Stablecoins • 1001-5000 employees • Est. 2014
HQ: San Salvador, El Salvador • Crypto Payments • Data Centers • Artificial Intelligence • Blockchain & Crypto • Cryptocurrencies • Stablecoins • 1001-5000 employees • Est. 2014
$102k – $215k per year (Estimated) • Remote (United Kingdom) • English: B2
Diffusion Models
Flash Attention
Multimodal AI
NLP
Quantization
Tokenization
Transformers
Mobile
Metal
Web3
Bitcoin
Apply
Report

Speechmatics

speechmatics.com
Speechmatics is a British company founded in 2006 that builds speech recognition designed to work across accents and dialects. It trains largely with self-supervised learning on unlabelled audio, which improved accuracy for speakers that mainstream systems transcribe poorly. Its engine is licensed by broadcasters, contact centres and software vendors and runs in the cloud or on-premise.
speechmatics.com • HQ: Cambridge, United Kingdom • Conversational AI • Speech & Audio AI • Artificial Intelligence • Est. 2006
HQ: Cambridge, United Kingdom • Conversational AI • Speech & Audio AI • Artificial Intelligence • Est. 2006
Open 460 daysVerified live · 2 days ago 1 year ago

Principal Machine Learning Engineer

$109k – $226k per year (Estimated) • Lead • Cambridge • Remote (United Kingdom) • Internship • Principal
Python
AI/ML
Flash Attention
PyTorch
Self-Supervised Learning
Speech Recognition
TensorFlow
Text-to-Speech
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
GDPR
Apply
Open 460 daysVerified live · 2 days ago 1 year ago

Principal Machine Learning Engineer

$109k – $226k per year (Estimated) • Lead • London • Remote (United Kingdom) • Internship • Principal
Python
AI/ML
Flash Attention
PyTorch
Self-Supervised Learning
Speech Recognition
TensorFlow
Text-to-Speech
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
GDPR
Apply
Report

Build

build.io
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.
build.io • HQ: London, United Kingdom • Web Development • Software • Data Centers • AI Agents • Artificial Intelligence • DevOps • Information Technology • Cloud Computing • Est. 2023
HQ: London, United Kingdom • Web Development • Software • Data Centers • AI Agents • Artificial Intelligence • DevOps • Information Technology • Cloud Computing • Est. 2023
Open 163 daysVerified live · 1 min ago 6 months ago

Member of Technical Staff - Inference Optimization

$62k – $136k per year (Estimated) • Staff+ • Remote/Hybrid (Yokohama, Japan) • Full-Time • Staff
C++
C++
LLVM
PyTorch C++
AI/ML
AWQ
CUDA
CUDA Toolkit
Flash Attention
GPTQ
LLM
OpenCL
OpenMP
PyTorch
Quantization
Triton
ROCm
Apply
Report

Tahoe Therapeutics

tahoebio.ai
Tahoe Therapeutics generates very large single-cell perturbation datasets to train foundation models of how cells respond to drugs. Founded in 2023 in San Francisco, it released the Tahoe-100M dataset openly to the research community. Its models aim to predict drug effects across cancer cell contexts before laboratory testing.
tahoebio.ai • HQ: San Francisco, United States • Artificial Intelligence • Biotechnology • Health Care • Est. 2023
HQ: San Francisco, United States • Artificial Intelligence • Biotechnology • Health Care • Est. 2023
Open 286 daysTop 25% pay 10 months ago

Senior Machine Learning Scientist

$200k – $275k per year • Senior • In office (South San Francisco, United States) • PhD • Senior
JAX
Multimodal AI
NumPy
Pandas
PyTorch
SciPy
TensorFlow
Contrastive Learning
Flash Attention
Self-Supervised Learning
FSDP
Apply
Report

Baseten

baseten.co
Baseten is an AI infrastructure platform designed to help developers and machine learning teams deploy, serve, and scale open-source and custom AI models. The platform provides performant, low-latency inference infrastructure alongside developer tools like Truss, an open-source model packaging framework. Headquartered in San Francisco, California, Baseten enables companies to run state-of-the-art models in production seamlessly without managing underlying cloud infrastructure.
baseten.co • HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
Open 409 daysTop 25% pay 1 year ago

Software Engineer - GPU Kernels

$180k – $360k per year • Remote/Hybrid (San Francisco, New York, United States, Toronto, Montreal, Canada) • Full-Time
C++
AI/ML
CUDA Toolkit
Cursor
Embeddings
Quantization
CUDA
Mixture of Experts
Flash Attention
Transformers
Triton
DevOps
AWS
Management
Notion
Apply
Report

Salary range

Seniority Jobs 25% Median 75%
All levels 8 $99K $168K $286K

Few employers in this selection publish pay, so these figures are Alion estimates built from comparable listings by role, level and country. Gross annual, converted to USD.

Fully remote 21%
3
Hybrid 14%
2
On site 64%
9
Visa sponsorship 7%
1
Equity 21%
3

The market right now

Alion currently lists 14 open Flash Attention jobs from 10 companies. 2 of them were posted or refreshed in the last seven days. 5 of the employers have posted something in the last three months, which is the pool worth watching if you are starting a search now. Every listing links straight to the employer, so you apply on their own board rather than through an intermediary.

What these roles pay

Most employers in this selection do not publish a figure, so the range below is Alion's own estimate, derived from 8 comparable listings priced by role, level and country. Across them the middle half of the market sits between $99K and $286K a year, with a median of $168K. Figures are gross annual amounts converted to US dollars, so roles in different currencies stay comparable.

Where the work is

21% of the listings are fully remote and a further 14% are hybrid, so a large part of this market is open to you regardless of where you live.

Who is hiring

The employers with the most open Flash Attention jobs right now are Scale AI (3), Tenstorrent (2), Speechmatics (2), Anthropic (1), Akrys.ai (1), and Modal (1). Each company page on Alion carries its size, funding stage, tech stack and every other position it has open, so you can judge the employer before you spend an evening on the application.

What employers ask for

Reading across the current listings, the tools that come up most often are Flash Attention (14), PyTorch (11), LLM (7), Quantization (7), CUDA Toolkit (7), CUDA (6), Multimodal AI (6), and TensorFlow (5). The counts are how many of these openings name each one, which is a better guide to what is actually being hired for than a generic skills list.

Relocation, visas and equity

Of the current Flash Attention jobs, 1 states visa sponsorship and 3 include equity. These are filters on the list above, so you can narrow it to the ones that make a move possible for you.

Flash Attention Jobs - Remote & On-site
Frequently asked questions
How many Flash Attention jobs are open right now?
There are 14 open Flash Attention jobs on Alion from 10 companies. The list was last refreshed on August 26, 2026.
What do Flash Attention jobs pay?
The median is $168K a year, and the middle half of the market falls between $99K and $286K. Few of these employers publish a figure, so this is Alion's estimate based on 8 comparable listings.
Are any of these roles remote?
3 of the 14 listings (21%) are fully remote, and you can filter the list down to them in one click.
Which companies are hiring?
The most active employers right now are Scale AI, Tenstorrent, Speechmatics, Anthropic, Akrys.ai, and Modal. Each has its own page on Alion with the rest of its open roles.
Can I get visa sponsorship or relocation?
Yes - of the current listings, 1 state visa sponsorship. Both are filters on the jobs page.
How do I apply?
Open any listing and apply on the employer's own board through the link on the page. No account is required to browse or to apply.