Tenstorrent is a Canadian semiconductor company founded in 2016 and led by veteran chip architect Jim Keller. It designs artificial intelligence accelerators and RISC-V processors, and sells intellectual property licences as well as hardware, with an open-source software stack. Investors and licensees include major automotive and technology groups seeking an alternative to closed accelerator ecosystems.
HQ: Toronto, Canada • Embedded Systems • Hardware • Machine Learning • Est. 2016
$100k – $500k per year • Middle • 4+ years exp • In office (Boston, United States) • PhD • Full-Time • Middle
Python
AI/ML
Flash Attention
LLM
PyTorch
Quantization
Edge AI
In office • Full-Time
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
CUDA
CUDA Toolkit
Flash Attention
PyTorch
Quantization
TensorFlow
Edge AI
Anthropic is an American artificial intelligence safety and research company founded in 2021 by former OpenAI researchers, among them the siblings Dario and Daniela Amodei. It develops the Claude family of large language models and ships them through a consumer assistant, an enterprise developer platform and the Claude Code agentic coding tool, alongside open standards such as the Model Context Protocol. Incorporated as a public benefit corporation and headquartered in San Francisco, the company concentrates on interpretability, alignment and reliability research and counts Google and Amazon among its largest investors.
HQ: San Francisco, United States • Code Intelligence • AI Agents • Machine Learning • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2021
$280k per year • In office (San Francisco, United States) • Visa sponsorship • Bachelor's Degree
Claude
CUDA Toolkit
Flash Attention
JAX
Multimodal AI
PyTorch
Quantization
CUDA
Triton
Anthropic
NCCL
NVLink
Akrys AI is a healthtech company building a clinical intelligence platform to convert doctor-patient conversations into structured records and clinical notes, offering AI-assisted support for healthcare workflows.
Bengaluru • Health Care • Digital Health • Medical AI • Est. 2025
≈ $37k – $90k per year (Estimated) • Lead • 8+ years exp • In office (Bengaluru, India) • Staff
Python
Bash
C++
Rust
SQL
C++
PyTorch C++
TensorFlow C++
Python
Dask
AI/ML
Fine-tuning
Hallucination
Knowledge Distillation
LoRA
QLoRA
RLHF
DPO
PPO
Pre-training
SFT
Speech Recognition
Accelerate
DeepSpeed
Encoder-Decoder
Flash Attention
Gemma
JAX
Kaldi
Llama
LLM
Mistral
MLFlow
NER
NLP
NumPy
ONNX
PEFT
Prompt Engineering
PyTorch
Qwen
RAG
Ray
Scikit-learn
SciPy
Spark
Synthetic Data
TensorFlow
TensorRT
Transformers
Triton Inference Server
TRL
vLLM
Whisper
xFormers
FSDP
Hugging Face
NCCL
NVIDIA NeMo
TorchServe
AI Agents
TGI
Weights & Biases
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
SLURM
Vector
Cybersecurity
GDPR
HIPAA
Scale AI is an American company founded in San Francisco in 2016 that supplies the training data, evaluation and tooling behind large artificial intelligence systems. It began with human-labelled annotation for autonomous driving and computer vision, then expanded into reinforcement learning from human feedback, expert data generation, model evaluation and full-stack deployment platforms for enterprises and governments. In 2025 Meta acquired a large minority stake and hired co-founder Alexandr Wang, after which the company continued under new leadership serving defence, public sector and commercial customers.
HQ: San Francisco, United States • Defense AI • Data & Analytics • Data Collection • LLM & Generative AI • Artificial Intelligence • 1001-5000 employees • Est. 2016
≈ $242k – $517k per year (Estimated) • Equity • Lead • In office (San Francisco, United States) • Full-Time • Staff
CUDA
CUDA Toolkit
Flash Attention
LLM
PyTorch
RLHF
GRPO
Post-training
PPO
Multimodal AI
Function Calling
≈ $232k – $508k per year (Estimated) • Equity • In office (San Francisco, United States) • Full-Time
CUDA Toolkit
Flash Attention
LLM
PyTorch
CUDA
Multimodal AI
RLHF
Post-training
Function Calling
≈ $171k – $400k per year (Estimated) • Equity • Junior • 1+ year exp • In office (San Francisco, United States) • Master's Degree • Full-Time • Junior
CUDA Toolkit
Flash Attention
LLM
PyTorch
RLHF
GRPO
Post-training
PPO
AI Agents
Modal is an AI infrastructure company headquartered in New York City and founded in 2021. It provides a serverless cloud platform featuring sub-second cold starts and instant autoscaling that enables developers to run GPU-accelerated workloads, including model inference and fine-tuning, using a Python-native SDK. The company operates a globally distributed compute network designed for AI applications and serves a diverse range of industries such as generative AI and biotechnology.
HQ: New York, United States • MLOps • LLM & Generative AI • Machine Learning • Information Technology • Hardware • Cloud Computing • AI Infrastructure • Artificial Intelligence • Est. 2021
$150k – $350k per year • Staff+ • In office (New York, San Francisco, United States) • Full-Time • Staff
Flash Attention
LLM
Multimodal AI
Quantization
SGLang
Analytics
Seaborn
Matplotlib
Tether is a digital asset company founded in 2014 that issues USDT, the largest stablecoin in circulation and the most heavily traded asset in crypto markets. Each token is intended to hold a one-to-one peg with the United States dollar and is backed by a reserve portfolio dominated by short-dated Treasury bills, with regular attestation reports published on the reserve composition. Headquartered in San Salvador, El Salvador, the group has expanded beyond stablecoin issuance into gold-backed tokens, Bitcoin mining, energy, peer-to-peer communications and local artificial intelligence infrastructure.
HQ: San Salvador, El Salvador • Crypto Payments • Data Centers • Artificial Intelligence • Blockchain & Crypto • Cryptocurrencies • Stablecoins • 1001-5000 employees • Est. 2014
≈ $102k – $215k per year (Estimated) • Remote (United Kingdom) • English: B2
Diffusion Models
Flash Attention
Multimodal AI
NLP
Quantization
Tokenization
Transformers
Mobile
Metal
Web3
Bitcoin
Speechmatics is a British company founded in 2006 that builds speech recognition designed to work across accents and dialects. It trains largely with self-supervised learning on unlabelled audio, which improved accuracy for speakers that mainstream systems transcribe poorly. Its engine is licensed by broadcasters, contact centres and software vendors and runs in the cloud or on-premise.
HQ: Cambridge, United Kingdom • Conversational AI • Speech & Audio AI • Artificial Intelligence • Est. 2006
≈ $109k – $226k per year (Estimated) • Lead • Cambridge • Remote (United Kingdom) • Internship • Principal
Python
AI/ML
Flash Attention
PyTorch
Self-Supervised Learning
Speech Recognition
TensorFlow
Text-to-Speech
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
GDPR
≈ $109k – $226k per year (Estimated) • Lead • London • Remote (United Kingdom) • Internship • Principal
Python
AI/ML
Flash Attention
PyTorch
Self-Supervised Learning
Speech Recognition
TensorFlow
Text-to-Speech
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
GDPR
Build is a cloud infrastructure and platform-as-a-service provider headquartered in London, United Kingdom, and founded in 2023. The company provides a full-stack platform for product teams to deploy and run production applications on its own bare-metal hardware rather than relying on rented hyperscaler capacity. It integrates AI-powered workflows for automated code deployment and infrastructure management, serving a global client base through data centers in the United States, Europe, and Japan.
HQ: London, United Kingdom • Web Development • Software • Data Centers • AI Agents • Artificial Intelligence • DevOps • Information Technology • Cloud Computing • Est. 2023
≈ $62k – $136k per year (Estimated) • Staff+ • Remote/Hybrid (Yokohama, Japan) • Full-Time • Staff
C++
C++
LLVM
PyTorch C++
AI/ML
AWQ
CUDA
CUDA Toolkit
Flash Attention
GPTQ
LLM
OpenCL
OpenMP
PyTorch
Quantization
Triton
ROCm
Tahoe Therapeutics generates very large single-cell perturbation datasets to train foundation models of how cells respond to drugs. Founded in 2023 in San Francisco, it released the Tahoe-100M dataset openly to the research community. Its models aim to predict drug effects across cancer cell contexts before laboratory testing.
HQ: San Francisco, United States • Artificial Intelligence • Biotechnology • Health Care • Est. 2023
$200k – $275k per year • Senior • In office (South San Francisco, United States) • PhD • Senior
JAX
Multimodal AI
NumPy
Pandas
PyTorch
SciPy
TensorFlow
Contrastive Learning
Flash Attention
Self-Supervised Learning
FSDP
Baseten is an AI infrastructure platform designed to help developers and machine learning teams deploy, serve, and scale open-source and custom AI models. The platform provides performant, low-latency inference infrastructure alongside developer tools like Truss, an open-source model packaging framework. Headquartered in San Francisco, California, Baseten enables companies to run state-of-the-art models in production seamlessly without managing underlying cloud infrastructure.
HQ: San Francisco, United States • Hardware • Machine Learning • AI Infrastructure • Artificial Intelligence
$180k – $360k per year • Remote/Hybrid (San Francisco, New York, United States, Toronto, Montreal, Canada) • Full-Time
C++
AI/ML
CUDA Toolkit
Cursor
Embeddings
Quantization
CUDA
Mixture of Experts
Flash Attention
Transformers
Triton
DevOps
AWS
Management
Notion
Salary range
| Seniority |
Jobs |
25% |
Median |
75% |
| All levels |
8 |
$99K |
$168K |
$286K |
Few employers in this selection publish pay, so these figures are Alion estimates built from comparable listings by role, level and country. Gross annual, converted to USD.
The market right now
Alion currently lists 14 open Flash Attention jobs from 10 companies. 2 of them were posted or refreshed in the last seven days. 5 of the employers have posted something in the last three months, which is the pool worth watching if you are starting a search now. Every listing links straight to the employer, so you apply on their own board rather than through an intermediary.
What these roles pay
Most employers in this selection do not publish a figure, so the range below is Alion's own estimate, derived from 8 comparable listings priced by role, level and country. Across them the middle half of the market sits between $99K and $286K a year, with a median of $168K. Figures are gross annual amounts converted to US dollars, so roles in different currencies stay comparable.
Where the work is
21% of the listings are fully remote and a further 14% are hybrid, so a large part of this market is open to you regardless of where you live.
Who is hiring
The employers with the most open Flash Attention jobs right now are Scale AI (3), Tenstorrent (2), Speechmatics (2), Anthropic (1), Akrys.ai (1), and Modal (1). Each company page on Alion carries its size, funding stage, tech stack and every other position it has open, so you can judge the employer before you spend an evening on the application.
What employers ask for
Reading across the current listings, the tools that come up most often are Flash Attention (14), PyTorch (11), LLM (7), Quantization (7), CUDA Toolkit (7), CUDA (6), Multimodal AI (6), and TensorFlow (5). The counts are how many of these openings name each one, which is a better guide to what is actually being hired for than a generic skills list.
Relocation, visas and equity
Of the current Flash Attention jobs, 1 states visa sponsorship and 3 include equity. These are filters on the list above, so you can narrow it to the ones that make a move possible for you.