1,419,246open jobs
82,989companies
210,832added this week
Browse all
Salary
≈ $164k – $363k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Oct 8, 2026. Hedra scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Create professional videos, images, and audio in one studio. Hedra gives you AI-powered tools to go from concept to production in minutes.

About Hedra

Hedra is the platform, models, and infrastructure for visual intelligence.

We build models and systems that push the frontier of visual intelligence, along with the infrastructure required to make those models fast, efficient, reliable, and accessible at scale.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Researchers and engineers at Hedra work closely across boundaries, own problems end to end, and have significant influence over both what we build and how we build it.

The Role

We’re looking for an Inference Optimization Engineer to work alongside our research team on making state-of-the-art visual models fast and efficient at inference time.

You’ll work at the boundary between research and systems, taking new model architectures and figuring out how to run them efficiently on modern hardware. That means understanding where time and memory are being spent, identifying opportunities for algorithmic and systems-level improvements, and implementing optimizations across model architecture, inference algorithms, runtimes, kernels, and distributed execution.

The problems rarely live neatly within one layer of the stack. Depending on what you find, you might modify how a model executes, develop a new inference technique, write a custom GPU kernel, rethink memory movement, or change how work is distributed across accelerators.

We care more about technical depth, curiosity, and demonstrated ability than years of experience. We’re open to experienced ML systems engineers as well as exceptional early-career engineers or researchers who have already gone unusually deep on model performance, GPU systems, or efficient inference.

What You’ll Do

  • Work directly with research scientists and engineers to make new visual models fast and efficient at inference time.

  • Profile model architectures and workloads to understand bottlenecks across compute, memory, communication, and model execution.

  • Develop and implement new approaches to improving inference latency, throughput, memory efficiency, and GPU utilization.

  • Explore algorithmic optimizations including quantization, sparsity, caching, compilation, attention optimizations, and alternative execution strategies.

  • Build or optimize GPU kernels using CUDA, Triton, or similar technologies when existing implementations leave performance on the table.

  • Optimize model execution across single-GPU, multi-GPU, and multi-node environments.

  • Reason about the interaction between model architecture and hardware, and work with researchers when architectural changes can unlock meaningful performance improvements.

  • Investigate communication, memory movement, parallelism, and distributed execution strategies for large visual models.

  • Build rigorous benchmarking, profiling, and performance-regression infrastructure to understand performance and evaluate new optimization ideas.

  • Evaluate new inference runtimes, compilers, frameworks, optimization techniques, and accelerator hardware.

  • Stay close to advances in efficient inference, GPU programming, model architectures, compilers, and ML systems research, and rapidly test promising ideas.

  • Help turn research breakthroughs into models that can be deployed and served efficiently at scale.

What We’re Looking For

  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research, demonstrated through research, production engineering, open-source contributions, or unusually ambitious independent work.

  • Strong understanding of how modern deep learning models execute on hardware, including the relationship between compute, memory, communication, and performance.

  • Experience profiling ML workloads, identifying bottlenecks, forming hypotheses, and driving measurable performance improvements.

  • Strong programming fundamentals in Python, C++, or another systems-oriented language.

  • Experience with some combination of PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.

  • Ability to reason across abstraction layers rather than treating model architecture, framework, runtime, kernel, and hardware boundaries as fixed.

  • Strong intuition for performance tradeoffs across latency, throughput, memory, numerical precision, model quality, and complexity.

  • Curiosity about how models work internally and a willingness to modify or rethink existing approaches when the performance problem calls for it.

  • Comfort working on ambiguous problems where the bottleneck, and sometimes even the right question, is not known in advance.

  • Ability to communicate technical ideas clearly and collaborate closely with research scientists and engineers.

We don’t expect every candidate to have experience across the entire stack. Exceptional depth in one or more relevant areas, combined with the ability and curiosity to reason across the others, matters more to us than checking every box.

Nice to Have

  • Experience optimizing large generative, multimodal, vision, or video models.

  • CUDA, Triton, CUTLASS, or other GPU kernel development.

  • Deep knowledge of GPU architecture, memory hierarchy, and hardware-aware optimization.

  • Experience with attention optimization, kernel fusion, memory-efficient execution, or custom operators.

  • Model compilation or graph optimization experience.

  • Quantization, sparsity, caching, speculative execution, or other efficient inference techniques.

  • Experience optimizing diffusion, autoregressive, transformer, or other large generative architectures.

  • Multi-GPU or multi-node model execution, including tensor, pipeline, sequence, or other forms of parallelism.

  • Experience optimizing communication or data movement between accelerators.

  • Experience with profiling tools such as Nsight Systems or Nsight Compute.

  • Contributions to ML systems, inference runtimes, compilers, GPU libraries, or performance-focused open-source projects.

  • Research or publications in efficient ML, ML systems, GPU computing, compilers, or related areas.

Benefits:

  • Competitive compensation and equity

  • 401k

  • Healthcare (Silver PPO Medical, Vision, Dental)

  • Lunch and snacks at the office

This role is based in San Francisco, and we work together in person five days a week.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,419,246 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$180k – $250k per year • In office • 10+ years exp • Bachelor's Degree • Chicago
DevOps
GCP
Azure
AWS
Platform Engineering
FinOps
Apply
$180k – $250k per year • In office • 10+ years exp • Bachelor's Degree • New York
DevOps
GCP
Azure
AWS
Platform Engineering
FinOps
Apply
AI Lab - Research 7 days ago
$200k per year • In office • New York
Python
AI/ML
CUDA Toolkit
Reinforcement Learning
JAX
AI Agents
Mamba
Transformers
PyTorch
Self-Supervised Learning
Time Series Forecasting
CUDA
Triton
Post-training
Pre-training
Machine Learning
Apply
AI Engineer 1 day ago
$200k per year • In office • 6+ years exp • Bachelor's Degree • Chicago
Python
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
LLM
Tool Use
DevOps
Platform Engineering
Apply
AI Engineer 1 day ago
$200k per year • In office • 6+ years exp • Bachelor's Degree • New York
Python
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
LLM
Tool Use
DevOps
Platform Engineering
Apply
$63k – $120k per year • In office • Secret • Full-Time • Bachelor's Degree • El Segundo
Python
C++
Perl
DevOps
RTOS
CI/CD
Git
GitLab
Linux
Management
Agile
Scrum
Apply
$76k – $144k per year • In office • Secret • Full-Time • Master's Degree • El Segundo
Python
C++
Perl
DevOps
RTOS
CI/CD
Git
GitLab
Linux
Management
Agile
Scrum
Apply
$63k – $120k per year • In office • Secret • Full-Time • Bachelor's Degree • El Segundo
Python
C++
Perl
DevOps
RTOS
CI/CD
Git
GitLab
Linux
Management
Agile
Scrum
Apply
$57k – $109k per year • In office • Secret • Full-Time • Bachelor's Degree • Tucson
Python
C++
Perl
DevOps
RTOS
CI/CD
Git
GitLab
Linux
Management
Agile
Scrum
Apply
≈ $37k – $64k per year (Estimated) • In office • Internship • Bachelor's Degree • Sindelfingen
Python
C++
Robotics
Sensor Fusion
Apply
Research Scientist 6 months ago
$200k – $325k per year • In office • Full-Time • PhD • San Francisco
AI/ML
RLHF
DPO
PPO
World Models
Physical AI
Embodied AI
Machine Learning
Robotics
Sim-to-Real
Apply
Research Engineer 6 months ago
$175k – $275k per year • In office • Full-Time • Bachelor's Degree • San Francisco
AI/ML
DeepSpeed
Fine-tuning
Reinforcement Learning
Multimodal AI
PyTorch
PPO
Post-training
Pre-training
FSDP
Vision-Language-Action
World Models
Physical AI
Embodied AI
Machine Learning
Robotics
Sim-to-Real
Reinforcement Learning
Apply
≈ $121k – $243k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
World Models
Marketing
LinkedIn
Apply
Head of Marketing 3 months ago
≈ $190k – $358k per year (Estimated) • In office • Full-Time • 7+ years exp • PhD • San Francisco
AI/ML
LLM Guardrails
World Models
Apply
Account Executive 3 months ago
≈ $136k – $287k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
AI/ML
World Models
Apply
AI Engineer (US) 1 day ago
$110k – $185k per year • In office • Full-Time • 1+ year exp • San Francisco
Python
AI/ML
LangChain
NLP
PyTorch
LLM
RAG
Hugging Face
Machine Learning
DevOps
Rest API
GCP
Azure
AWS
Apply
$4k – $14k per year • Equity • In office • 2+ years exp • PhD • San Francisco
Python
Go
Java
Kotlin
C++
AI/ML
Machine Learning
DevOps
Kubernetes
Apply
$4k – $14k per year • Equity • In office • 2+ years exp • Bachelor's Degree • San Francisco
Python
Databases
Databricks
AI/ML
Spark
Airflow
Dagster
MLFlow
Transformers
TensorFlow
Pandas
PyTorch
BERT
Machine Learning
DevOps
AWS
Apply
$4k – $14k per year • Equity • In office • 3+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
Cursor
Qwen
DeepSeek
Claude Code
LoRA
Model Context Protocol
vLLM
Fine-tuning
Embeddings
RLHF
Reinforcement Learning
Quantization
AI Agents
SGLang
AWQ
GPTQ
TensorRT
PEFT
TensorRT-LLM
LLM
RAG
OpenAI Codex
DPO
SFT
Post-training
LLM Guardrails
KV Cache
Machine Learning
DevOps
GCP
AWS
Kubernetes
Apply
$115k – $130k per year • Hybrid • San Francisco
Design
AutoCAD
Management
Intercom
Microsoft Office
Apply
See all jobs
This is one of many
1,419,246 more open roles from verified company boards, updated every day.