821,231open jobs
52,883companies
133,999added this week
Browse all
Salary
≈ $155k – $335k per year (Estimated)
Location
In office (San Mateo)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Dec 24, 2024.

Overview
Company
Impact
Profile match
OpenInfer is an innovation lab building one unified inference stack, from kernel to cloud — reimagining inference around the economics of AI demand: cost, reliability, and sovereignty.

Position Overview

We are looking for an experienced AI Acceleration Engineer who can dive deep into large model (eg. transformer) architectures and blocks such as self/cross/multi-attention, and perform research and development of advanced techniques to accelerate these areas. The ideal candidate will have a deep understanding of large model design, AI acceleration techniques, and will integrate these advancements into the PyTorch stack. Familiarity with Python is essential, and experience with CUDA programming is highly desirable.

Key Responsibilities

  • Innovate on AI model components, such as attention blocks, KV-cache strategies, layer streaming, tokenization, layer norms, and more, to improve AI model performance and scalability.
  • Optimize and integrate AI acceleration techniques into the PyTorch stack, enabling efficient use across diverse hardware platforms.
  • Own & drive features end to end to push the limits of large model architecture, ensuring seamless integration with existing frameworks.
  • Benchmark and profile AI models to evaluate performance improvements, ensuring optimal execution on target hardware.
  • Write and maintain clean, efficient code in Python, with a focus on integration with PyTorch.
  • Leverage CUDA for GPU-based acceleration when necessary, optimizing the attention blocks for maximum performance.
  • Work on cross-functional teams to design, implement, and test new features.

Qualifications

  • Extensive experience with large AI model architectures, particularly with attention blocks and transformer models.
  • Proficiency in Python and hands-on experience with the PyTorch framework.
  • Strong understanding of AI acceleration techniques and their application in real-world use cases.
  • Familiarity with CUDA for GPU programming is highly desirable.
  • Demonstrated ability to optimize complex models for performance across different hardware environments.
  • Experience in developing and deploying AI models at scale is a plus.

What You’ll Gain

  • Opportunity to work alongside industry experts in AI optimization, high-performance computing, and hardware acceleration.
  • Hands-on experience with cutting-edge technologies at the intersection of AI and hardware acceleration.
  • Exposure to open-source development and collaboration with a vibrant community.

Benefits We Offer:

At OpenInfer we offer comprehensive benefits, some include:

  • Medical, Dental, and Vision benefits
  • Flexible Paid Time Off, 10 days
  • Parental Leave
  • 401(k) Plan with company matching
  • Snacks and coffee to keep you energized

These benefits are further detailed in OpenInfer policies and are subject to change at any time, consistent with the terms of any applicable compensation or benefits plans.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
821,231 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Mateo
AI Engineer 1 hour ago
$100k – $120k per year • Equity 0.5–0.5% • In office • Full-Time • San Francisco
Python
TypeScript
Databases
PostgreSQL
Supabase
AI/ML
Claude Code
Multimodal AI
Function Calling
AI Agents
LLM
Tool Use
DevOps
Vercel
Apply
$103k – $145k per year • In office • Full-Time • Bachelor's Degree • Alameda
Python
JavaScript
TypeScript
AI/ML
Prompt Engineering
AI Agents
Machine Learning
DevOps
Rest API
Git
GitHub
Apply
$129k – $182k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Alameda
Python
JavaScript
TypeScript
Databases
Databricks
AI/ML
Copilot
Claude Code
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
LLM
OpenAI Codex
Context Engineering
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
Rest API
CI/CD
Git
AWS
Apply
$152k – $215k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Alameda
Python
JavaScript
TypeScript
Databases
Databricks
AI/ML
Copilot
Claude Code
Model Context Protocol
Function Calling
AI Agents
AWS Bedrock
LLM
RAG
OpenAI Codex
LLM Evaluation
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Rest API
CI/CD
AWS
Cybersecurity
Threat Modeling
Apply
$103k – $145k per year • In office • Full-Time • Bachelor's Degree • Alameda
Python
JavaScript
TypeScript
AI/ML
Prompt Engineering
AI Agents
Machine Learning
DevOps
Rest API
Git
GitHub
Apply
$78k – $119k per year • Equity • In office • 5+ years exp • High School Diploma
AI/ML
PyTorch
Ignite
Analytics
Microsoft Excel
Management
Outlook
Apply
$42k – $56k per year • Equity • In office • 5+ years exp
AI/ML
PyTorch
Ignite
Apply
APU Team Lead 1 day ago
$52k – $68k per year • Equity • In office
AI/ML
PyTorch
Ignite
Apply
In office • 5+ years exp • Bachelor's Degree
Python
C++
DevOps
CI/CD
Jenkins
Git
Bitbucket
Linux
Windows
QA
Pytest
Apply
Hybrid • 3+ years exp • Bachelor's Degree
Python
SQL
AI/ML
Machine Learning
Analytics
Microsoft Excel
Master Data Management
Management
Monday.com
Apply
≈ $100k – $226k per year (Estimated) • In office • Full-Time • 1+ year exp
Python
Java
Rust
Kotlin
C++
Swift
Rust
Axum
AI/ML
llama.cpp
LocalAI
Ollama
LLM
Mobile
Android NDK
DevOps
Rest API
GitHub Actions
CI/CD
Docker
Linux
Windows
Apply
≈ $165k – $300k per year (Estimated) • In office • Full-Time • 5+ years exp • San Mateo
C++
AI/ML
CUDA Toolkit
CUDA
Edge AI
Apply
$397k – $456k per year • Equity • Hybrid • 10+ years exp • Bachelor's Degree • San Mateo
AI/ML
Embeddings
Multimodal AI
LLM
RAG
Feature Store
Recommender Systems
Machine Learning
Apply
$295k – $345k per year • Equity • Hybrid • 5+ years exp • Bachelor's Degree • San Mateo
AI/ML
vLLM
CUDA Toolkit
Quantization
AI Agents
SGLang
CUDA
FSDP
KV Cache
Machine Learning
Apply
≈ $85k – $223k per year (Estimated) • In office • Part-Time • PhD • San Mateo
Apply
Sales Engineer 1 day ago
$90k – $110k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Mateo
Python
Management
Confluence
Apply
$48k – $56k per year • In office • Internship • Bachelor's Degree • San Mateo
PHP
PHP
WordPress
Design
Adobe Photoshop
Adobe After Effects
Marketing
HubSpot
YouTube
Apply
See all jobs
This is one of many
821,231 more open roles from verified company boards, updated every day.