368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$210k – $320k per year
Location
In office (San Mateo)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Fireworks AI is an artificial intelligence infrastructure company headquartered in Redwood City, California, and founded in 2022. The company provides a high-performance inference and training platform that enables developers to deploy, fine-tune, and scale open-source generative models with optimized speed and cost. It operates globally as a cloud-based service provider, catering to technology firms and enterprises seeking to integrate specialized intelligence into their applications through a serverless or dedicated API.

About Us:

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.

We are seeking a Member of Technical Staff, Evals & Post-Training Product to help define how developers improve models on Fireworks. This role sits at a unique intersection of scalable system design, deep data science, and model quality.

You will build the infrastructure and workflows that connect evaluation and post-training. Our evaluation setup has grown past its original scope, and we need someone who can take it to the next stage, improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to make it easier to author evals, understand results, and iterate quickly.

Key Responsibilities

  • Scale Eval Infrastructure: Take ownership of our internal eval setup and evolve it for the future. You will design systems to eliminate single-host coordination bottlenecks, resolve log-syncing latency, and build seamless programmatic access.

  • Benchmark Obsession & Reproduction: Act as a "metrics obsessive." Track state-of-the-art (SOTA) benchmarks, read the latest research papers, dig deep into data discrepancies, and insist on rigorously reproducing published results internally.

  • Pioneer New Benchmarks: Design and build entirely new benchmarks to measure model performance on complex, emerging, or domain-specific use cases.

  • Own Fine-Tuning Product Experiences: Build and improve user-facing workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.

  • Work Closely With Users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, triage issues, and convert bespoke workflows into productized, reusable solutions.

Minimum Requirements

  • Experience: 1-7 years of software engineering or data science experience (We are hiring at multiple levels for this role).

  • Strong System Design Skills: You know how to architect scalable, programmatic systems and transition legacy setups into robust infrastructure.

  • Sandbox Infrastructure: Hands-on experience building or working with sandbox environments and sandbox infrastructure for secure code execution and testing.

  • Analytical & Data Science Mindset: You possess a deep understanding of LLM evaluations, how to design them, and how to use the results to guide model improvement. You are meticulous about data and metrics.

  • Understanding of the GenAI Lifecycle: You understand the end-to-end workflow-from prompting a base model to curating a dataset, fine-tuning, and productionizing agents.

Preferred Qualifications

  • Experience: 3+ years of software engineering or applied data science experience.

  • Frameworks & Orchestration: Experience working with the Harbor framework or similar container registry and orchestration tools.

  • Public Writing & Analysis: A strong interest in discovering where different models excel and fall short, with a desire to write up and publish these insights publicly (e.g., technical blogs, whitepapers).

  • Inference & Hardware Knowledge: Interest in the hardware side of AI-understanding GPU constraints, inference optimization techniques, and how they relate to model performance.

  • Startup DNA: Experience in fast-paced environments where you own features end-to-end.

Why Fireworks?

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.

  • Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.

  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI-no bureaucracy, just results.

  • Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.

Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Mateo
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$22k – $57k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
Python
TypeScript
JavaScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
Claude
Claude Code
LLM
Model Context Protocol
Frontend
Angular
Next.js
React.js
DevOps
AWS
AWS Lambda
CI/CD
Docker
GitHub Actions
Terraform
Amazon ECS
Amazon S3
GitHub
IAM
Apply
Data & AI Engineer 2 days ago
In office • Full-Time • 3+ years exp • Luxembourg City
Java
Python
TypeScript
C#
Java
Spring Boot
Python
FastAPI
C#
.NET
AI/ML
Computer Vision
LLM
Mistral
NLP
Ollama
ONNX
RAG
Semantic Search
Transformers
Edge AI
OCR
Semantic Search
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GitHub Actions
Rest API
Terraform
Vector
GitHub
Apply
$64k – $150k per year (Estimated) • Equity • Remote • Full-Time • PhD
Bash
Ruby
Ruby
Ruby on Rails
AI/ML
LLM
DevOps
CI/CD
Git
Kubernetes
OpenShift
GitLab
Cybersecurity
FedRAMP
Apply
$132k – $290k per year (Estimated) • In office • Full-Time • 3+ years exp • London
Python
AI/ML
Fine-tuning
Fireworks AI
Multimodal AI
PyTorch
Post-training
AI Agents
Apply
$77k – $214k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • London
Python
AI/ML
Fine-tuning
Fireworks AI
Multimodal AI
PyTorch
Quantization
AI Agents
Apply
GRC Analyst 10 days ago
$160k – $170k per year • In office • Full-Time • 3+ years exp • San Mateo
AI/ML
Fireworks AI
Multimodal AI
PyTorch
ISO 42001
DevOps
AWS
Azure
GCP
Cybersecurity
GDPR
HIPAA
ISO 27001
Least Privilege
NIST CSF
SOC 2
Management
ServiceNow
Apply
$180k – $210k per year • In office • Full-Time • 7+ years exp • San Mateo • New York
Python
AI/ML
Fireworks AI
Multimodal AI
PyTorch
Red Teaming
DevOps
AWS
Azure
GCP
Incident Management
PagerDuty
Splunk
Cybersecurity
Crowdstrike
MITRE ATT&CK
SentinelOne
Sumo Logic
Apply
$170k – $300k per year • In office • Full-Time • 8+ years exp • San Mateo
AI/ML
Fireworks AI
LLM
Multimodal AI
PyTorch
Apply
$235k per year • Remote/Hybrid • 2+ years exp • Master's Degree • San Mateo
Python
SQL
DevOps
Git
Analytics
A/B Testing
Apply
$216k – $388k per year (Estimated) • Equity • In office • Full-Time • San Mateo
AI/ML
AI Agents
Semantic Search
Apply
$192k – $346k per year (Estimated) • Equity • In office • Full-Time • 6+ years exp • San Mateo
Go
JavaScript
Node JS
Python
Ruby
AI/ML
AI Agents
Cybersecurity
FedRAMP
Apply
$150k – $320k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Mateo
Databases
Apache Kafka
AI/ML
Flink
Feature Store
Apply
$222k – $416k per year (Estimated) • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • San Mateo
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.