368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$210k – $320k per year
Location
Remote/Hybrid (San Mateo, New York, United States)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Fireworks AI is an artificial intelligence infrastructure company headquartered in Redwood City, California, and founded in 2022. The company provides a high-performance inference and training platform that enables developers to deploy, fine-tune, and scale open-source generative models with optimized speed and cost. It operates globally as a cloud-based service provider, catering to technology firms and enterprises seeking to integrate specialized intelligence into their applications through a serverless or dedicated API.

About Us:

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.

The Role:

As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure that powers our large-scale model training operations. Your work will be essential to developing high-performance AI training infrastructure. You'll collaborate with AI researchers and engineers to create robust training pipelines, optimize distributed training workloads, and ensure reliable model development.

Key Responsibilities:

  • Design and implement scalable infrastructure for large-scale model training workloads

  • Develop and maintain distributed training pipelines for LLMs and multimodal models

  • Optimize training performance across multiple GPUs, nodes, and data centers

  • Implement monitoring, logging, and debugging tools for training operations

  • Architect and maintain data storage solutions for large-scale training datasets

  • Automate infrastructure provisioning, scaling, and orchestration for model training

  • Collaborate with researchers to implement and optimize training methodologies

  • Analyze and improve efficiency, scalability, and cost-effectiveness of training systems

  • Troubleshoot complex performance issues in distributed training environments

Minimum Qualifications:

  • Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience

  • 3+ years of experience with distributed systems and ML infrastructure

  • Experience with PyTorch

  • Proficiency in cloud platforms (AWS, GCP, Azure)

  • Experience with containerization, orchestration (Kubernetes, Docker)

  • Knowledge of distributed training techniques (data parallelism, model parallelism, FSDP)

Preferred Qualifications:

  • Master's or PhD in Computer Science or related field

  • Experience training large language models or multimodal AI systems

  • Experience with ML workflow orchestration tools

  • Background in optimizing high-performance distributed computing systems

  • Familiarity with ML DevOps practices

  • Contributions to open-source ML infrastructure or related projects

Why Fireworks?

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.

  • Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.

  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI-no bureaucracy, just results.

  • Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.

Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Mateo
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
$19k – $47k per year (Estimated) • Remote • Full-Time • Perm
C#
C#
ASP.NET Core
Dapper
Entity Framework Core
Databases
Apache Kafka
ClickHouse
ElasticSearch
PostgreSQL
RabbitMQ
Redis
DevOps
CI/CD
Docker
Docker Compose
GitHub Actions
Kubernetes
TeamCity
GitHub
GitLab
Apply
$17k – $43k per year (Estimated) • In office • Full-Time • Tomsk
C#
Java
Python
DevOps
CI/CD
Docker
Git
Grafana
Graylog
Kubernetes
Prometheus
Splunk
Apply
$19k – $47k per year (Estimated) • Remote • Full-Time • Tomsk
C#
C#
ASP.NET Core
Dapper
Entity Framework Core
Databases
Apache Kafka
ClickHouse
ElasticSearch
PostgreSQL
RabbitMQ
Redis
DevOps
CI/CD
Docker
Docker Compose
GitHub Actions
Kubernetes
TeamCity
GitHub
GitLab
Apply
$17k – $43k per year (Estimated) • In office • Full-Time • Perm
C#
Java
Python
DevOps
CI/CD
Docker
Git
Grafana
Graylog
Kubernetes
Prometheus
Splunk
Apply
$132k – $290k per year (Estimated) • In office • Full-Time • 3+ years exp • London
Python
AI/ML
Fine-tuning
Fireworks AI
Multimodal AI
PyTorch
Post-training
AI Agents
Apply
$77k – $214k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • London
Python
AI/ML
Fine-tuning
Fireworks AI
Multimodal AI
PyTorch
Quantization
AI Agents
Apply
GRC Analyst 10 days ago
$160k – $170k per year • In office • Full-Time • 3+ years exp • San Mateo
AI/ML
Fireworks AI
Multimodal AI
PyTorch
ISO 42001
DevOps
AWS
Azure
GCP
Cybersecurity
GDPR
HIPAA
ISO 27001
Least Privilege
NIST CSF
SOC 2
Management
ServiceNow
Apply
$180k – $210k per year • In office • Full-Time • 7+ years exp • San Mateo • New York
Python
AI/ML
Fireworks AI
Multimodal AI
PyTorch
Red Teaming
DevOps
AWS
Azure
GCP
Incident Management
PagerDuty
Splunk
Cybersecurity
Crowdstrike
MITRE ATT&CK
SentinelOne
Sumo Logic
Apply
$170k – $300k per year • In office • Full-Time • 8+ years exp • San Mateo
AI/ML
Fireworks AI
LLM
Multimodal AI
PyTorch
Apply
$235k per year • Remote/Hybrid • 2+ years exp • Master's Degree • San Mateo
Python
SQL
DevOps
Git
Analytics
A/B Testing
Apply
$216k – $388k per year (Estimated) • Equity • In office • Full-Time • San Mateo
AI/ML
AI Agents
Semantic Search
Apply
$192k – $346k per year (Estimated) • Equity • In office • Full-Time • 6+ years exp • San Mateo
Go
JavaScript
Node JS
Python
Ruby
AI/ML
AI Agents
Cybersecurity
FedRAMP
Apply
$150k – $320k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 8+ years exp • San Mateo
Databases
Apache Kafka
AI/ML
Flink
Feature Store
Apply
$222k – $416k per year (Estimated) • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • San Mateo
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.