429,021open jobs
14,503companies
63,779added this week
Browse all
Salary
$123k – $234k per year (Estimated)
Location
Remote/Hybrid (Bellevue, United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Designworks Talent is a specialist recruitment agency headquartered in Austin, Texas. The agency places product designers, user experience researchers, brand designers, and creative leaders into permanent and contract roles at technology companies and studios. It works across the United States on design specific searches rather than general technical recruitment, and maintains its own network of vetted creative professionals.

AI Training Infrastructure Engineer

Location: Hybrid | Bellevue, WA Area

Titles: Senior and Staff (multiple roles available)

Build the Training Infrastructure Powering Next-Generation AI Models

About the Opportunity

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads-including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

The Opportunity

This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.

You'll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.

What You'll Do

  • Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.

  • Design and improve systems that increase training reliability, efficiency, and resource utilization.

  • Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.

  • Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.

  • Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.

  • Build tools and automation that improve the developer experience for AI researchers and engineers.

  • Establish best practices for training infrastructure, operational processes, and platform reliability.

  • Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.

What We're Looking For

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.

  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.

  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.

  • Experience integrating training systems with production machine learning pipelines.

  • Strong programming skills and experience working with complex distributed systems.

  • Ability to independently own technically challenging projects in a fast-moving engineering environment.

  • Comfortable operating with high ownership and limited process overhead.

Preferred Qualifications

  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.

  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.

  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.

  • Experience optimizing GPU utilization, training performance, or distributed system reliability.

  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.

Compensation

  • Competitive base pay for Bellevue market

  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays

Location

  • Hybrid role based in the Bellevue, WA area.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

Why Join?

  • Build the infrastructure powering the next generation of AI models and applications.

  • Work directly on distributed training systems, GPU clusters, and large-scale AI platforms.

  • Solve some of the industry's most challenging problems around AI scalability, reliability, and efficiency.

  • Join early enough to influence architecture, tooling, and engineering practices.

  • Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

  • Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
429,021 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
$156k – $234k per year • Remote/Hybrid • Full-Time • 10+ years exp • Irving • Jacksonville
Python
Python
Flask
FastAPI
Databases
Chroma
Milvus
Pinecone
AI/ML
LangChain
Vertex AI
Gemma
Fine-tuning
Prompt Engineering
AI Agents
NeMo Guardrails
Llama
Mistral
Pandas
NumPy
PyTorch
LLM
RAG
Google ADK
Hallucination
Hugging Face
NVIDIA NeMo
LLM Guardrails
Edge AI
Agentic Workflows
DevOps
OpenShift
CI/CD
Docker
Kubernetes
Vector
Apply
$191k – $334k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
Databases
Apache Kafka
AI/ML
Fine-tuning
Quantization
Prompt Engineering
AI Agents
TensorFlow
PyTorch
RAG
Anomaly Detection
Feature Store
LLM Guardrails
Cybersecurity
Zero Trust
Management
ServiceNow
Apply
$120k – $130k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
AI/ML
AI Agents
Apply
$191k – $334k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Python
Java
Databases
ElasticSearch
OpenSearch
AI/ML
Embeddings
AI Agents
RAG
Semantic Search
Semantic Search
Recommender Systems
Mobile
Algolia
DevOps
Platform Engineering
Vector
Management
ServiceNow
Apply
$25k – $59k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Malaysia
Python
Verilog
C++
SystemVerilog
VHDL
AI/ML
AI Agents
DevOps
CI/CD
Apply
$134k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • Bellevue
AI/ML
Fine-tuning
AI Agents
InfiniBand
DevOps
Platform Engineering
HPC
Apply
$169k – $313k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Bellevue
AI/ML
AI Agents
Agentic Workflows
Apply
$129k – $239k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Tampa • Orlando
Design
AutoCAD
Apply
$133k – $248k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin
Design
AutoCAD
Apply
$86k – $164k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Indianapolis
Python
PowerShell
DevOps
VMWare
Azure
Windows Server
Apply
$100k – $130k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bellevue
Apply
$92k – $120k per year • Remote • Full-Time • 7+ years exp • Bellevue
Apply
$44k – $46k per year • In office • Bellevue
Apply
$42k per year • In office • Bellevue
Apply
$77k – $225k per year (Estimated) • In office • Confidential • Internship • PhD • Bellevue
Apply
See all jobs
This is one of many
429,021 more open roles from verified company boards, updated every day.