428,564open jobs
14,652companies
61,154added this week
Browse all
Salary
$131k – $266k per year (Estimated)
Location
Remote/Hybrid (Bellevue, United States)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Designworks Talent is a specialist recruitment agency headquartered in Austin, Texas. The agency places product designers, user experience researchers, brand designers, and creative leaders into permanent and contract roles at technology companies and studios. It works across the United States on design specific searches rather than general technical recruitment, and maintains its own network of vetted creative professionals.

Data Center Infrastructure Software Engineer

Location: Hybrid | Bellevue, WA Area

Titles: Senior | Staff | Principal (multiple roles available)

Build the Data Center Software Infrastructure

About the Opportunity

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking Data Center Software Engineers to lead the design, development, configuration, and automation of AI infrastructure clusters.

The Opportunity

Your responsibility begins once servers and racks are installed in the data center and extends through software deployment, networking, configuration, cluster bring-up, and automation, ensuring the platform is fully operational and ready for customer workloads.

What You'll Do

  • Develop infrastructure-as-code, automation, and provisioning systems for compute, networking, and storage.

  • Deploy and optimize Kubernetes, container, and distributed computing platforms.

  • Optimize GPU, networking, storage, and system performance for large-scale AI workloads.

  • Troubleshoot complex issues across hardware, operating systems, networking, storage, and software stacks.

  • Build reliability, observability, and operational excellence practices for mission-critical infrastructure.

What We're Looking For

  • 5+ years of experience designing, building, or operating large-scale Linux-based infrastructure.

  • Hands-on experience with Kubernetes, containerization, and distributed systems in production environments.

  • Experience with infrastructure-as-code and automation tools such as Terraform, Ansible, or similar frameworks.

  • Strong experience operating cloud or datacenter-scale infrastructure.

Preferred Qualifications

  • Experience with bare-metal provisioning and hardware lifecycle management platforms (e.g., MAAS, Ironic, xCAT, xCAT, or similar).

  • Experience with IPMI, Redfish, PXE boot, and automated operating system deployment at scale.

  • Experience managing GPU clusters in datacenter or cloud environments.

Compensation

  • Competitive base pay for Bellevue market

  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Location

  • Hybrid role based in the Bellevue, WA area.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

Why Join?

  • Ground-floor opportunity: you'll be among the earliest engineers on the team, directly shaping architecture, tooling, and culture.

  • Work directly on cutting-edge AI infrastructure at real scale - from data center design through GPU clusters to production inference.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
428,564 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
$156k – $234k per year • Remote/Hybrid • Full-Time • 10+ years exp • Irving • Jacksonville
Python
Python
Flask
FastAPI
Databases
Chroma
Milvus
Pinecone
AI/ML
LangChain
Vertex AI
Gemma
Fine-tuning
Prompt Engineering
AI Agents
NeMo Guardrails
Llama
Mistral
Pandas
NumPy
PyTorch
LLM
RAG
Google ADK
Hallucination
Hugging Face
NVIDIA NeMo
LLM Guardrails
Edge AI
Agentic Workflows
DevOps
OpenShift
CI/CD
Docker
Kubernetes
Vector
Apply
$25k – $73k per year (Estimated) • Remote • Full-Time • 4+ years exp • High School Diploma • Brazil
Python
PowerShell
DevOps
Terraform
Azure DevOps
GitHub Actions
Kibana
Datadog
FluxCD
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Amazon EKS
Azure AKS
GitHub
Cybersecurity
OWASP ZAP
SonarQube
Mend
Apply
$191k – $334k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
Databases
Apache Kafka
AI/ML
Fine-tuning
Quantization
Prompt Engineering
AI Agents
TensorFlow
PyTorch
RAG
Anomaly Detection
Feature Store
LLM Guardrails
Cybersecurity
Zero Trust
Management
ServiceNow
Apply
$120k – $130k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
AI/ML
AI Agents
Apply
$193k – $462k per year (Estimated) • Remote/Hybrid • 10+ years exp • Singapore
DevOps
OpenShift
VMWare
AWS
Kubernetes
Platform Engineering
Amazon EKS
Marketing
LinkedIn
Apply
$116k – $221k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bellevue
AI/ML
Fine-tuning
AI Agents
Edge AI
DevOps
Terraform
Ansible
Kubernetes
Apply
$95k – $187k per year (Estimated) • Remote/Hybrid • Full-Time • Bellevue
AI/ML
Fine-tuning
AI Agents
DevOps
SLURM
Kubernetes
Platform Engineering
HPC
Apply
$134k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • Bellevue
AI/ML
Fine-tuning
AI Agents
InfiniBand
DevOps
Platform Engineering
HPC
Apply
$123k – $234k per year (Estimated) • Remote/Hybrid • Full-Time • Bellevue
AI/ML
DeepSpeed
Fine-tuning
RLHF
Reinforcement Learning
AI Agents
PyTorch
Ray
SFT
Post-training
Megatron-LM
DevOps
Kubernetes
Apply
$169k – $313k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Bellevue
AI/ML
AI Agents
Agentic Workflows
Apply
Equity • In office • 3+ years exp • Bellevue
Python
Cybersecurity
Threat Modeling
Apply
$128k – $192k per year • In office • Full-Time • 4+ years exp • Master's Degree • Bellevue
DevOps
GitHub
Apply
$29k – $69k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bellevue
Apply
$120k – $285k per year (Estimated) • Equity • In office • Bellevue
Python
Go
AI/ML
AI Agents
DevOps
GCP
AWS
Kubernetes
Apply
$105k – $140k per year • In office • 3+ years exp • PhD • Bellevue
Go
AI/ML
Text-to-Speech
Voice Agents
DevOps
Vercel
Management
Slack
Google Docs
Stripe
Marketing
LinkedIn
Apply
See all jobs
This is one of many
428,564 more open roles from verified company boards, updated every day.