655,248open jobs
38,129companies
91,967added this week
Browse all
Salary
$127k – $274k per year (Estimated)
Location
Remote/Hybrid (Bellevue, United States)
Seniority
Architect · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Designworks Talent is a specialist recruitment agency headquartered in Austin, Texas. The agency places product designers, user experience researchers, brand designers, and creative leaders into permanent and contract roles at technology companies and studios. It works across the United States on design specific searches rather than general technical recruitment, and maintains its own network of vetted creative professionals.

Vice President AI Infrastructure Engineering

Bellevue, WA Area | Hybrid | Senior Leadership

A rapidly growing, well-funded technology company is seeking a Vice President AI Infrastructure Engineering Leader to lead the engineering organization responsible for transforming newly deployed data center hardware into reliable, production-ready compute infrastructure.

This is a high-impact leadership opportunity for someone who combines deep technical expertise in AI/HPC infrastructure with strong engineering leadership. The ideal candidate understands both the physical infrastructure layer and the software automation required to operate large-scale GPU and compute environments efficiently.

You’ll operate at the intersection of servers, GPUs, Linux, networking, Kubernetes, distributed systems, automation, and infrastructure software, helping establish the architecture, standards, tooling, and engineering practices required to deploy and operate infrastructure at scale.

What You'll Do

  • Lead the team of 15+ responsible for data center infrastructure bring-up and production readiness.

  • Own the platform lifecycle from installed hardware through automated provisioning, configuration, validation, and workload readiness.

  • Build and scale automation for bare-metal provisioning, Linux deployment, configuration management, and infrastructure validation.

  • Lead deployment and configuration of GPU clusters, Kubernetes environments, and distributed compute infrastructure.

  • Establish engineering standards for servers, GPUs, networking, storage, firmware, and system configuration.

  • Drive infrastructure automation using Terraform, Ansible, Bash, and similar technologies.

  • Oversee integration with technologies such as Redfish, IPMI, BMCs, PXE, MAAS, Ironic, Foreman, or comparable platforms.

  • Partner closely with Network, Hardware/GPU, Data Center Operations, SRE, and Software Engineering teams to deliver production-ready infrastructure.

  • Establish automated testing, health checks, monitoring, and validation processes to identify infrastructure issues before workloads reach production.

  • Improve deployment speed, reliability, automation, scalability, and operational efficiency.

  • Build and develop a highly capable engineering organization while establishing processes that can scale with the business.

What We're Looking For

  • 12+ years of experience across infrastructure, software, systems, platform engineering, or related technical disciplines.

  • 5+ years of engineering leadership experience, including managing and developing highly technical teams.

  • Proven experience building and operating large-scale data center, cloud, HPC, or AI infrastructure.

  • Strong technical understanding of Linux, distributed systems, networking, and infrastructure automation.

  • Hands-on understanding of Kubernetes, containers, and infrastructure-as-code.

  • Demonstrated ability to lead complex infrastructure deployments and bring new environments into production.

  • Ability to operate comfortably across both hardware and software organizations.

  • Strong communication and cross-functional leadership skills.

  • A hands-on, high-ownership leadership style with the ability to operate effectively in a fast-moving, build-from-the-ground-up environment.

Preferred Experience

Experience in one or more of the following areas is highly valued:

  • GPU infrastructure, NVIDIA platforms, AI or HPC environments.

  • Bare-metal provisioning technologies such as MAAS, Ironic, xCAT, Foreman, or similar.

  • Hardware management technologies including Redfish, IPMI, BMC, PXE, and firmware management.

  • NVIDIA technologies such as CUDA, NVML, DCGM, NVIDIA drivers, or GPU Operator.

  • High-performance networking including InfiniBand, RoCE, RDMA, or high-speed Ethernet.

  • Cluster orchestration and scheduling technologies such as Kubernetes, Slurm, or similar.

  • Automated infrastructure validation and hardware health testing.

  • Experience scaling infrastructure across thousands of servers or GPUs.

The Opportunity

This is an opportunity to join an organization at an early and highly consequential stage of its growth. You’ll have significant influence over architecture, automation, engineering standards, tooling, and team development, rather than simply inheriting an established infrastructure environment.

The company operates with a startup mentality of fast, lean, highly collaborative, and high ownership while having the resources to build infrastructure for significant scale.

The role is particularly well suited to a leader who enjoys building something new, moving quickly, solving complex infrastructure challenges, and creating software-driven systems that replace manual processes with scalable automation.

Location

  • Bellevue, WA area

  • Hybrid work model with three days per week in the office

  • Candidates currently outside the area may be considered if they are willing to relocate

  • U.S. work authorization required; visa sponsorship is not currently available

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
655,248 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
$21k – $56k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
Bash
AI/ML
AI Agents
DevOps
Terraform
Ansible
Red Hat
Zabbix
OpenTelemetry
Datadog
Prometheus
AWS
Docker
Kubernetes
Grafana
Configuration Management
AWX
Amazon EC2
CentOS Stream
IAM
Apply
$31k – $68k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Bash
Databases
MySQL
PostgreSQL
DevOps
Terraform
Ansible
Azure
AWS
Docker
Kubernetes
Apply
$94k – $211k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Sydney • Melbourne
Python
Go
JavaScript
Java
Java
Maven
Frontend
npm
DevOps
Ansible
Red Hat
OpenShift
CI/CD
Kubernetes
Platform Engineering
Cybersecurity
SBOM
SLSA
Apply
$56k – $126k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Tokyo • Sydney
Python
Go
JavaScript
Java
Java
Maven
Frontend
npm
DevOps
Ansible
Red Hat
OpenShift
CI/CD
Kubernetes
Platform Engineering
Cybersecurity
SBOM
SLSA
Apply
$18k – $41k per year (Estimated) • In office • Full-Time • 3+ years exp • Moscow
DevOps
Terraform
Ansible
Zabbix
OpenShift
Prometheus
VictoriaMetrics
CI/CD
Jenkins
Kubernetes
Grafana
GitLab
Management
Telegram
Apply
$175k – $325k per year (Estimated) • Remote/Hybrid • Full-Time • 15+ years exp • Bachelor's Degree • Bellevue
AI/ML
Agentic Workflows
Apply
$132k – $244k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Tampa • Orlando
Design
AutoCAD
Apply
$137k – $255k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin
Design
AutoCAD
Apply
$87k – $169k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Indianapolis
Python
PowerShell
DevOps
VMWare
Azure
Windows Server
Apply
$122k – $246k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bellevue
AI/ML
Fine-tuning
AI Agents
Edge AI
DevOps
Terraform
Ansible
Kubernetes
Apply
$135k – $297k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Bellevue
Apply
$109k – $164k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Bellevue
Python
Java
SQL
Java
Maven
Databases
PostgreSQL
DevOps
Splunk
Prometheus
CI/CD
Jenkins
Git
Docker
Kubernetes
Management
Agile
Apply
$89k – $117k per year • Remote • Full-Time • 4+ years exp • Bachelor's Degree • Bellevue
SQL
Analytics
Power BI
Apply
$118k – $259k per year (Estimated) • Equity • In office • 5+ years exp • Bellevue
Databases
Apache Kafka
AI/ML
Flink
DevOps
Kubernetes
Platform Engineering
Apply
$98k – $198k per year (Estimated) • In office • 5+ years exp • Bellevue
Apply
See all jobs
This is one of many
655,248 more open roles from verified company boards, updated every day.