368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$87k – $182k per year (Estimated)
Location
In office (Seattle)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Elastix AI delivers scalable, energy-efficient AI inference through machine learning, system software, and reconfigurable hardware.

Location: Seattle, WA (Hybrid - 3 days/week in office)

About ElastixAI:

ElastixAI is an early-stage Software startup on a mission to reinvent AI inference infrastructure from the ground up. We're building a next-generation inference platform that delivers unprecedented efficiency by tightly integrating machine learning, software stack, and custom hardware. Our philosophy is simple: the best performance comes from holistic co-design, where every layer, from model architecture to kernels to silicon, works in harmony.

If you're excited about pushing AI performance to physical limits and shaping the future of large-scale inference, we'd love to meet you.

Role Summary:

We're looking for an Inference Infrastructure Software Engineer to own and evolve the cloud and Kubernetes backbone behind our Token-as-a-Service platform. You'll be the connective tissue between our inference engine and the production environments where customers actually consume tokens - making sure our accelerated workloads run reliably, scale predictably, and deploy seamlessly across managed and self-hosted clusters.

This is a hands-on role with broad surface area. You'll touch everything from cluster bring-up, automating the software releases, and AI Accelerator scheduling to service reliability and cost optimization, working closely with our ML, runtime, and hardware teams to expose the full performance of our co-designed stack to end users.

Key Responsibilities:

  • Build, operate, and evolve ElastixAI's Kubernetes infrastructure powering our Token-as-a-Service capability.

  • Run accelerated inference workloads in production at scale, with strong SLAs around latency, throughput, and availability.

  • Manage and harden our AWS, GCP, and on-prem infrastructure, including networking, storage, IAM, and observability layers tied to our services.

  • Develop tooling and automation in Python, Bash, Rust, and Go to streamline deployments, rollouts, autoscaling, and incident response.

  • Partner with the ML and runtime teams to productionize new inference capabilities, model deployments, and routing strategies.

  • Contribute to capacity planning, cost optimization, and reliability engineering across multi-cloud and self-hosted environments.

  • Help define the platform roadmap as we scale from early customers to broad production deployments.

  • Be a member of the Elastix On-Call rotation

Required Qualifications:

  • Minimum BS in Computer Science, Software Engineering, or a related field.

  • 3-5 years of hands-on Kubernetes experience, including EKS, GKE, and/or self-hosted clusters.

  • 2-3 years of production experience operating workloads on AWS or GCP.

  • Proven track record running ML or inference services at scale on Kubernetes in production.

  • Strong experience running accelerated workloads in Kubernetes, scheduling, drivers, device plugins, MIG, networking, and storage considerations.

  • Solid coding skills in Python, Bash and proficiency in Go

  • Proficient in configuring and leveraging Linux OS in production

  • Experience with infrastructure-as-code (Terraform, Pulumi), OS configuration state (Ansible, Puppet, Salt) and GitOps workflows (Argo CD, Flux).

  • Experience in OS configuration tooling.

  • Familiarity with AI inference and/or training workflows and the operational patterns around them.

  • Pragmatic, ownership-oriented mindset; comfortable operating in early-stage ambiguity and shipping iteratively.

Preferred/Bonus Qualifications:

  • MS/PhD in Computer Science, Software Engineering, or a related field.

  • Experience with inference servers and runtimes (e.g., Triton, vLLM, TGI) and model-serving patterns (batching, streaming, KV-cache aware routing).

  • Exposure to heterogeneous accelerators beyond GPUs (FPGAs, custom ASICs).

  • Background in observability, SRE, or performance engineering for latency-sensitive services.

  • Experience building customer facing API platforms including onboarding, API keys/auth management, and usage metering.

What We Offer:

  • A chance to be a foundational engineer in an innovative AI startup.

  • A dynamic and collaborative work environment and the change to have a significant impact on new technology

  • The opportunity to work on challenging problems at the intersection of ML, software, and systems.

  • Competitive compensation and startup equity package

  • Comprehensive medical, dental, and vision coverage (premiums 100% paid by employer)

  • Flexible Time Off (FTO)

  • Paid parental leave

  • Gym or fitness benefit

  • Commuter benefit

  • Investment in employee learning & development

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seattle
$21k – $35k per year • In office • Moscow
Bash
Python
Databases
MySQL
PostgreSQL
DevOps
Ansible
Debian
Grafana
Proxmox VE
Ubuntu
Zabbix
Apply
DevOps Engineer 3 hours ago
$49k – $88k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Warsaw
Bash
Python
DevOps
AWS
Azure
Azure AKS
CI/CD
Docker
GCP
Git
GitHub
GitHub Actions
GitOps
Grafana
Kubernetes
Loki
Prometheus
Terraform
Thanos
Apply
$45k – $114k per year (Estimated) • Remote • Full-Time • São Paulo • Vitoria-Gasteiz • Fortaleza • Rio de Janeiro • Recife
DevOps
AWS
Azure
Azure DevOps
Bicep
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
GitLab
GitLab CI
IAM
Incident Management
Jenkins
Kubernetes
Terraform
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
IT Administrator 2 hours ago
$83k – $113k per year • In office • Full-Time
Node JS
Python
JavaScript
DevOps
AWS
GCP
GitHub
Cybersecurity
Okta
SOC 2
Management
Confluence
Jira
Slack
Apply
$141k – $308k per year (Estimated) • Equity • In office • Full-Time • Bachelor's Degree • Seattle
C++
Python
C++
LLVM
PyTorch C++
TensorFlow C++
AI/ML
JAX
LLM
PyTorch
Quantization
TensorFlow
Triton
TPU
Apply
$135k – $252k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • PhD • Seattle
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
JAX
PyTorch
TensorFlow
Edge AI
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$135k – $296k per year (Estimated) • Equity • In office • Full-Time • Master's Degree • Seattle
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
JAX
PyTorch
TensorFlow
Edge AI
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$113k – $224k per year (Estimated) • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
SystemVerilog
Verilog
AI/ML
Quantization
Edge AI
Apply
AI Software Engineer 6 months ago
$138k – $259k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
DeepSpeed
LLM
PyTorch
Quantization
SGLang
TensorRT
TensorRT-LLM
vLLM
CUDA
DevOps
Docker
Kubernetes
Apply
$180k – $225k per year • Equity • In office • Full-Time • 4+ years exp • Bachelor's Degree • Seattle
C#
Go
Java
Kotlin
TypeScript
JavaScript
AI/ML
Claude
Claude Code
OpenAI Codex
Frontend
React.js
DevOps
AWS
AWS Step Functions
Apply
$172k – $314k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Seattle
Python
SQL
Python
pySpark
Databases
Databricks
Snowflake
AI/ML
Spark
DevOps
AWS
Apply
$191k – $297k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Seattle
AI/ML
AI Agents
DevOps
AWS
Azure
CI/CD
GCP
IAM
Platform Engineering
Cybersecurity
Least Privilege
Microsoft Entra ID
PCI DSS
Threat Modeling
Zero Trust
Apply
$144k – $180k per year • Equity • Remote/Hybrid • Full-Time • 6+ years exp • Seattle
Java
Python
Scala
SQL
Databases
Amazon Redshift
Apache Kafka
DynamoDB
Snowflake
AI/ML
Feature Store
Spark
DevOps
Amazon Kinesis
Amazon S3
AWS
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.