1,138,035open jobs
65,468companies
202,917added this week
Browse all
Salary
≈ $79k – $156k per year (Estimated)
Location
In office (Las Vegas)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Oct 1, 2026.

Overview
Company
Impact
Profile match
TensorWave is an American cloud infrastructure company headquartered in Las Vegas, Nevada, and founded in 2023. The company provides a specialized AI cloud platform powered by AMD Instinct accelerators, offering bare metal instances, managed Kubernetes, and high-speed storage for large-scale model training and inference. It serves enterprise clients and AI research teams globally, positioning itself as a high-performance alternative to NVIDIA-based cloud providers through its focus on open-source ROCm software and cost-efficient scaling.

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

About the Role

We are looking for a Hardware Diagnostics Engineer to run burn-in, triage what fails, work servers out-of-band, and own RMAs end to end. If you like hardware that misbehaves in ways that take real work to explain, this is a good seat.

Before any GPU server carries a customer workload, it has to prove it works - under load, at temperature, for hours. When it doesn't, somebody has to figure out why, get replacement hardware in, and send the failed part back to the vendor.

What You’ll Do

  • Run server and GPU burn-in and stress testing, interpret the results, and decide whether hardware is production-ready

  • Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling: reproduce the failure, isolate the faulty component, and document what proved it

  • Work servers out-of-band through IPMI and Redfish for power control, boot configuration, BIOS settings, and sensor and event log collection

  • Apply firmware updates across the fleet following the team's qualified baselines and rollout process

  • Drive RMAs with vendors from ticket through replacement, installation, and return of the failed part

  • Keep asset, serial, and replacement history accurate in NetBox so we know what's actually in every rack

  • Track failure patterns across the fleet and raise them when the same part or firmware version keeps turning up

  • Improve the runbooks you work from, and script the steps you find yourself repeating

  • Partner with datacenter operations on hands-on work during turn-ups and expansions

  • Take part in an on-call and escalation rotation for hardware issues

Who You Are

Required Qualifications

  • 3-6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering

  • Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, and reading hardware behavior at the rack

  • Practical experience with BMCs and out-of-band management: IPMI, Redfish, iDRAC, iLO, or equivalent

  • Strong Linux troubleshooting: boot process, driver and device issues, and diagnostic tools such as {{dmesg}}, {{lspci}}, {{ipmitool}}, and SMART

  • Comfort reading sensor data, event logs, and thermal and power telemetry well enough to tell a real failure from noise

  • Working scripting ability in Bash or Python - enough to automate a repetitive task and read someone else's tooling

  • Experience running hardware RMAs with vendors, or a clear track record of driving issues to closure with outside parties

  • A methodical troubleshooting habit: you isolate variables, you don't change three things at once, and you can say what evidence led to your conclusion

  • Clear written communication for tickets, runbooks, and vendor cases

Preferred Qualifications

  • GPU server experience, especially AMD GPUs and ROCm

  • Burn-in, stress testing, or node validation tooling in a GPU or HPC environment

  • Familiarity with firmware update processes and why fleet-wide changes get staged

  • NetBox or other DCIM and IPAM tooling

  • Ansible, or Python against REST APIs

  • Prior work in a high-volume hardware environment: hyperscaler, colo, integrator, or manufacturing test

First Six Months

  • By 90 days you'll run burn-in cycles and triage failures independently from our runbooks, and you'll have driven at least one RMA to closure. By six months you're the person who spots the pattern before anyone else - this batch, this firmware, this part - and you've automated at least one step you used to do by hand.

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,138,035 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Hardware
Similar stack
Same company
Las Vegas
≈ $115k – $216k per year (Estimated) • In office • Secret • 1+ year exp • Bachelor's Degree • United States
Apply
$104k – $166k per year • In office • Palestine
Apply
$99k – $206k per year • In office • TS/SCI • Full-Time • 6+ years exp • Sterling
Apply
$108k – $167k per year • Equity • In office • 3+ years exp • Bachelor's Degree • Santa Cruz
C++
Apply
$52k – $80k per year • In office • Full-Time • 3+ years exp • Associate's Degree • Phoenix
Apply
$88k – $130k per year • Hybrid • Bachelor's Degree • San Diego
Java
SQL
Bash
Java
Apache Tomcat
Databases
Oracle
DevOps
Azure
Jenkins
Linux
Cybersecurity
Active Directory
Apply
≈ $17k – $37k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Bengaluru
PowerShell
Bash
Databases
MySQL
PostgreSQL
DevOps
Terraform
Ansible
Zabbix
VMWare
Azure
Windows Server
Git
AWS
Docker
Kubernetes
Nagios
Proxmox VE
Incident Management
GitHub
GitLab
Windows
TCP/IP
DNS
DHCP
Cybersecurity
Microsoft Entra ID
Active Directory
Management
Jira
ServiceNow
ITIL
ITSM
Apply
SaaS Technical Lead 2 days ago
≈ $57k – $117k per year (Estimated) • In office • 8+ years exp
Python
PowerShell
DevOps
Splunk
Dynatrace
AWS
Platform Engineering
Amazon EC2
SLI/SLO/SLA
IAM
Management
ITIL
Apply
≈ $47k – $80k per year (Estimated) • In office • Internship • Bloomington
Python
AI/ML
TensorFlow
PyTorch
Machine Learning
Apply
≈ $70k – $179k per year (Estimated) • In office • TS/SCI • 3+ years exp • Bachelor's Degree • Quantico
DevOps
Linux
Windows
Cybersecurity
Nessus
Active Directory
SIEM
Robotics
Digital Twin
Analytics
Power BI
Management
ServiceNow
ITSM
Apply
≈ $143k – $238k per year (Estimated) • Equity • Remote (United States) • Full-Time • 5+ years exp
Python
Go
JavaScript
Node JS
Python
FastAPI
Node JS
Express
Databases
Redis
Frontend
GraphQL
Next.js
React.js
DevOps
Rest API
OpenTelemetry
Docker
Kubernetes
Grafana
Apply
≈ $125k – $231k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • Master's Degree • Las Vegas
AI/ML
ROCm
Apply
≈ $49k – $88k per year (Estimated) • Equity • In office • Full-Time • 1+ year exp • Bachelor's Degree • Las Vegas
AI/ML
ROCm
Design
Figma
Canva
Marketing
Salesforce
HubSpot
Apply
≈ $108k – $197k per year (Estimated) • Equity • Hybrid • Full-Time • 8+ years exp • Las Vegas
Apply
≈ $125k – $228k per year (Estimated) • Equity • Hybrid • Full-Time • 5+ years exp • Las Vegas
Apply
$56k – $84k per year • Equity • In office • 10+ years exp • Bachelor's Degree • Las Vegas
Apply
Service Supervisor 10 hours ago
≈ $44k – $105k per year (Estimated) • Equity • In office • High School Diploma • Las Vegas
Apply
$34k per year • In office • Full-Time • Las Vegas
Apply
$68k – $70k per year • In office • Full-Time • 2+ years exp • Associate's Degree • Las Vegas
Management
Microsoft Office
Apply
≈ $97k – $213k per year (Estimated) • In office • Part-Time • Las Vegas
Apply
See all jobs
This is one of many
1,138,035 more open roles from verified company boards, updated every day.