368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$150k – $275k per year
Location
In office (San Jose)
Employment
Full-Time
Overview
Company
Impact
Profile match
Etched. We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads.

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

Join our team and take the lead in illuminating the performance landscape of our cutting-edge ML accelerator. We are seeking a highly skilled engineer to design and develop a sophisticated performance analysis tool, tailored specifically for our hardware. You will be instrumental in creating the essential tooling that enables our ML engineers and customers to understand workload behavior, identify performance bottlenecks, and unlock the full potential of our hardware, accelerating the most demanding ML applications in the world. This is a unique opportunity to shape performance analysis for novel hardware from the ground up.

Key responsibilities

  • Lead the design and architecture of a comprehensive performance analysis suite, including data collection mechanisms, data processing pipelines, analysis engines, and user interfaces (CLI and/or GUI).

  • Develop robust methods to capture performance data directly from our custom ML accelerator hardware (e.g., hardware performance counters, execution unit status, memory access patterns) via driver interfaces or other mechanisms.

  • Implement tracing for host-side API calls (runtime libraries, driver interactions) and system-level events (CPU activity, PCIe traffic, memory usage, network contention) related to our workloads.

  • Design and implement techniques to accurately correlate performance events across the host CPU, device driver, PCIe bus, multiple accelerators, and multiple hosts, ensuring precise time synchronization.

  • Build analysis modules to automatically interpret collected trace and counter data, identifying key performance limiters (e.g., compute-bound, memory bandwidth-bound, latency-bound, PCIe-bound, specific hardware bottlenecks).

  • Develop intuitive visualizations (timelines, dependency graphs, resource utilization charts, statistical summaries) to clearly communicate performance characteristics and bottlenecks to users.

  • Work closely with hardware architects, firmware engineers, driver developers, compiler engineers, and ML application engineers to understand their needs, define tool requirements, and provide expert guidance on performance analysis and optimization using the tool.

Representative projects

  • Architect and implement the core data collection framework for hardware performance counters on a custom PCIe-based accelerator.

  • Develop a kernel driver module or user-space service for low-overhead tracing of accelerator activity.

  • Design and build a correlated timeline view visualizing CPU API calls, driver submissions, PCIe transfers, and accelerator execution units.

  • Create an analysis pass to detect and quantify memory access inefficiencies or PCIe bandwidth saturation while transacting on a PCIe-attached accelerator.

You may be a good fit if you have

  • Strong proficiency in C++ or Rust

  • Proficiency in Python is a plus

  • Deep understanding of computer architecture (CPU, GPU, accelerators), memory hierarchies (caches, DRAM), and interconnects (especially PCIe).

  • Proven experience in low-level performance analysis, profiling, and bottleneck identification on complex hardware systems (GPUs, CPUs, FPGAs, or custom accelerators).

  • Experience with performance analysis tools (e.g., NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW).

  • Experience working close to hardware, potentially reading performance counters or interacting directly with device drivers.

Strong candidates may also have experience with (Nice-to-have qualifications)

  • Direct experience developing performance analysis or debugging tools.

  • Experience with ML accelerator architectures (GPUs, TPUs, etc.).

  • Experience with kernel-mode driver development (Linux or Windows).

  • Understanding of compiler internals, code generation, and optimization.

  • In-depth knowledge of the PCIe protocol and analysis tools (PCIe analyzers).

  • Experience with multi-chip or multi-host accelerator systems (e.g., TPU pods, or NVidia DGX clusters)

  • Experience with firmware or embedded systems development.

  • Experience with hardware description languages (Verilog, VHDL) or hardware verification.

Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2k per month for those living within walking distance of the office

  • Relocation support for those moving to San Jose (Santana Row)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch + dinner in our office

  • Unlimited compute budget subject to ROI justification

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$25k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Chennai
Java
PowerShell
Python
TypeScript
JavaScript
Java
Spring Boot
Databases
Amazon Redshift
DynamoDB
Frontend
Angular
DevOps
Amazon EC2
Amazon EKS
Ansible
ArgoCD
AWS
AWS Lambda
Azure
Azure DevOps
Bicep
Chef
CI/CD
CloudFormation
Configuration Management
Docker
GCP
GitLab CI
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
Amazon S3
GitLab
IAM
Apply
$120k – $240k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • New York
C#
C++
Go
JavaScript
TypeScript
Frontend
React.js
Redux
DevOps
AWS
Azure
GCP
IAM
Cybersecurity
FedRAMP
Apply
Sr. Analyst 4 hours ago
$56k – $113k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Atlanta
Python
SQL
Apply
Data Architect 4 hours ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$23k – $62k per year (Estimated) • In office • Full-Time • 3+ years exp • Navi Mumbai • Pune
Python
Apply
$150k – $275k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
AI/ML
LLM
DevOps
HPC
Chips/EDA
Ansys HFSS
Ansys SIwave
Cadence Sigrity
Keysight ADS
Apply
Equity • In office • Full-Time • 5+ years exp • Taipei
Apply
$175k – $225k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • San Jose
C#
Go
Java
Python
TypeScript
DevOps
CI/CD
Apply
$200k – $300k per year • In office • Full-Time • 10+ years exp • Austin
Perl
Python
AI/ML
LLM
Chips/EDA
Formal Verification
Synopsys Fusion Compiler
Synopsys PrimeTime
Apply
In office • Full-Time • 2+ years exp • Bachelor's Degree • Taipei
Python
C#
C#
.NET
Chips/EDA
Cadence Allegro
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.