368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$150k – $250k per year
Location
In office (San Jose)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Etched. We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads.

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

Building cutting-edge model-specific ASICs requires crafting custom infrastructure and toolchains to support ultra-fast, reliable, and scalable development across the stack - from simulation to silicon. We build this infrastructure as software - and we engineer it with the same best practices we apply to our products. We use the same rigor, design discipline, and quality standards and testing as we do to our ASIC, software, and platform.

You will lead the development and adoption of next-generation infrastructure tooling, enabling Etched ASIC, Software, and Platform engineers to iterate faster, build more reliably, and push the boundaries of AI performance. This includes building and scaling our hybrid high-performance compute (HPC) cluster, optimized for massively parallel CI, EDA workflows, Emulation, and hardware-aware job execution.

You’ll also architect and implement a state-of-the-art observability stack with LLM integration and a strong emphasis on streaming health and performance telemetry, log aggregation, distributed tracing, insight generation, synthetic testing, and smart alerting - across CI pipelines, simulation clusters, and service endpoints.

This role demands a strong software engineering mindset, quality instincts, and deep understanding of systems. It’s not just about writing scripts - it’s about writing code that builds and manages infrastructure with precision, repeatability, and intent.

Key responsibilities

  • Design and build the orchestration layers that drive our hybrid high-performance clusters-enabling simulation, synthesis, and continuous integration of AI ASICs at unprecedented scale.

  • Develop and maintain a fully programmable infrastructure control plane to ensure reproducibility, auditability, and rapid iteration across the entire stack.

  • Create tools and abstractions that empower engineers to harness massive parallelism without worrying about the underlying complexity..

  • Prototype and execute workload orchestration and migration strategies between on-premise and cloud environments, balancing performance, storage availability and replication, uptime, and cost across heterogeneous hardware and compute backends.

  • Implement real-time telemetry, tracing systems that surface insights from millions of metrics, enabling proactive debugging and system optimization.

  • Build a full observability stack that includes dashboards, alerting, automated responses, and a synthetic testing framework to proactively test infrastructure performance and reliability for various application and data flows, ensuring we remain proactive against issues impacting development and productivity workflows.

Representative projects

  • Design and deploy a fully automated, scalable hybrid HPC cluster, combining bare-metal servers and switches with cloud instances, provisioned through MaaS and orchestrated via SLURM and Kubernetes, optimized for mixed EDA workloads and parallel CI pipelines.

  • Develop a real-time observability system for ASIC toolchain jobs and distributed builds, integrating Prometheus, Grafana, and VictoriaMetrics with streaming telemetry, tracing, and alerting to detect performance regressions before they hit silicon.

  • Architect and implement a programmable infrastructure-as-code control plane, using Terraform, Ansible, and Puppet, to version, audit, and redeploy every layer of Etched's development stack with deterministic reproducibility.

  • Create a zero-downtime interactive development environment that provisions and connects Jupyter and VS Code sessions to GPUs and high-memory nodes via a secure zero-trust network, abstracting away cluster state and machine failures.

  • Prototype and evaluate dynamic workload migration strategies between on-premise and cloud environments to optimize for latency, reliability, and cost across simulation and synthesis pipelines.

  • Design a synthetic testing and fault injection framework to validate the behavior of infrastructure under high-load, degraded hardware, and intermittent network partitions - before they happen in production.

You may be a good fit if you

  • Are a systems-minded software engineer who loves building foundational platforms, working close to the metal and cloud, solving high-leverage problems at scale.

  • Are a deeply technical engineer who treats infrastructure as a software problem - prioritizing clean abstractions, version control,small change lists, easy roll backs, testing, and long-term maintainability over ad hoc configuration.

  • Have strong programming skills in languages such as Python, Go, Rust, and C++, and are comfortable building production-grade tooling.

  • Possess expert-level knowledge of Linux, virtualization, containerization, and CI/CD pipelines, with a deep understanding of how to debug, optimize, and scale complex systems.

  • Are familiar with Infrastructure as Code tools like OpenTofu, Ansible, or Puppet, and enjoy designing declarative, reproducible infrastructure systems.

  • Understand and use PromQL and other telemetry/query languages and have used LLM to extract insight from real-time metrics, and know how to architect and tune observability stacks.

  • Have a track record of debugging and resolving difficult hardware-software integration problems across bare-metal systems, networks, and distributed workloads.

  • Can lead and mentor technical teams, guiding design decisions and helping others develop sound engineering instincts.

  • Have 8+ years of experience in infrastructure engineering, systems programming, or backend software development - ideally in environments where performance, scale, or hardware interaction mattered.

  • Are driven by curiosity, take initiative, and have an innate sense of ownership - you thrive in uncharted territory, design for edge cases, and love making systems more powerful, reliable, and elegant.

Strong candidates may also have experience with

  • Familiarity with Bazel build system

  • Deep understanding of ASIC development flows, especially those involving Synopsys, Cadence, and Verilator, including how EDA tools interact with infrastructure for simulation, synthesis, and verification.

  • Hands-on experience architecting systems with AWS, GCP, or Azure, including hybrid on-prem/cloud deployments, workload migration strategies, and cloud-native orchestration tooling.

  • Experience monitoring, provisioning, and debugging bare-metal servers, network hardware, and high-performance storage systems in rack-scale environments.

  • Comfortable in profiling and optimizing compute environments for single-threaded latency, memory-bound workloads, or I/O throughput, especially in the context of simulation or CI performance.

  • Proficiency building or operating telemetry systems at scale using Prometheus, Grafana, Loki, VictoriaMetrics, and tools for distributed tracing, log aggregation, and real-time alerting across heterogeneous mediums (SMS, email, push alerts, etc.)

Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2k per month for those living within walking distance of the office

  • Relocation support for those moving to San Jose (Santana Row)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch + dinner in our office

  • Unlimited compute budget subject to ROI justification

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$71k – $120k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Kraków
DevOps
AWS
Azure
CI/CD
GCP
Apply
Principal Engineer 11 hours ago
$27k – $71k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Groovy
JavaScript
Python
Java
Java
Gradle
Spring Boot
Databases
DynamoDB
ElasticSearch
MySQL
Redis
Frontend
React.js
Redux
Mobile
JUnit
DevOps
AWS
AWS Lambda
CI/CD
Docker
Git
Jenkins
Amazon ECS
Amazon Kinesis
Amazon S3
API Gateway
Design
AutoCAD
Fusion 360
Management
Jira
QA
JMeter
Apply
Lead Product Engineer 11 hours ago
Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • South Africa
COBOL
SQL
COBOL
Easytrieve
JCL
Databases
Db2
IMS
DevOps
Azure
Azure DevOps
Incident Management
Apply
Remote • Full-Time • 7+ years exp • Bucharest
JavaScript
Node JS
TypeScript
AI/ML
LLM
A2A
Model Context Protocol
DevOps
AWS
Platform Engineering
Apply
$78k – $141k per year (Estimated) • Remote • Full-Time • 7+ years exp • Kraków
JavaScript
Node JS
TypeScript
AI/ML
LLM
A2A
Model Context Protocol
DevOps
AWS
Platform Engineering
Apply
$150k – $275k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
AI/ML
LLM
DevOps
HPC
Chips/EDA
Ansys HFSS
Ansys SIwave
Cadence Sigrity
Keysight ADS
Apply
Equity • In office • Full-Time • 5+ years exp • Taipei
Apply
$175k – $225k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • San Jose
C#
Go
Java
Python
TypeScript
DevOps
CI/CD
Apply
$200k – $300k per year • In office • Full-Time • 10+ years exp • Austin
Perl
Python
AI/ML
LLM
Chips/EDA
Formal Verification
Synopsys Fusion Compiler
Synopsys PrimeTime
Apply
In office • Full-Time • 2+ years exp • Bachelor's Degree • Taipei
Python
C#
C#
.NET
Chips/EDA
Cadence Allegro
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
$68k – $85k per year • In office • Full-Time • Master's Degree • San Jose
Python
AI/ML
AI Agents
DevOps
Amazon EC2
AWS
AWS Lambda
Bitbucket
CI/CD
CloudFormation
Docker
Git
Kubernetes
Terraform
Amazon S3
IAM
HPC
Cybersecurity
Least Privilege
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.