394,870open jobs
13,845companies
76,904added this week
Browse all
Location
In office (San Jose)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
Etched. We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads.

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking talented Fall '26, Spring '27, and Summer '27 Inference Architecture interns to join our team and contribute to the design of next-generation AI accelerators. This role focuses on developing and optimizing compute architectures that deliver exceptional performance and efficiency for inference workloads. You will work on cutting-edge architectural problems and performance modeling over the course of your internship.

Key responsibilities

  • Support porting state-of-the-art models to our architecture. Help build programming abstractions and testing capabilities to rapidly iterate on model porting.

  • Assist in building, enhancing, and scaling our runtime, including multi-node inference, intra-node execution, state management, and robust error handling.

  • Contribute to optimizing routing and communication layers using our collectives.

  • Utilize performance profiling and debugging tools to identify bottlenecks and correctness issues.

  • Develop and leverage a deep understanding of our architecture to co-design both HW instructions and model architecture operations to maximize model performance

  • Implement high-performance software components for the Model Toolkit

You may be a good fit if you have

  • Progress towards a Bachelor’s, Master’s, or PhD degree in computer science, computer engineering, applied mathematics, or a related field

  • Proficiency in Python, C++

  • Understanding of performance-sensitive or complex distributed software systems, e.g. Linux internals, accelerator architectures (e.g. GPUs, TPUs), Compilers, or high-speed interconnects (e.g. NVLink, InfiniBand).

  • Ported applications to non-standard accelerator hardware or hardware platforms.

  • Deep knowledge of transformer model architectures and/or inference serving stacks (vLLM, SGLang, etc.)

Strong candidates may have some experience with

  • Proficiency in Rust

  • Low-latency, high-performance applications using both kernel-level and user-space networking stacks.

  • Deep understanding of distributed systems concepts, algorithms, and challenges, including consensus protocols, consistency models, and communication patterns.

  • Solid grasp of Transformer architectures, particularly Mixture-of-Experts (MoE).

  • Built applications with extensive SIMD (Single Instruction, Multiple Data) optimizations for performance-critical paths.

  • Familiarity with PyTorch or JAX.

  • Math competitions (AIME, AMC, etc)

We encourage you to apply even if you do not believe you meet every qualification.

Program details

  • 12-week paid internship

  • Generous housing support for those relocating

  • Daily lunch and dinner in our office

  • Based at our office in San Jose, CA

  • Direct mentorship from industry leaders and world-class engineers

  • Opportunity to work on one of the most important problems of our time

For any questions, contact [email protected].

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
394,870 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
C++
C
C++
LLVM
PyTorch C++
TensorFlow C++
C
MPI
AI/ML
CUDA
CUDA Toolkit
CUTLASS
DeepSpeed
JAX
MLIR
NCCL
PyTorch
ROCm
TensorFlow
TensorRT
Triton
vLLM
DevOps
AWS
Azure
GCP
HPC
Apply
Algorithm Engineer 5 min ago
Remote/Hybrid • Full-Time • 3+ years exp • Netanya
C++
MATLAB
Python
AI/ML
Computer Vision
Quantization
DevOps
CI/CD
HPC
Analytics
A/B Testing
Apply
$74k – $98k per year • Remote • 7+ years exp • Bachelor's Degree
C++
Python
Rust
AI/ML
Knowledge Distillation
LLM
Model Distillation
Quantization
TensorRT
TensorRT-LLM
vLLM
DevOps
FinOps
Kubernetes
Platform Engineering
Apply
$175k – $200k per year • Remote • 10+ years exp • Bachelor's Degree
C#
C++
Go
Java
DevOps
AWS
Azure
CI/CD
FinOps
GCP
Kubernetes
Service Mesh
Apply
$95k – $120k per year • Remote • 8+ years exp • Bachelor's Degree
C++
Python
Databases
InfluxDB
TimescaleDB
PostgreSQL
AI/ML
Time Series Forecasting
DevOps
AWS
Azure
GCP
Robotics
Digital Twin
IoT
MQTT
OPC UA
Apply
$150k – $275k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
AI/ML
LLM
DevOps
HPC
Chips/EDA
Ansys HFSS
Ansys SIwave
Cadence Sigrity
Keysight ADS
Apply
IT Engineer 10 days ago
$150k – $250k per year • In office • Full-Time • 7+ years exp • San Jose
Bash
PowerShell
Python
DevOps
AWS
Azure
GCP
Platform Engineering
Cybersecurity
Crowdstrike
ISO 27001
Okta
SentinelOne
SOC 2
Zero Trust
Management
Google Workspace
Apply
$175k – $275k per year • In office • Full-Time • San Jose
AI/ML
CUDA
CUDA Toolkit
Apply
In office • Full-Time • San Jose
Bash
Python
Apply
In office • Contractor • 2+ years exp • San Jose
Apply
In office • Full-Time • 15+ years exp • PhD • San Jose
Python
SystemVerilog
AI/ML
Claude
Chips/EDA
UVM
Apply
$115k – $228k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Jose
Perl
Python
SystemVerilog
AI/ML
Claude
Chips/EDA
UVM
Apply
$144k – $198k per year • In office • 2+ years exp • Bachelor's Degree • San Jose
Apply
$173k – $324k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • San Jose
Go
Java
Python
JavaScript
Java
Spring Boot
Python
FastAPI
Databases
Apache Kafka
PostgreSQL
RabbitMQ
Redis
Kafka
AI/ML
Dagster
LLM
Prefect
RAG
AI Agents
Function Calling
Tool Use
Frontend
Material UI
Next.js
React.js
DevOps
CI/CD
Docker
gRPC
Kubernetes
Cybersecurity
Zero Trust
Zscaler
Apply
$172k – $308k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Jose • Seattle • San Francisco
Apply
See all jobs
This is one of many
394,870 more open roles from verified company boards, updated every day.