435,822open jobs
15,110companies
67,093added this week
Browse all
Salary
$150k – $350k per year
Location
In office (San Jose, Santa Barbara)
Employment
Full-Time
Overview
Company
Impact
Profile match
ChipAgents builds artificial intelligence agents for chip design and verification. Its systems write testbenches and debug register transfer level code. The company serves semiconductor engineering teams.

About ChipAgents

ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier-1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.

Position Overview

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Key Responsibilities

  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.

  • Implement and benchmark concrete inference optimizations.

  • Profile and analyze inference bottlenecks at the systems level-from GPU kernel execution to memory bandwidth constraints.

  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.

  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.

  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.

Qualifications

  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).

  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.

  • Strong proficiency in Python and C++/CUDA; hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.

  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.

  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.

  • Strong systems-level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.

  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.

  • Self-directed problem solver who is interested in working on ambitious optimization challenges.

Why Join Us

  • Work on cutting-edge LLM inference optimization problems with real-world production impact.

  • Access to substantial GPU compute resources for experimentation and benchmarking.

  • Collaborate with a world-class team spanning AI research, systems engineering, and EDA.

  • Shape the performance characteristics of AI systems used by leading semiconductor companies.

What we offer

  • $150K/yr - $350K/yr + Offers Equity. We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.

  • Unlimited PTO and full benefits (medical, vision, dental, 401k).

  • Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,822 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$148k – $185k per year • Remote/Hybrid • Full-Time • 2+ years exp • Zurich
Python
TypeScript
AI/ML
Spark
Apply
$22k – $27k per year • Remote • 1+ year exp • Moscow
Python
JavaScript
Python
FastAPI
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
LLM
RAG
DevOps
Docker
Apply
$102k per year • Remote/Hybrid • Full-Time • Reutlingen
C++
VHDL
Apply
Security Architect 8 hours ago
$27k – $64k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Python
AI/ML
LangChain
LlamaIndex
Model Context Protocol
Function Calling
AI Agents
Promptfoo
LLM
RAG
OpenAI
Anthropic
Red Teaming
LLM Evaluation
LLM Guardrails
NIST AI RMF
Agentic Workflows
Tool Use
DevOps
CI/CD
Cybersecurity
Least Privilege
Threat Modeling
Apply
$19k – $45k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Indore
Python
AI/ML
Copilot
AI Agents
RAG
OpenAI
DevOps
Rest API
Terraform
Ansible
Azure
AWS
Configuration Management
AIOps
Cybersecurity
Zscaler
Zero Trust
Management
ServiceNow
Apply
$150k – $200k per year • In office • Full-Time • 4+ years exp • San Jose
SQL
AI/ML
AI Agents
LLM
Chips/EDA
Formal Verification
Management
Google Sheets
Apply
$175k – $225k per year • In office • Full-Time • 5+ years exp • San Jose
AI/ML
AI Agents
LLM
Chips/EDA
Formal Verification
Apply
$180k – $350k per year • In office • Full-Time • San Jose
AI/ML
AI Agents
Design
Figma
Apply
$150k – $350k per year • In office • Full-Time • 3+ years exp • San Jose
AI/ML
AI Agents
Apply
$180k – $450k per year • In office • Full-Time • 7+ years exp • San Jose
AI/ML
AI Agents
Chips/EDA
UVM
Formal Verification
Apply
FP&A Manager 5 hours ago
$84k – $167k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Jose
Analytics
Microsoft Excel
Apply
$203k – $290k per year • Remote/Hybrid • Full-Time • San Jose
Cybersecurity
Zscaler
Zero Trust
Apply
$64k – $157k per year (Estimated) • In office • Full-Time • San Jose
Apply
$64k – $157k per year (Estimated) • In office • Full-Time • San Jose
Apply
$64k – $157k per year (Estimated) • In office • Full-Time • San Jose
Apply
See all jobs
This is one of many
435,822 more open roles from verified company boards, updated every day.