826,151open jobs
53,202companies
135,845added this week
Browse all
Salary
$200k – $275k per year
Location
In office (San Francisco)
Seniority
Staff · 8+ years exp

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Apr 30, 2026.

Overview
Company
Impact
Profile match
Efficient Computer builds the world's most energy-efficient general-purpose AI chips, delivering up to 100× better efficiency than low-power CPUs. Enabling edge AI, on-device AI inference, IoT devices, and industrial automation that run for years on a single charge.

Efficient is developing the world’s most energy-efficient general-purpose computer processor. Efficient’s patented technology uses 100x less energy than state of the art commercially available ultra-low-power processors and is programmable using standard high-level programming languages and AI/ML frameworks. This level of efficiency makes perpetual, pervasive intelligence possible: run AI/ML continuously on a AA battery for 5-10 years. Our platform’s unprecedented level of efficiency enables IoT devices to intelligently capture and curate first-party data to drive the next major computing revolution

We are looking for a Lead RTL Design Engineer to own microarchitecture definition and RTL implementation across the dataflow execution fabric, memory subsystem, on-chip interconnect/NoC, low-power logic, and standard peripheral IP (RiscV, NVM, I2S, I2C) integration. You will work from architecture spec through synthesis-ready RTL, collaborating with architects, microarchitects, DV leads, physical design, and firmware teams to tape out an industry leading power-efficient SoC.

This is a unique opportunity to be a part of a newly formed HW engineering org and have an influence on our products and processes as we move from the initial stages of product development to market release and scaled volume production.  Join our team and help us shape the future of computing at the edge and beyond!

Key Responsibilities 

  • Microarchitecture definition (core focus): Own the design and definition of processor and compute-unit microarchitecture, including dataflow pipelines, execution units, and interfaces. Set performance, power, and area targets, and guide the team toward achieving them.
  • On-chip interconnects and system integration: Define and drive the design of on-chip networks and data movement across the fabric, balancing performance, scalability, and implementation constraints in collaboration with physical design.
  • Memory subsystem & system architecture: Define the interface to the memory subsystem, including data movement, ordering, and synchronization behavior, ensuring a clean and scalable model for software and future system expansion.
  • Reconfiguration and execution model: Lead the architecture of configuration, scheduling, and execution of workloads on the fabric, including multi-kernel support and interaction with host systems.
  • Power management and low-power design: Drive power architecture across the design, including clocking, reset, power domains, and low-power strategies to meet aggressive energy and efficiency goals.
  • HW/SW co-design: Collaborate closely with compiler and software teams to define the hardware execution model, ensuring efficient mapping of workloads onto the architecture.

Specifications, Documentations and Reviews: Author and own uArch specification documents for assigned blocks; drive design reviews with architecture, compiler, DV, and physical design stakeholders.

  • Mentoring and process improvement: Mentor senior and junior RTL engineers; review RTL, flag microarchitecture risks, and enforce coding style and lint-clean standards across the team.
  • Driving PPA Metrics: Participate in PPA analysis loops: synthesize blocks regularly, review area/timing/power reports, and make data-driven tradeoffs against performance and feature requirements.
  • DV Collaboration: Collaborate with DV leads to define/review verification plans; provide directed test scenarios for graph execution corner cases, back-pressure conditions, and power state transitions.
  • Silicon Bring-up: Support silicon bring-up: contribute scan/ATPG guidelines, review DFT insertion, and provide RTL-level debug assistance during lab validation.

Required Qualifications & Experience

  • 8+ years of RTL design experience with tape-out ownership of dataflow based design, on chip networks, memory subsystems  or peripheral integration on a processor or accelerator SoC.
  • Deep proficiency in SystemVerilog for RTL - synthesis-clean, lint-clean, timing-aware; able to design complex state machines, arbiters, token flow controllers, and datapath logic from scratch.
  • Solid understanding of parallel execution models: dataflow, SIMD, or systolic array architectures; familiarity with the hardware challenges of token-based firing-rule evaluation and producer-consumer synchronization.
  • Hands-on experience with on-chip memory design: SRAM wrappers, scratchpad/TCM, banking, and memory-mapped register interfaces.
  • Experience with low-power RTL techniques: UPF-driven flows, clock gating, power domains, retention registers, and AON wakeup logic.
  • Familiarity with at least one standard on-chip bus protocol (AXI, AHB, APB, TileLink, or NoC equivalent) at the RTL implementation level.
  • Experience taking RTL through synthesis and timing closure; ability to read and act on SDC constraints, STA reports, and synthesis QoR summaries.
  • Strong written communication skills; able to produce uArch specs and design review material independently.
  • Experience with memory compiler toolchains

Desired Qualifications & Experience Requirements 

  • Prior RTL ownership of a dataflow engine, neural processing unit (NPU), or streaming DSP architecture with explicit producer-consumer token management.
  • Experience collaborating with compiler or graph-optimization teams to co-design hardware execution models and graph IR representations.
  • Familiarity with NVM controller RTL (MRAM, RRAM) including ECC, program/erase sequencing, and model weight storage use cases.
  • Experience with IoT-class power budgets (sub-10 mW active, sub-100 µW standby) and the RTL design choices they necessitate.
  • Familiarity with functional safety standards (ISO 26262, IEC 61508) as applied to execution fabric error detection and power domain isolation.
  • Exposure to AI framework graph formats (ONNX, TFLite) and understanding of how graph compilation maps to hardware execution primitives.
  • Tape-out credits on an edge-AI, IoT, or wearable SoC at 12nm or below.
  • Experience with formal verification of flow-control logic, deadlock freedom, or bus protocol compliance.

We offer a competitive base salary for this role, generally ranging from $200,000 to $275,000, plus a 10% annual bonus, meaningful equity, and comprehensive benefits. Final compensation will depend on experience, level, and location, with flexibility for the right candidate.

Why Join Efficient?

Efficient offers a competitive compensation and benefits package, including  401K match, company-paid benefits, equity program, paid parental leave, and flexibility. We are committed to personal and professional development and strive to grow together as people and as a company.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
826,151 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Hardware
Similar stack
Same company
San Francisco
$115k – $173k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Budd Lake
MATLAB
Chips/EDA
OrCAD
Apply
≈ $153k – $285k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Sunnyvale
SystemVerilog
AI/ML
Vertex AI
TPU
Edge AI
DevOps
GCP
CI/CD
Chips/EDA
Formal Verification
Synopsys VC Formal
Apply
≈ $133k – $248k per year (Estimated) • In office • 5+ years exp • Austin
Apply
≈ $137k – $256k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Austin
Python
Verilog
SystemVerilog
Perl
DevOps
HPC
Linux
Unix
Chips/EDA
Synopsys Design Compiler
Synopsys PrimeTime
Synopsys Fusion Compiler
Synopsys SpyGlass
Synopsys VC Formal
Synopsys PrimePower
Apply
≈ $139k – $268k per year (Estimated) • In office • 8+ years exp • Austin
SystemVerilog
AI/ML
Machine Learning
Chips/EDA
Siemens ModelSim
UVM
Apply
≈ $55k – $142k per year (Estimated) • In office • Full-Time • PhD • Seongnam
Verilog
SystemVerilog
Chips/EDA
Xilinx Vivado
Formal Verification
Apply
≈ $113k – $247k per year (Estimated) • Remote (likely United States) • 5+ years exp • PhD
Python
Java
Rust
SQL
C++
Scala
C++
TensorFlow C++
PyTorch C++
Databases
Apache Kafka
AI/ML
Copilot
CatBoost
Spark
XGBoost
Triton Inference Server
AI Agents
LightGBM
ONNX
TensorRT
Flink
TensorFlow
PyTorch
Ray
Google AI Studio
Feature Store
Recommender Systems
Agentic Workflows
Machine Learning
DevOps
CI/CD
Apply
≈ $55k – $142k per year (Estimated) • In office • Full-Time • PhD • Seongnam
Verilog
SystemVerilog
Chips/EDA
Xilinx Vivado
Synopsys ZeBu
Siemens Veloce
OrCAD
Mentor PADS
Apply
≈ $48k – $101k per year (Estimated) • In office • 10+ years exp • Bengaluru
Python
Go
JavaScript
AI/ML
Model Context Protocol
MLFlow
AI Agents
NLP
TensorFlow
PyTorch
LLM
Few-Shot Learning
Synthetic Data
LiteRT
Machine Learning
DevOps
Rest API
GCP
Azure
AWS
Apply
Hybrid • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Verilog
SystemVerilog
TCL Scripting
Chips/EDA
UVM
Apply
≈ $91k – $235k per year (Estimated) • In office • Internship • Bachelor's Degree • Austin
Python
Verilog
C++
SystemVerilog
Perl
Apply
$160k – $245k per year • In office • 10+ years exp • Bachelor's Degree • San Francisco
Python
C++
SystemVerilog
AI/ML
AI Agents
Chips/EDA
UVM
Cadence JasperGold
Formal Verification
Apply
$160k – $220k per year • In office • 8+ years exp • San Francisco
C++
Apply
$160k – $230k per year • In office • 5+ years exp • Bachelor's Degree • Pittsburgh
Python
SystemVerilog
Chips/EDA
Cadence Innovus
Synopsys PrimeTime
Synopsys Fusion Compiler
Cadence Xcelium
UVM
Formal Verification
Apply
$160k – $230k per year • In office • 6+ years exp • Bachelor's Degree • San Francisco
Verilog
C
C++
SystemVerilog
VHDL
C
GCC
C++
LLVM
AI/ML
MLIR
Machine Learning
Apply
In office • 2+ years exp • Bachelor's Degree • San Francisco
Python
Verilog
C++
SystemVerilog
Bash
DevOps
Git
Linux
Chips/EDA
UVM
Xilinx Vivado
Vitis
Apply
$130k – $175k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Apply
≈ $105k – $207k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$140k – $295k per year • In office • 5+ years exp • San Francisco
Apply
$100k – $125k per year • In office • 2+ years exp • San Francisco
Python
Apply
See all jobs
This is one of many
826,151 more open roles from verified company boards, updated every day.