368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$198k – $370k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 10+ years exp
Overview
Company
Impact
Profile match
d-Matrix is a semiconductor technology company based in Santa Clara, California, and was founded in 2019. The company develops specialized AI inference hardware and software platforms, including its Corsair platform and 3DIMC architecture, designed to accelerate generative AI workloads in data centers. It operates as a venture-backed enterprise serving global hyperscalers and enterprises with a focus on reducing latency, power consumption, and total cost of ownership for large language models.

At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI.

D-Matrix Frontier Group sits at the leading edge of what's possible with LLM inference on heterogeneous hardware. Our charter spans the full stack: from pathfinding emerging use cases and novel deployment patterns to deep optimization of inference kernels, to building proof-of-concept systems that showcase D-Matrix's unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly with product and hardware teams to shape the roadmap.

We build the tools, runtimes, and frameworks that let frontier AI models run efficiently and cost-effectively across heterogeneous deployments - combining D-Matrix silicon with CPUs, GPUs, and custom accelerators. Our work powers everything from benchmarking and evaluation pipelines to production-grade inference serving.

This Role

We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack - from kernel-level optimization to distributed orchestration to high-level serving APIs.

This role could be a great match for you if you:

  • tHave deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.
  • tAre familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.
  • tEnjoy pathfinding new use cases - exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.
  • tAre results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.
  • tAre energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.
  • tValue clear communication and thrive in a small, high-ownership team environment.

Responsibilities

  • tIdentify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.
  • tBuild compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.
  • tDevelop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.
  • tDrive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.
  • tBuild and maintain inference runtimes, serving frameworks, and evaluation tooling.
  • tContribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.
  • tWork closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.
  • tPartner with product and business development to translate POCs into customer-facing demonstrations.
  • tContribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.

Required Qualifications

  • tBachelor's degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.

t

  • tMaster's or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.
  • tStrong proficiency in Python and C/C++.
  • tHands-on experience optimizing LLM inference - attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).
  • tExperience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.
  • tFamiliarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.

Preferred Qualifications

  • tExperience with heterogeneous compute deployments - scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).
  • tFamiliarity with custom silicon or ASIC-based inference (beyond GPU-only environments).
  • tExperience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.
  • tContributions to open-source inference or ML systems projects.
  • tExperience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).
  • tFamiliarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.
  • tWorking familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.

Why D-Matrix Frontier Group

  • tWork on genuinely novel hardware - D-Matrix in-memory compute architecture opens up inference optimization problems that don't exist anywhere else.
  • tEnd-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.
  • tSmall, senior team with high autonomy and direct influence on product direction.
  • tCompetitive compensation, equity, and benefits in Santa Clara, CA.

Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We're committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
Remote • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
Function Calling
LLM
Prompt Engineering
Structured Outputs
Management
n8n
Zapier
Apply
$27k – $110k per year (Estimated) • Remote • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
Function Calling
LLM
Prompt Engineering
Structured Outputs
Management
n8n
Zapier
Apply
$36k – $74k per year (Estimated) • Remote • 3+ years exp • Bachelor's Degree • Yekaterinburg
Bash
C++
Python
DevOps
CI/CD
Cybersecurity
OWASP SAMM
Apply
$67k – $173k per year (Estimated) • In office • Contractor • 3+ years exp • Bachelor's Degree • Singapore
JavaScript
Python
DevOps
AWS
Azure
GCP
Cybersecurity
ISO 27001
OWASP Top 10
Apply
$65k – $169k per year (Estimated) • In office • Contractor • 3+ years exp • Bachelor's Degree • Singapore
Python
Analytics
Power BI
Management
Power Automate
Apply
$155k – $250k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Verilog
Chips/EDA
Cadence Virtuoso
Apply
$155k – $250k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Verilog
Chips/EDA
Cadence Virtuoso
Apply
$175k – $285k per year • In office • Full-Time • 6+ years exp • Santa Clara
C++
Python
AI/ML
Multimodal AI
Apply
$195k – $285k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
C++
Python
AI/ML
CUDA
CUDA Toolkit
JAX
LLM
ONNX
Quantization
SGLang
TensorRT
TensorRT-LLM
Triton
vLLM
Mixture of Experts
Apply
Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Taipei
Python
Apply
$72k – $99k per year • Equity • In office • Full-Time • Santa Clara
Apply
$166k – $290k per year • Equity • In office • Full-Time • 8+ years exp • Santa Clara
Management
ServiceNow
Apply
$133k – $272k per year (Estimated) • In office • Santa Clara
Go
Python
AI/ML
Edge AI
LLM
RAG
Apply
$80k – $110k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Apply
$142k – $256k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara • Toronto
C++
Go
IoT
MQTT
OPC UA
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.