574,821open jobs
24,300companies
79,096added this week
Browse all
Salary
$146k – $296k per year (Estimated)
Location
In office (San Jose)
Employment
Full-Time
Overview
Company
Impact
Profile match

Who we are:

Persimmons is building the infrastructure that will power the next decade of AI. Founded in 2023 by veteran technologists from the worlds of semiconductors, AI systems, and software innovation, We’re on a mission to enable smarter devices, more sustainable data centers, and entirely new applications the world hasn’t imagined yet.

Why join us:

We’re growing fast and looking for bold thinkers, builders, and curious problem-solvers who want to push the limits of AI hardware and software. If you're ready to join a world-class team and play a critical role in making a global impact - we want to talk to you.

Summary of Role:

This role focuses on transforming higher-level MLIR-based large language models by applying sophisticated mid- and backend compiler techniques to target Persimmons.ai's custom accelerator hardware. You will help design and optimize the Persimmons Compiler mid- and backend, integrate it with custom operations and kernels, as well as implement compiler passes that convert higher-level intermediate representations into runtime-oriented code and libraries. This position offers the opportunity to directly shape Persimmons.ai’s innovative AI hardware and software stack through close collaboration with teams across hardware, systems, and software.

What you’ll do:

  • Develop and enhance MLIR-based compiler pipelines targeting Persimmons' custom spatial accelerator hardware.
  • Design and optimize the Persimmons Compiler mid- and backend techniques for efficient lowering, graph-to-resources mapping, and code generation.
  • Implement transformations to convert Python, PyTorch, and similar kernel representations to LLVM IR and runtime-ready libraries.
  • Architect and implement efficient support for SPMD-based, distributed collective operations and lower them through specialized MLIR compiler dialects (e.g., MESH, SHARDY).
  • Drive advanced loop optimizations leveraging polyhedral analysis: loop tiling, fusion, interchange, skewing, and related techniques.
  • Apply and optimize techniques such as bufferization, padding, inlining, and integration of custom operations and kernels within the compilation workflow.
  • Work on register allocation and instruction scheduling for Persimmons’ spatial hardware, ensuring high resource utilization, throughput, and low latency.
  • Contribute to graph and tensor partitioning logic for optimal hardware-targeted execution.
  • Collaborate across teams to deliver performant compilation flows from high-level ML representations to low-level executable artifacts.

Requirements

What You Bring To The Table:

  • We do not expect candidates to meet all of the requirements listed below; strong candidates will demonstrate expertise in several key areas.
  • Solid understanding and experience with underlying principles and methods of the MLIR framework (SSA representation, interfaces, rewriting, dialect hierarchy, etc.).
  • Hands-on experience with developing MLIR-based compiler infrastructure, algorithms, and techniques for non-GPU/custom spatial hardware architectures.
  • Working experience with lowering SIMD operations from PyTorch, Triton, xDSL, pyDSL, or similar Python-based frontends toward LLVM IR and, further, to SIMD kernel library.
  • Extensive experience and understanding of loop optimization based on polyhedral principles.
  • Experience and understanding of SPMD-based, distributed collective operations, specialized MLIR compiler dialects (e.g., MESH, SHARDY), and collective operation lowering in compilers for spatial hardware.
  • Experience with techniques such as padding, bufferization, inlining, and other lowering techniques.
  • Knowledge of register allocation and instruction scheduling in spatial architectures.
  • Experience in lowering and integration of custom operations and kernels at the compiler mid- and backend.
  • Familiarity with graph and tensor partitioning and mapping optimization algorithms and their integration in the compiler workflow.
  • High level of understanding and 5+ years of experience with C++ and appreciation for writing clean and maintainable code. Good knowledge of Python is a big plus.
  • Demonstrated fluency with modern AI tools and workflows (e.g., leveraging AI assistants for research, analysis, or productivity).

Benefits

  • Competitive salary and benefits package.
  • Flexible PTO
  • 401k

Please note: Our organization does not accept unsolicited candidate submissions from external recruiters or agencies. Any such submissions, regardless of form (including but not limited to email, direct messaging, or social media), shall be deemed voluntary and shall not create any express or implied obligation on the part of the organization to pay any fees, commissions, or other compensation. Direct contact of employees, officers, or board members regarding employment opportunities is strictly prohibited and will not receive a response.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
574,821 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$28k – $59k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Gurgaon
Python
C++
DevOps
Bitbucket
GitHub
Apply
$17k – $42k per year (Estimated) • In office • Tomsk
Python
C
C++
C
U-Boot
C++
STL
Management
Scrum
Apply
$16k – $41k per year (Estimated) • In office • Full-Time • Tomsk
Python
C++
Cython
C++
CMake
Protobuf
Qt
Cython
PyBind11
DevOps
gRPC
Git
Management
Scrum
QA
Pytest
Apply
$47k – $118k per year (Estimated) • In office • 7+ years exp • Master's Degree • Bengaluru
Python
Java
Scala
Databases
Apache Kafka
AI/ML
Cursor
Spark
OpenCV
Claude Code
Model Context Protocol
Scikit-learn
Multimodal AI
Computer Vision
AI Agents
NLP
spaCy
Flink
TensorFlow
Keras
PyTorch
BERT
ResNet
EfficientNet
NLTK
CLIP
DevOps
Amazon S3
Apply
$68k – $136k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Madison
Python
Java
MATLAB
DevOps
AWS
Apply
$103k – $244k per year (Estimated) • In office • Full-Time • San Jose
Apply
$111k – $224k per year (Estimated) • In office • Full-Time • Bachelor's Degree • San Jose
Apply
$111k – $262k per year (Estimated) • In office • Full-Time • San Jose
Verilog
SystemVerilog
AI/ML
AI Agents
Chips/EDA
Formal Verification
Apply
$139k – $311k per year (Estimated) • In office • Full-Time • San Jose
Python
C++
C++
TensorFlow C++
PyTorch C++
LLVM
AI/ML
JAX
Diffusion Models
TensorFlow
PyTorch
MLIR
Apache TVM
XLA
IREE
Apply
$56k – $93k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • San Jose
Apply
$163k – $205k per year • Equity • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Apply
$243k – $364k per year • Equity • Remote/Hybrid • Full-Time • 15+ years exp • San Jose
Apply
ASIC Engineer 1 day ago
$142k – $270k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Python
Java
C++
Apply
$163k – $205k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
Apply
See all jobs
This is one of many
574,821 more open roles from verified company boards, updated every day.