368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$150k – $250k per year
Location
Remote (United States)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Positron AI builds purpose-built inference servers for large language models, aiming for much higher energy efficiency than general-purpose GPU systems. Its Atlas hardware uses field-programmable logic and a memory-centric design to serve transformer models at high token throughput. The company was founded in 2023 and assembles its systems in the United States.

About Positron AI

Positron AI specializes in developing custom hardware systems to accelerate AI inference. These inference systems offer significant performance and efficiency gains over traditional GPU-based systems, delivering advantages in both performance per dollar and performance per watt. Positron exists to create the world's best AI inference systems.

Role Overview

Senior Software Engineer - Machine Learning Systems & High-Performance LLM Inference

We are seeking a Senior Software Engineer to contribute to the development of high-performance software that powers execution of open-source large language models (LLMs) on our custom appliance. This appliance leverages a combination of FPGAs and x86 CPUs to accelerate transformer-based models. The software stack is written primarily in modern C++ (C++17/20) and heavily relies on templates, SIMD optimizations, and efficient parallel computing techniques.

Key Responsibilities

  • Design and implement high-performance inference software for LLMs on custom hardware.
  • Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
  • Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
  • Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
  • Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
  • Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads.
  • Implement ML system-level optimizations such as token streaming, KV cache optimizations, and efficient batching for transformer execution.
  • Collaborate with ML researchers and software engineers to integrate model quantization techniques, sparsity optimizations, and mixed-precision execution.
  • Ensure all code contributions include unit, performance, acceptance, and regression tests as part of a continuous integration-based development process.

Required Qualifications

  • 7+ years of professional experience in C++ software development, with a focus on performance-critical applications.
  • Strong understanding of C++ templates and modern memory management.
  • Hands-on experience with SIMD programming (AVX-512, SSE, or equivalent) and intrinsics-based vectorization.
  • Experience in high-performance computing (HPC), numerical computing, or ML inference optimization.
  • Experience with ML model execution optimizations, including efficient tensor computations and memory access patterns.
  • Knowledge of multi-threading, NUMA architectures, and low-level CPU optimization.
  • Proficiency with systems-level software development, profiling tools (perfetto, VTune, Valgrind), and benchmarking.
  • Experience working with hardware accelerators (FPGAs, GPUs, or custom ASICs) and designing efficient software-hardware interfaces.

Preferred Qualifications

  • Familiarity with LLVM/Clang or GCC compiler optimizations.
  • Experience in LLM quantization, sparsity optimizations, and mixed-precision computation.
  • Knowledge of distributed inference techniques and networking optimizations.
  • Understanding of graph partitioning and execution scheduling for large-scale ML models.

Leveling & Scope

While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.

Why Join Us?

  • Work on a cutting-edge ML inference platform that redefines performance and efficiency for LLMs.
  • Tackle challenging low-level performance engineering problems in AI and HPC.
  • Collaborate with a team of hardware, software, and ML experts building an industry-first product.
  • Opportunity to contribute to and shape the future of open-source AI inference software.

Compensation and Benefits

The base salary range for this role is $150,000 - $250,000.

Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.

At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.

Benefits & Perks

We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.

Health and wellness

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available

Time off and flexibility

  • Unlimited paid time off - we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup

Compensation and future

  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one

Visa Support

This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.

Equal Opportunity Employer. If you’re excited about the role but don’t meet every bullet, we’d still love to hear from you.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$123k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas • Denver • Birmingham
Java
SQL
Java
Gradle
Hibernate
Maven
Spring Boot
Spring Framework
Databases
Apache Kafka
MySQL
Redis
DevOps
CI/CD
Dynatrace
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Apply
Platform Engineer 1 day ago
$87k – $140k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
$70k – $105k per year • In office • Full-Time • 3+ years exp
Python
SQL
TypeScript
AI/ML
LLM
RAG
Function Calling
LLM Guardrails
Cybersecurity
GDPR
Management
n8n
Apply
$93k – $140k per year • In office • Full-Time • 3+ years exp
Node JS
TypeScript
JavaScript
Node JS
Nest.JS
AI/ML
LLM
EU AI Act
Frontend
React.js
Cybersecurity
GDPR
Apply
Founding Engineer 1 day ago
$93k – $140k per year • In office • Full-Time • 3+ years exp • Munich
Python
Python
FastAPI
AI/ML
Fine-tuning
LLM
VLM
Apply
$200k – $350k per year • Remote • Full-Time • 10+ years exp
AI/ML
Quantization
SGLang
TPOT
vLLM
AI Agents
Mixture of Experts
Apply
$200k – $350k per year • In office • Full-Time • 12+ years exp • Austin
SystemVerilog
DevOps
CI/CD
HPC
Chips/EDA
Cadence Palladium
Formal Verification
Siemens Veloce
Synopsys ZeBu
UVM
Apply
$200k – $300k per year • Remote • Full-Time • 12+ years exp
Cybersecurity
Threat Modeling
Apply
$150k – $300k per year • Remote • Full-Time • 12+ years exp
Chips/EDA
Ansys RedHawk
Synopsys PrimeTime
Apply
$200k – $300k per year • Remote • Full-Time • 12+ years exp
AI/ML
LLM
DevOps
Vector
Chips/EDA
Synopsys PrimePower
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.