368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$155k – $300k per year (Estimated)
Location
Remote/Hybrid (Hillsboro, Santa Clara, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Intel is an American semiconductor company founded in 1968 by Robert Noyce and Gordon Moore and headquartered in Santa Clara, California. It created the x86 instruction set that still underpins most personal computers and servers, and designs and manufactures processors, chipsets, discrete graphics, networking silicon and AI accelerators. Unlike most of its competitors the company owns its fabrication plants, and it is investing heavily in Intel Foundry to manufacture chips for external customers while rebuilding its process technology leadership.

Job Details:

Job Description:

About the Role

The Software and AI (SAI) organization is seeking a highly skilled software engineer to contribute to the development and low-level optimization of oneDNN, a complex, cross-platform, open-source performance library that serves as the foundation for deep learning applications (github.com/uxlfoundation/oneDNN).

Please Note: This is a low-level software engineering and hardware-acceleration role. It does not involve building, training, or tuning machine learning models. Instead, you will focus on developing highly optimized math primitives, parallel algorithms, and GPU kernels that power industry-leading AI frameworks (such as OpenVINO, TensorFlow, PyTorch, and ONNX Runtime) on Intel hardware.

Key Responsibilities

Kernel Development and Architecture

  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads
  • Design scalable JIT and codegen infrastructure for GPU kernel generation

Low-Level Optimization

  • Implement fusion and memory-traffic optimizations to maximize hardware utilization
  • Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)

Performance Modeling and Profiling

  • Build analytical and empirical performance models for kernel dispatch and tuning
  • Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths

Hardware and Software Co-Design

  • Co-design GPU primitives and kernel architectures for next-generation Intel GPUs
  • Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks

Infrastructure and Validation

  • Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads

Why Join Us

Massive Scale

  • Work on a global, high-impact open-source library that scales AI performance across millions of devices worldwide

Cutting-Edge Hardware

  • Get early access to and influence the software stack for Intel's roadmap of next-generation discrete GPUs

Expert Collaboration

  • Work alongside industry-leading experts in GPU compilers, hardware architecture, and performance libraries

Total Rewards

  • Enjoy a competitive package including stock programs, quarterly bonuses, robust healthcare, and highly flexible hybrid/remote working options

What We're Looking For

To be successful in this role, you should demonstrate the following professional traits:

  • A strong ownership mindset - you take initiative on complex, ambiguous technical problems and drive them to resolution
  • A collaborative approach - you work effectively across hardware, compiler, and framework teams to align on shared technical goals
  • A performance-driven curiosity - you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems

Qualifications:

Minimum Qualifications

  • Education: BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field
  • Core Language: 5+ years of professional software development experience with expert-level modern C++
  • Performance Optimizations: 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs
  • Hardware Architecture: Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization)

Preferred Qualifications

  • Math Libraries:Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)
  • Low-Level Tuning:Hands-on experience with GPU assembly-level tuning or compiler optimization
  • Parallel APIs:Familiarity with parallel programming APIs such as OpenMP or oneTBB
  • AI Workload Context:Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks

Job Type:

Experienced Hire

Shift:

Shift 1 (United States of America)

Primary Location:

US, Oregon, Hillsboro

Additional Locations:

US, California, Santa Clara

Posting Statement:

All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.

Position of Trust

N/A

Benefits

We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.

Annual Salary Range for jobs which could be performed in the US: $195,200.00-275,580.00 USDThe range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.

Work Model for this Role

This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.

*

ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Hillsboro
Data Scientist 1 day ago
$25k – $56k per year (Estimated) • Remote • Bachelor's Degree • Moscow
C++
Python
SQL
AI/ML
Computer Vision
CUDA
CUDA Toolkit
TensorRT
DevOps
Docker
Git
Kubernetes
Apply
$55k – $157k per year (Estimated) • Remote • Full-Time • Sydney
C++
Go
Lua
Python
C++
CMake
Databases
ActiveMQ
Aerospike
Apache Kafka
Cassandra
DevOps
Docker
gRPC
Kubernetes
Apply
$39k per year (gross) • In office • Full-Time • Voronezh
C++
Chips/EDA
Altium Designer
Apply
$39k per year (gross) • In office • Full-Time • Nizhny Novgorod
C++
Chips/EDA
Altium Designer
Apply
$39k per year (gross) • In office • Full-Time • Tomsk
C++
Chips/EDA
Altium Designer
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Malaysia
Apply
$30k – $72k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru
Perl
Python
SystemVerilog
Chips/EDA
Formal Verification
UVM
Apply
In office • Internship • Bachelor's Degree • Kulim
C#
Python
Apply
$94k – $212k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Chandler
JavaScript
SQL
Databases
MS SQL
Apply
$176k – $334k per year (Estimated) • Remote/Hybrid • Full-Time • 14+ years exp • Master's Degree • Hillsboro • Santa Clara • Phoenix
Apply
RTL Design Engineer 1 hour ago
$87k – $168k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • Phoenix • Hillsboro
Assembly
Perl
Python
SystemVerilog
Verilog
VHDL
Chips/EDA
Synopsys Design Compiler
Apply
Business Analyst 4 days ago
$118k – $126k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Sunnyvale • Austin • Hillsboro
SQL
Analytics
Power BI
Tableau
Apply
$176k – $334k per year (Estimated) • Remote/Hybrid • Full-Time • 14+ years exp • Master's Degree • Hillsboro • Santa Clara • Phoenix
Apply
$142k – $255k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Master's Degree • Hillsboro • Santa Clara • Austin • Phoenix
Apply
$114k – $229k per year (Estimated) • In office • Contractor • 4+ years exp • Bachelor's Degree • Hillsboro • Folsom • Phoenix
C#
SQL
Databases
PostgreSQL
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.