1,321,836open jobs
77,909companies
210,930added this week
Browse all
Salary
≈ $224k – $427k per year (Estimated)
Location
In office (San Jose)
Seniority
Senior · 3+ years exp
Visa
H-1B filings in 12 months: 747 · for this role: 506 · green card filings: 451
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 7, 2026. First seen by Alion on Oct 1, 2026. ByteDance scores A on the Alion truth index.

Overview
Company
Impact
Profile match
ByteDance is a Chinese internet technology company founded in Beijing in 2012 that operates some of the world's largest content and commerce platforms. Its portfolio includes the short-video apps TikTok and Douyin, the news aggregator Toutiao, the video editor CapCut, the workplace suite Lark and the Doubao family of AI assistants and models. Recommendation systems trained on user behaviour sit at the centre of every product, and the company has grown into one of the highest-revenue private technology businesses in the world.

岗位职责 / Responsibilities

About the Team

We are a systems software team building the foundational software for large-scale GPU computing platforms. We work at the hardware/software boundary across the Linux kernel, accelerators, storage, firmware, and platform validation. We value rigorous engineering, clear interfaces, measurable performance and reliability, and upstream collaboration where appropriate. The team partners closely with hardware, architecture, product, validation, and production engineering groups to move new capabilities from design through dependable deployment.

About the Role

You will build and optimize low-level software that makes modern GPU accelerators usable, observable, and reliable in production. Your work may span kernel drivers, runtime components, firmware interfaces, resource management, telemetry, debugging tools, and performance-critical paths. You will own technically difficult features end to end and collaborate across the stack while remaining primarily accountable for code, validation, and production outcomes.

Responsibilities

- Design, implement, and maintain production GPU system software, including kernel-driver, runtime, firmware-interface, host-management components, debugging tracing and profiling tools, and GPU system performance measurement tools

- Own accelerator features from proof of concept and architecture through implementation, pre-silicon or emulation validation, bring-up, qualification, and deployment.

- Debug complex failures involving GPUs, CPUs, memory, PCIe or interconnects, IOMMU, firmware, operating systems, runtimes, and distributed workloads.

- Profile workloads and remove bottlenecks in initialization, memory movement, scheduling, synchronization, communication, recovery, and device utilization.

- Build automated tests, telemetry, dashboards, health checks, and diagnostic tools that make failures reproducible and regressions visible.

- Partner with silicon, firmware, compiler, library, machine-learning framework, platform, validation, and production teams to deliver compatible system behavior.

- Improve resilience through error detection, isolation, retry, reset, repair, graceful degradation, and clear operational procedures.

- Review low-level designs and code, document hardware/software contracts, and share practical performance and debugging methods with peers.

任职要求 / Requirements

Minimum Qualifications

- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.

- 3+ years of professional programming experience in C or C++ for low-level, embedded, kernel, driver, runtime, or performance-critical software.

- Solid understanding of computer architecture, operating systems, concurrency, memory hierarchies, DMA, interrupts, and device I/O.

- Hands-on experience with GPU or accelerator software in at least one layer, such as kernel drivers, runtimes, firmware, libraries, collective communication, or performance tooling.

- Demonstrated ability to diagnose system failures using traces, logs, profilers, debuggers, counters, and controlled experiments.

- Experience delivering and maintaining production-quality features across hardware and software teams.

Preferred Qualifications

- Experience with CUDA, ROCm, Level Zero, OpenCL, Triton, CUTLASS, or comparable accelerator programming and runtime environments.

- Knowledge of GPU scheduling, memory management, virtualization, PCIe, coherent interconnects, NUMA, or multi-GPU topology.

- Experience with pre-silicon development, board bring-up, firmware communication, baseboard management controllers, or fleet qualification.

- Familiarity with distributed training or inference and topology-aware communication libraries; this is preferred rather than a universal requirement.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,321,836 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Jose
$176k – $308k per year • Equity • In office • Full-Time • 8+ years exp • Santa Clara
Python
Go
Java
AI/ML
Cursor
Windsurf
Claude Code
Prompt Engineering
LLM
DevOps
Kubernetes
Management
ServiceNow
Apply
$201k – $352k per year • Equity • In office • Full-Time • 12+ years exp • Santa Clara
Python
Go
Java
AI/ML
Cursor
Windsurf
Claude Code
Prompt Engineering
LLM
DevOps
Kubernetes
Management
ServiceNow
Apply
$135k – $160k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Chicago
Python
SQL
Databases
Databricks
Azure SQL Database
AI/ML
Scikit-learn
AI Agents
TensorFlow
PyTorch
RAG
Hugging Face
Machine Learning
Frontend
GraphQL
DevOps
Rest API
Azure
CI/CD
Apply
$116k – $174k per year • In office • Full-Time • 8+ years exp • Columbus
Apply
$183k – $246k per year • Remote (United States) • Full-Time • 7+ years exp • Austin • New York
Python
AI/ML
AI Agents
NLP
LLM
Knowledge Graph
Recommender Systems
Multi-Agent Systems
Machine Learning
DevOps
Kubernetes
Apply
$200k – $400k per year • In office • Full-Time • San Jose
Python
AI/ML
Multimodal AI
Computer Vision
Midjourney
PyTorch
OpenAI
World Models
Embodied AI
Machine Learning
Apply
$195k – $329k per year • Equity • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • San Jose
AI/ML
Model Context Protocol
AI Agents
Management
Agile
Apply
≈ $105k – $248k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • San Jose • Kansas City
Python
AI/ML
AI Agents
Edge AI
DevOps
Splunk
GCP
AWS
Kubernetes
Cybersecurity
Crowdstrike
Microsoft Sentinel
SentinelOne
Microsoft Defender
Google SecOps
SIEM
Apply
≈ $84k – $161k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • San Jose
AI/ML
LLM
DevOps
GCP
Apply
$187k – $262k per year • In office • Full-Time • 8+ years exp • London • San Jose • Washington
Analytics
A/B Testing
Apply
See all jobs
This is one of many
1,321,836 more open roles from verified company boards, updated every day.