696,071open jobs
40,735companies
104,724added this week
Browse all
Salary
$125k – $275k per year (Estimated)
Location
In office (Salt Lake City)
Employment
Full-Time
Overview
Company
Impact
Profile match
Blackrock Neurotech makes the implanted electrode arrays and signal processing hardware behind much of the published brain-computer interface research. Founded in 2008 in Salt Lake City, its systems have been used by participants for over a decade of continuous recording. The company works on decoders that restore movement, speech and touch.

Build the systems that expand human capability

At Blackrock Neurotech, we’ve spent decades making the impossible possible - helping people move, speak, and reconnect with the world when they otherwise could not.We’ve seen that restoring function restores more than ability. It restores independence, identity, and agency.

Today, we are building the next generation of human capability: brain-computer interfaces that are designed to be safe, scalable, and trusted in the real world. Our work is not only about reconnecting people to what was lost, but about expanding what is possible - creating a seamless interface between human intent and technology.

This is foundational work in a category-defining field. You will help build the infrastructure for a future where neural interfaces are invisible, reliable, and deeply human-centered.

Working at Blackrock Neurotech means:

  • Owning meaningful, high-impact problems at the frontier of science and engineering
  • Building alongside experienced, thoughtful peers across disciplines
  • Solving technically complex challenges grounded in real human outcomes
  • Contributing to a culture that values rigor, clarity, and long-term thinking over noise

The Role

The ML Training Performance Engineer will own the efficiency and scalability of training models on GPU and cloud infrastructure. You will turn available compute into faster, more capable experiments as our model training scales in complexity and compute requirements.

As a hands-on individual contributor on a small research team, you will work across the training stack, from Python and model execution to GPU kernels, distributed communication, and runtime environments. Partnering with model researchers, data engineers, and infrastructure and IT teams, you will identify and implement performance improvements while preserving numerical correctness and scientific intent.

You will have significant ownership over how we measure, optimize, and scale training performance, establishing the baselines, tooling, and technical approaches that will support our AI/ML work as it grows.

What You'll Do

  • Own training performance across single-GPU, multi-GPU, and multi-node workloads, establishing reproducible baselines for throughput, memory use, utilization, time to target quality, and cost
  • Profile the full training path to distinguish compute, memory, communication, CPU, and I/O bottlenecks and prioritize changes with measurable end-to-end impact
  • Optimize tensor layouts, precision, memory allocation, activation checkpointing, operator fusion, and execution graphs to fit larger or longer-context models within available resources
  • Write, tune, and validate custom GPU kernels using CUDA, Triton, HIP/ROCm, or the appropriate platform tools when existing implementations limit performance
  • Improve Python training scripts, framework and compiler settings, batching, gradient accumulation, and optimizer execution while preserving intended training behavior
  • Design and tune distributed training strategies, including data, tensor, pipeline, or sharded parallelism, based on model structure, memory limits, and interconnect topology
  • Partner with model researchers on hardware-aware architecture and hyperparameter changes, measuring their effects on convergence, model quality, and compute requirements
  • Coordinate with the neural data infrastructure engineer on prefetching, pinned memory, host-to-device transfer, and I/O overlap so data delivery keeps pace with training
  • Work with infrastructure and IT on GPU selection, cloud instance configurations, networking, drivers, containers, scheduling, and capacity planning as training needs and compute capacity scale
  • Build robust checkpoint, restart, and recovery workflows and performance regression checks that keep long-running experiments reproducible and productive
  • Communicate benchmark evidence, numerical tradeoffs, scaling limits, and resource recommendations clearly to researchers and organizational stakeholders

What You Bring

  • Demonstrated experience improving the performance of substantial deep learning training workloads, with measured gains in speed, memory efficiency, or compute cost
  • A measurement-driven approach to performance optimization, using profiling and benchmarking to validate meaningful end-to-end improvements
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience
  • Exceptional programming ability in Python and C++ or a comparable systems language, with strong debugging, testing, and performance analysis practices
  • Deep understanding of GPU execution, including memory hierarchies, memory coalescing, thread blocks, warps or wavefronts, occupancy, synchronization, and bandwidth limits
  • Hands-on experience developing and profiling GPU kernels with CUDA or HIP/ROCm, and the ability to diagnose correctness and performance at the hardware level
  • Deep knowledge of PyTorch or an equivalent framework, including automatic differentiation, computation graphs, tensor storage, compilation, and mixed-precision training
  • Strong understanding of deep learning architectures and the underlying computations that drive training performance
  • Experience with distributed training, collective communication, sharding, and the interaction between model partitioning and GPU interconnects
  • Strong understanding of numerical stability and the ability to validate gradients, convergence, and model quality after performance changes
  • Experience configuring and diagnosing Linux-based GPU environments, containers, cloud compute, and high-throughput storage or networking
  • Ability to collaborate closely with researchers and infrastructure teams and make clear tradeoffs between implementation effort, performance, reliability, and scientific value
  • Experience with Triton, compiler optimization, advanced GPU profiling tools, multiple accelerator generations, long-sequence or multimodal models, neural time series, or large-scale model training is a plus

Working Location:

This is an on-site role based at Blackrock Neurotech's headquarters in Salt Lake City, Utah. Occasional travel may be required.

How We Work

We are a small, experienced team working on consequential problems.

  • We take ownership of outcomes and follow through with clarity and accountability
  • We prioritize sustained, high-quality work over performative urgency
  • We value rigor, sound judgement and thoughtful decision-making
  • We collaborate deliberately: low ego, high trust and high context

This is a high-ownership role, but it is not an "always-on" one. We expect strong work and our people to have a life outside of it.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
696,071 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Salt Lake City
$92k – $138k per year • In office • Full-Time • 3+ years exp • Berlin
Python
JavaScript
TypeScript
SQL
Node JS
Node JS
Fastify
Prisma
AI/ML
Prefect
Multimodal AI
AI Agents
Cohere SDK
Langfuse
Gemini
LLM
OpenAI
OCR
Semantic Search
DevOps
Datadog
Azure
AWS
Amazon S3
Analytics
ETL/ELT
Apply
$39k – $112k per year (Estimated) • Remote • Contractor • Master's Degree
Python
AI/ML
SciPy
NumPy
Gemini
LLM
Apply
$39k – $112k per year (Estimated) • Remote • Contractor • Master's Degree
Python
AI/ML
SciPy
NumPy
Gemini
LLM
Apply
$145k – $218k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • Mississauga
Python
Java
Kotlin
SQL
Databases
Weaviate
Pinecone
Apache Kafka
AI/ML
Copilot
AutoGen
LangChain
Fine-tuning
Prompt Engineering
AI Agents
Flink
LLM
RAG
Anomaly Detection
Feature Store
DevOps
CI/CD
AWS
Docker
Kubernetes
Trunk-Based Development
Chaos Engineering
Progressive Delivery
AIOps
SLI/SLO/SLA
Cybersecurity
Threat Modeling
Management
Agile
Apply
$97k – $206k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Singapore
Python
SQL
AI/ML
Pandas
Analytics
Alteryx
Microsoft Excel
Apply
$101k – $222k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Salt Lake City
Python
AI/ML
PyTorch
Machine Learning
Apply
$80k – $175k per year (Estimated) • In office • Part-Time • Salt Lake City
Apply
$150k – $278k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Salt Lake City
Python
AI/ML
Machine Learning
DevOps
CI/CD
Linux
Apply
$36k – $63k per year (Estimated) • In office • Full-Time • High School Diploma • Salt Lake City
Management
Microsoft Office
Apply
$150k – $200k per year • Remote • Bachelor's Degree • Salt Lake City
Databases
Snowflake
Google BigQuery
BigQuery
AI/ML
AI Agents
Agentic Workflows
Management
Smartsheet
Scrum
Kanban
Microsoft Office
Apply
$289k – $352k per year • Equity • In office • Salt Lake City
JavaScript
Frontend
React.js
Mobile
React Native
Apply
$101k – $222k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Salt Lake City
Python
AI/ML
PyTorch
Machine Learning
Apply
up to $100k per year • In office • Full-Time • 1+ year exp • Salt Lake City
Apply
$82k – $157k per year (Estimated) • Equity • Remote • 5+ years exp • Bachelor's Degree • Salt Lake City
AI/ML
Multimodal AI
Apply
See all jobs
This is one of many
696,071 more open roles from verified company boards, updated every day.