665,268open jobs
38,876companies
100,271added this week
Browse all
Salary
$130k per year
Location
Remote (Argentina, Brazil, Colombia, Mexico, Spain, Chile, Ecuador, Portugal, Uruguay)
Seniority
Middle · 3+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match
Anyone AI is an AI talent company that trains software developers, mainly from Latin America, for careers in artificial intelligence and supplies experts who create coding and STEM training data for frontier AI models. Its founders were members of the founding team of Deep Vision AI, and it is venture-backed by investors including Global Founders Capital, Canvas Ventures and Latitud. Its openings are remote, project-based contract roles for full-stack and Python developers and physics experts who design and review tasks for leading AI labs, recruited in countries across Europe and Asia.

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at least two of the following:

    • CUDA

    • Triton

    • NKI / AWS Neuron

    • Pallas / JAX

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

    • Memory bandwidth

    • Compute throughput

    • GPU occupancy

    • Shared memory

    • Register pressure

    • Memory coalescing

    • Bank conflicts

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

Nice to Have

  • Experience across both NVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote

Engagement: Part-time, project-based consulting

Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
665,268 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$193k – $339k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • United States
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Apache Iceberg
Delta Lake
Amazon Redshift
AI/ML
LangChain
Spark
Vertex AI
Fine-tuning
Prompt Engineering
AWS Bedrock
LLM
RAG
OpenAI
Amazon SageMaker
Knowledge Graph
Agentic Workflows
DevOps
Azure
AWS
Vector
Analytics
ETL/ELT
AWS Glue
Apply
$21k – $47k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
AI/ML
Copilot
Claude
Frontend
Angular
Storybook
Sass
DevOps
Rest API
CI/CD
Git
AWS
Apply
$121k – $257k per year (Estimated) • Remote/Hybrid • Full-Time • 13+ years exp • Master's Degree • Toronto
DevOps
GCP
Azure
AWS
Configuration Management
Management
Agile
Apply
$230k – $262k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • McLean
Python
AI/ML
LangChain
Spark
AI Agents
Ray
Agentic Workflows
DevOps
AWS
Apply
$45k – $85k per year (Estimated) • In office • Full-Time • 5+ years exp • High School Diploma • Tulsa
DevOps
AWS
Management
Outlook
Apply
$130k per year • Remote • Contractor • 3+ years exp
Python
SQL
AI/ML
RLHF
DevOps
Docker Compose
CI/CD
Docker
Cybersecurity
CVE
CWE
CVSS
pwntools
Apply
$130k per year • Remote • Contractor • 3+ years exp
AI/ML
RLHF
Synthetic Data
Apply
$130k per year • Remote • Contractor • 2+ years exp
AI/ML
CUDA Toolkit
RLHF
CUDA
Triton
AWS Trainium
XLA
DevOps
AWS
Apply
$90k – $160k per year • Remote • Contractor • 2+ years exp
Python
Go
JavaScript
TypeScript
C#
Mobile
JUnit
QA
Jest
Pytest
Apply
$90k – $160k per year • Remote • Contractor • 2+ years exp
Python
Go
JavaScript
TypeScript
C#
Mobile
JUnit
QA
Jest
Pytest
Apply
See all jobs
This is one of many
665,268 more open roles from verified company boards, updated every day.