368,634open jobs
9,437companies
50,578added this week
Browse all
Location
Remote (United States, Canada, Hong Kong, Singapore)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match

Location: Remote (Global)

Type: Internship

Company: Yotta Labs

Apply: [email protected]

About Yotta Labs

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware. Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.

Role Overview

We are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization. Over a 12-16 week internship, you will own a well-scoped project at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, GPU/Accelerator Kernel Development, and Large Language Model Infrastructure - taking it from design to working, profiled code running on real hardware. Your work will ship to production or open source and directly impact the performance of AI applications deployed on our platform. Strong interns receive return offers for full-time roles.

Responsibilities

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium.

  • Build custom operators using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.

  • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes - kernel fusion, scheduling, KV-cache and memory optimizations.

  • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.

  • Ship code upstream to open-source AI infrastructure projects, with tests and documentation.

Qualifications

  • Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field.

  • Solid programming skills in Python and familiarity with C++.

  • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects.

  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels - class projects and personal projects count.

  • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler).

  • Strong problem-solving skills and the ability to work independently in a collaborative, remote environment.

Preferred Experience

  • Contributions to open-source AI infra projects like vLLM, SGLang, PyTorch, or Triton.

  • Familiarity with LLM inference internals - FlashAttention, PagedAttention, continuous batching, speculative decoding, MoE, or quantization.

  • Experience with profiling tools (e.g. Nsight, ROCm Profiler, Neuron Profiler, or PyTorch Profiler) and performance debugging on real workloads.

  • Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA

Why Join Yotta Labs?

  • Be part of a visionary team aiming to redefine AI infrastructure and influence the future of multi-silicon AI computing.

  • Work on frontier AI infrastructure problems with access to serious hardware - latest-generation NVIDIA GPUs, AMD accelerators, and AWS Trainium at scale.

  • Get direct mentorship from engineers from leading institutions and tech companies.

  • Competitive internship compensation, a flexible remote work environment, and a fast path to a full-time return offer for top performers.

How to Apply

Interested candidates should apply directly or send their resume to [email protected]. Please include links to any relevant projects or contributions (GitHub, open-source PRs, course projects) - for internships, these matter more to us than a cover letter.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Hong Kong
Staff Engineer 1 day ago
$105k – $233k per year • In office • Full-Time • 8+ years exp • Munich
Python
TypeScript
JavaScript
AI/ML
LLM
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Terraform
Apply
Platform Engineer 1 day ago
$87k – $140k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
$70k – $105k per year • In office • Full-Time • 3+ years exp
Python
SQL
TypeScript
AI/ML
LLM
RAG
Function Calling
LLM Guardrails
Cybersecurity
GDPR
Management
n8n
Apply
$93k – $140k per year • In office • Full-Time • 3+ years exp
Node JS
TypeScript
JavaScript
Node JS
Nest.JS
AI/ML
LLM
EU AI Act
Frontend
React.js
Cybersecurity
GDPR
Apply
Founding Engineer 1 day ago
$93k – $140k per year • In office • Full-Time • 3+ years exp • Munich
Python
Python
FastAPI
AI/ML
Fine-tuning
LLM
VLM
Apply
$136k – $298k per year (Estimated) • Remote • Full-Time • Hong Kong • Singapore
C++
Python
C++
PyTorch C++
AI/ML
CUDA Toolkit
LLM
PyTorch
Quantization
RLHF
SGLang
TensorRT
TensorRT-LLM
vLLM
AWS Trainium
ROCm
Mixture of Experts
DevOps
AWS
Apply
$139k – $249k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Hong Kong • Singapore
Go
Python
AI/ML
CUDA Toolkit
Quantization
Ray
SGLang
vLLM
CUDA
Triton
AWS Trainium
Hugging Face
NCCL
NVLink
DevOps
AWS
Azure
Docker
GCP
Grafana
gRPC
Helm
kubectl
Kubernetes
Prometheus
Rancher
SLURM
GitHub
Apply
In office • 5+ years exp • Bachelor's Degree • Hong Kong
Apply
In office • Hong Kong
DevOps
AWS
Azure
GCP
Apply
In office • Hong Kong
Databases
Databricks
DevOps
AWS
Azure
GCP
Apply
In office • 3+ years exp • Hong Kong
Java
Node JS
Python
SQL
JavaScript
AI/ML
BERT
Fine-tuning
Hadoop
LangChain
LangGraph
LLM
RAG
Spark
Edge AI
Knowledge Graph
AI Agents
Google ADK
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
In office • 3+ years exp • Hong Kong
Java
Node JS
Python
SQL
JavaScript
AI/ML
AI Agents
BERT
Edge AI
Fine-tuning
Google ADK
Hadoop
Knowledge Graph
LangChain
LangGraph
LLM
RAG
Spark
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.