368,941open jobs
9,452companies
47,951added this week
Browse all
Salary
$136k – $298k per year (Estimated)
Location
Remote (United States, Canada, Hong Kong, Singapore)
Employment
Full-Time
Overview
Company
Impact
Profile match

Location: Remote (Global)

Type: Full-time

Company: Yotta Labs

Apply: [email protected]

About Yotta Labs

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. We enable training and inference across NVIDIA GPUs, AMD GPUs, and AWS Trainium, helping AI companies achieve the best performance and economics across heterogeneous hardware. Our mission is to provide high-performance AI computing and Model API services, enabling AI companies, research labs, and enterprises to train, deploy and integrate cutting-edge models at scale.

Role Overview

We are seeking a highly motivated AI Systems Research Engineer specializing in Trainium, GPU kernels, and LLM systems optimization. You will work at the intersection of AI Systems, Compiler and Runtime Optimization, Distributed Training & Inference, GPU/Accelerator Kernel Development, and Large Language Model Infrastructure. Your work will directly impact the scalability and performance of AI applications deployed on our platform.

Responsibilities

  • Design and implement high-performance kernels for Attention, MoE, GEMM, collective communication, and quantization.

  • Optimize kernels for NVIDIA, AMD, and AWS Trainium.

  • Develop custom operators and graph optimizations using Neuron SDK, PyTorch/XLA, Torch Dynamo, and Neuron Compiler.

  • Improve performance of vLLM, SGLang, TensorRT-LLM, and custom inference runtimes.

  • Design scalable distributed training and inference solutions across thousands of accelerators.

  • Contribute to open-source projects, publish technical findings and engage with the developer community.

Qualifications

  • Proficiency in AI programming languages such as Python and C++.

  • Deep understanding of GPU architecture and performance optimization.

  • Experience with CUDA, Triton, ROCm/HIP, or AWS Neuron.

  • Strong understanding of AI frameworks (e.g., PyTorch, Dynamo, LMCache), model architectures and profiling tools (e.g. Nsight, ROCm Profiler, or Neuron Profiler).

  • Strong problem-solving skills and the ability to work in a collaborative, remote environment.

  • A background in computer science, engineering, or a related field is preferred.

Preferred Experience

  • Contributions to open-source AI infra projects like vLLM, SGLang, PyTorch, or Triton.

  • Experience with with FlashAttention, PagedAttention, MoE, RLHF, or distributed AI systems.

  • Publications in top-tier conferences like MLSys, OSDI, SOSP, NSDI, SC, HPCA, or ISCA

Why Join Yotta Labs?

  • Be part of a visionary team aiming to redefine AI infrastructure and influence the future of multi-silicon AI computing.

  • Work on cutting-edge technologies that solves frontier AI infrastructure problems.

  • Collaborate with experts from leading institutions and tech companies.

  • Competitive compensation with equity. Enjoy a flexible, remote work environment that values innovation and autonomy.

How to Apply

Interested candidates should apply directly or send their resume and a brief cover letter to [email protected]. Please include links to any relevant projects or contributions.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Hong Kong
$17k – $49k per year (Estimated) • Remote/Hybrid • Kaliningrad
Node JS
Python
JavaScript
Node JS
BullMQ
Fastify
Python
FastAPI
Flask
Databases
Apache Kafka
Chroma
Pinecone
PostgreSQL
Qdrant
RabbitMQ
Redis
AI/ML
Claude
Copilot
Cursor
LangChain
LlamaIndex
LLM
Model Context Protocol
Prompt Engineering
RAG
Anthropic
OpenAI Codex
Structured Outputs
Function Calling
Frontend
Next.js
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Gitflow
Rest API
WebSockets
GitHub
Management
Jira
Apply
$32k – $72k per year (Estimated) • In office • Full-Time • 6+ years exp • Bengaluru
C++
Java
Python
YARA
Databases
Amazon Aurora
DevOps
Azure
GCP
Kubernetes
Cybersecurity
MITRE ATT&CK
Suricata
YARA
Zeek
Apply
$23k – $58k per year (Estimated) • In office • Full-Time • 10+ years exp • Pune
PowerShell
Python
SQL
Databases
MS SQL
Oracle
PostgreSQL
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Configuration Management
GCP
Grafana
Helm
Kubernetes
OpenShift
Prometheus
Apply
$23k – $58k per year (Estimated) • In office • Full-Time • 3+ years exp • Chennai • Bengaluru • Hyderabad
JavaScript
Python
AI/ML
Copilot
AI Agents
DevOps
CI/CD
Platform Engineering
Rest API
Terraform
Management
ServiceNow
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Gurgaon
Python
Visual Basic
AI/ML
AI Agents
Edge AI
Analytics
Power BI
Tableau
ETL/ELT
Apply
Remote • Internship • Bachelor's Degree • Hong Kong • Singapore
C++
Python
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
LLM
PyTorch
Quantization
SGLang
Triton
vLLM
AWS Trainium
ROCm
Mixture of Experts
DevOps
AWS
GitHub
Apply
$139k – $249k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Hong Kong • Singapore
Go
Python
AI/ML
CUDA Toolkit
Quantization
Ray
SGLang
vLLM
CUDA
Triton
AWS Trainium
Hugging Face
NCCL
NVLink
DevOps
AWS
Azure
Docker
GCP
Grafana
gRPC
Helm
kubectl
Kubernetes
Prometheus
Rancher
SLURM
GitHub
Apply
In office • Full-Time • 10+ years exp • Hong Kong
Go
DevOps
Platform Engineering
Web3
Smart Contracts
Apply
In office • 5+ years exp • Bachelor's Degree • Hong Kong
Apply
In office • Hong Kong
DevOps
AWS
Azure
GCP
Apply
In office • Hong Kong
Databases
Databricks
DevOps
AWS
Azure
GCP
Apply
In office • 3+ years exp • Hong Kong
Java
Node JS
Python
SQL
JavaScript
AI/ML
BERT
Fine-tuning
Hadoop
LangChain
LangGraph
LLM
RAG
Spark
Edge AI
Knowledge Graph
AI Agents
Google ADK
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Cybersecurity
GDPR
Apply
See all jobs
This is one of many
368,941 more open roles from verified company boards, updated every day.