694,856open jobs
40,803companies
104,316added this week
Browse all
Salary
$116k – $140k per year
Location
In office (San Francisco)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

About The Role  

The Inference Research team is dedicated to building the next generation of efficient, scalable, and reliable serving systems for large foundation models, directly contributing to the mission of advancing open and transparent AI. Our work operates at the critical intersection of cutting-edge model architectures, high-performance systems engineering, and deep hardware optimization. We focus on co-designing software, algorithms, and models to significantly lower the cost and latency of modern AI systems.

As a research intern, you will dive into the complexities of distributed inference, compiler-aware optimization, and novel inference-time computation strategies (such as speculative decoding and phase-aware execution). You will be tasked with co-designing and implementing cross-layer optimizations across models, systems, and hardware, with a focus on areas like KV cache design and large-scale serving architectures.

Projects aim to unlock unprecedented performance and scale for foundation models, enabling faster serving, larger model deployment (e.g., Mixture-of-Experts), and robust, reproducible evaluation under realistic serving workloads.

Responsibilities

  • Design and conduct rigorous experiments to validate hypotheses
  • Communicate the plans, progress, and results of projects to the broader team
  • Document findings in scientific publications and blog posts

Requirements

  • Currently pursuing a final year of Bachelor's, Master's, or Ph.D. degree in Computer Science, Electrical Engineering, or a related field
  • Strong knowledge of Machine Learning and Deep Learning fundamentals
  • Experience with deep learning frameworks (PyTorch, JAX, etc.)
  • Strong programming skills in Python
  • Familiarity with Transformer architectures and recent developments in foundation models

Preferred Qualifications

  • Prior research experience in foundation models, efficient machine learning, or ML systems.
  • Publications at leading conferences in machine learning or systems (i.e., MLSys, ICLR).
  • Experience with CUDA programming (for kernel development)
  • Understanding of model optimization techniques and hardware acceleration approaches
  • Contributions to open-source machine learning projects

About Together AI

Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.

Internship Program Details

Our internship program runs 12 to 14 weeks, giving you the opportunity to work alongside industry-leading engineers and researchers across multiple teams. This cohort's internship dates span January 4th to April 9th.

Compensation

We offer competitive compensation, housing stipends, and other competitive benefits. The estimated US hourly rate for this role is $58 to $70 an hour. Our hourly rates are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
694,856 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
AI Engineer 6 hours ago
$22k – $56k per year (Estimated) • In office • Full-Time • 4+ years exp
Python
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
FAISS
Apache Kafka
AI/ML
Copilot
Cursor
LangChain
Ray Serve
Qwen
Spark
Claude Code
LlamaIndex
LoRA
vLLM
Fine-tuning
Embeddings
Quantization
Knowledge Distillation
Computer Vision
NLP
AWQ
GPTQ
ONNX
PEFT
QLoRA
TGI
Llama
Mistral
TensorFlow
PyTorch
LLM
RAG
Ray
Hallucination
Triton
OpenAI
Anthropic
Feature Store
LLM Guardrails
Recommender Systems
Agentic Workflows
Model Distillation
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$34k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
Python
TypeScript
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
RAG
LLM Guardrails
DevOps
Azure
CI/CD
Platform Engineering
Management
Agile
Apply
AI Engineer 5 hours ago
$62k – $181k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Toronto
Python
AI/ML
LLM
DevOps
Git
Apply
DevOps Engineer 6 hours ago
$37k – $39k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree
Python
JavaScript
Node JS
Bash
Databases
PostgreSQL
TimescaleDB
DevOps
Terraform
Ansible
GCP
GitLab CI
Azure
CI/CD
Jenkins
AWS
Hetzner
Linux
Apply
Full Stack Engineer 6 hours ago
$22k – $55k per year (Estimated) • In office • Full-Time • 5+ years exp
Python
JavaScript
TypeScript
Python
FastAPI
Django
Databases
MySQL
PostgreSQL
AI/ML
Copilot
Cursor
Claude Code
LLM
RAG
Frontend
React.js
DevOps
Rest API
Terraform
GCP
Datadog
Azure
CI/CD
AWS
Docker
Grafana
QA
Sentry
Apply
$116k – $140k per year • In office • Internship • PhD • San Francisco
Python
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
JAX
AI Agents
NLP
Transformers
PyTorch
Together AI
Post-training
Pre-training
Machine Learning
Apply
$200k – $250k per year • Remote • Full-Time • 5+ years exp • San Francisco
Python
SQL
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
InfiniBand
DevOps
HPC
Apply
$116k – $140k per year • In office • Internship • Bachelor's Degree • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
JAX
NLP
Transformers
PyTorch
Together AI
Post-training
Pre-training
Machine Learning
Apply
$116k – $140k per year • In office • Internship • Bachelor's Degree • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
JAX
NLP
Transformers
PyTorch
Together AI
Post-training
Pre-training
Machine Learning
Apply
$116k – $140k per year • In office • Internship • San Francisco
AI/ML
Cursor
CUDA Toolkit
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
CUDA
Triton
Pre-training
Apply
Account Executive 1 hour ago
$150k – $250k per year • In office • Full-Time • 3+ years exp • San Francisco
Marketing
LinkedIn
Apply
$100k – $120k per year • In office • Full-Time • 3+ years exp • San Francisco
Marketing
LinkedIn
Apply
$55k – $120k per year • In office • Full-Time • San Francisco
Marketing
LinkedIn
Apply
$100k – $200k per year • Equity 0.1–2% • In office • Full-Time • San Francisco
Python
JavaScript
TypeScript
AI/ML
LLM
OpenAI
Anthropic
Frontend
React.js
Apply
$120k – $200k per year • Equity 0.5–2% • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
Prompt Engineering
Apply
See all jobs
This is one of many
694,856 more open roles from verified company boards, updated every day.