436,673open jobs
15,382companies
63,600added this week
Browse all
Salary
$270k – $300k per year
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

About the role

As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers - production AI teams looking to leverage high quality models and do inference at scale. For us, FDE is not a replacement for a Solutions Architect; you will partner with our SAs as a deep-domain specialist in inference optimization, fine-tuning pipelines, and production deployment. As key contributors to both the CX, Engineering, and Sales organizations, FDEs add tremendous value by ensuring we can meet the requirements of our most complex POCs, facilitate successful platform adoption, and guide tailored optimization efforts - directly impacting customer success, company growth, and the hardening of our core platform.

Responsibilities

  • Inference Engine Optimization: Select, configure, and optimize inference engine based on hardware, model architecture, and workload profile
  • Configuration & Performance Tuning: Develop configuration updates to win critical POCs, benchmarks, and optimize customer deployments; tune KV cache, apply speculative decoding, determine optimal tensor parallelism, and determine quantization strategy to hit throughput and latency targets.
  • Post-Training & Fine-Tuning: Drive hands-on RL training runs and optimize system design; guide customers through LoRA, SFT, DPO, RLHF, and GRPO pipelines from experimentation through production.
  • Strategic Customer Alignment: Act as the primary technical point of contact for aligned strategic accounts - monitoring and optimizing endpoint configurations, helping customers get the most out of the platform, and collaborating to ensure we hit critical milestones.
  • Opinionated Onboarding: Establish direct alignment with strategic customers at onboarding; ensure the right inference and post-training configurations are in place from day one to improve time-to-value.
  • Product Feedback Loop: Directly influence our software and model roadmap by surfacing insights from the field. Contribute back to the product where needed to support customer requirements or drive a better experience. Drive early feature and research adoption with strategic logos.

Qualifications

  • Experience: 5+ years in a technical role, with a strong focus on inference systems, open-source LLM deployment, or post-training workflows.
  • Inference Engine Depth: Expert-level, hands-on experience with inference engines (e.g., vLLM, TensorRT-LLM, SGLang); ability to diagnose and resolve performance issues at the engine level.
  • Inference Optimization: Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques
  • Post-Training Knowledge: Hands-on experience with fine-tuning and post-training pipelines, including LoRA, SFT, DPO, RLHF, and GRPO; ability to advise on system design
  • Model Landscape Awareness: Broad knowledge of state-of-the-art open-source models and strong judgment on model selection for specific customer use cases, hardware profiles, and performance targets.
  • Coding Proficiency: Strong Python skills; comfortable working in production environments

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers on our journey in building the next generation of AI infrastructure. 

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $270,000 - $300,000 OTE + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge. 

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our Privacy Policy at https://www.together.ai/privacy

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
436,673 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$156k – $234k per year • Remote/Hybrid • Full-Time • 10+ years exp • Irving • Jacksonville
Python
Python
Flask
FastAPI
Databases
Chroma
Milvus
Pinecone
AI/ML
LangChain
Vertex AI
Gemma
Fine-tuning
Prompt Engineering
AI Agents
NeMo Guardrails
Llama
Mistral
Pandas
NumPy
PyTorch
LLM
RAG
Google ADK
Hallucination
Hugging Face
NVIDIA NeMo
LLM Guardrails
Edge AI
Agentic Workflows
DevOps
OpenShift
CI/CD
Docker
Kubernetes
Vector
Apply
$191k – $334k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
Databases
Apache Kafka
AI/ML
Fine-tuning
Quantization
Prompt Engineering
AI Agents
TensorFlow
PyTorch
RAG
Anomaly Detection
Feature Store
LLM Guardrails
Cybersecurity
Zero Trust
Management
ServiceNow
Apply
$90k – $95k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Chicago
Python
SQL
Databases
MySQL
DevOps
AWS
Analytics
Tableau
ETL/ELT
Alteryx
Apply
$25k – $73k per year (Estimated) • Remote • Full-Time • 4+ years exp • High School Diploma • Brazil
Python
PowerShell
DevOps
Terraform
Azure DevOps
GitHub Actions
Kibana
Datadog
FluxCD
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Amazon EKS
Azure AKS
GitHub
Cybersecurity
OWASP ZAP
SonarQube
Mend
Apply
$24k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Mexico City
Python
SQL
Analytics
Microsoft Excel
Apply
$190k – $225k per year • In office • Full-Time • San Francisco
AI/ML
Cursor
ElevenLabs
Reinforcement Learning
Together AI
Pre-training
Analytics
Tableau
Marketing
Salesforce
HubSpot
Marketo
Apply
$92k – $192k per year (Estimated) • In office • 5+ years exp • Amsterdam
Python
Go
Rust
TypeScript
Databases
NATS
Apache Kafka
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Function Calling
AI Agents
LLM
RAG
Semantic Search
Together AI
Pre-training
Semantic Search
Knowledge Graph
Tool Use
DevOps
Prometheus
GitOps
ArgoCD
Kubernetes
Grafana
Incident Management
Management
Slack
Apply
$180k – $250k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Python
Go
Java
TypeScript
C++
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
Apply
$155k – $195k per year • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
Apply
$165k – $210k per year • Remote • Full-Time • 5+ years exp • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
Edge AI
Marketing
Salesforce
Apply
$76k – $92k per year • Remote • Internship • Bachelor's Degree • San Francisco
SQL
Analytics
SSIS
SSAS
Management
Microsoft Project
Apply
$70k – $82k per year • Remote • Internship • San Francisco
Analytics
Microsoft Excel
Apply
$70k – $82k per year • In office • Internship • San Francisco
Apply
$79k – $95k per year • Remote • Full-Time • San Francisco
Apply
$213k – $374k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Chicago • New York • Atlanta • San Francisco
AI/ML
AI Agents
Agentforce
Agentic Workflows
Marketing
Salesforce
Apply
See all jobs
This is one of many
436,673 more open roles from verified company boards, updated every day.