411,951open jobs
14,442companies
73,805added this week
Browse all
Salary
$60k – $300k per year
Location
Remote (United States)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Headquartered in San Francisco, California, Moss is a developer tools and AI infrastructure provider. The company offers a high-performance runtime for real-time semantic search designed specifically for conversational AI agents, voice assistants, and copilots. By deploying local-first vector indexes directly across browsers, mobile devices, and cloud servers with drop-in SDKs, it enables software engineers and AI developers to eliminate network latency, execute sub-10 millisecond retrieval queries, and deliver instant, responsive AI interactions.

Founding ML Engineer

Moss is building the retrieval runtime for real-time AI. We help agents access the right knowledge, conversation history, and user context in milliseconds, and use that context to decide what to do next.

We’re looking for a Founding ML Engineer to own the models and machine learning systems behind that experience. You’ll work across embeddings, retrieval, reranking, multilingual understanding, agent intelligence, and our Action Layer; taking ideas from experiments into production.

The work comes with real constraints: limited memory, CPU execution, changing context, multiple languages, and latency budgets that leave little room for error. Your job is to improve intelligence and quality while making the models practical to run.

What You’ll Do

  • Train and fine-tune embedding and reranking models for real-world retrieval workloads.
  • Build multilingual embedding models that retrieve accurately across languages, regions, and mixed-language conversations.
  • Improve the intelligence behind our Founding Agent. From understanding intent and retrieving context to choosing better responses and converting conversations into meaningful outcomes.
  • Help build Moss’s Action Layer, enabling agents to move from retrieving context to determining and executing the right next action.
  • Own the full model-development cycle: dataset creation, training, evaluation, optimization, deployment, and iteration.
  • Build evaluation pipelines that measure retrieval relevance, multilingual quality, agent outcomes, latency, memory usage, and inference cost.
  • Improve model efficiency through distillation, quantization, and inference optimization, particularly for CPU and ARM devices.
  • Work with runtime and SDK engineers to ship models across cloud, browser, edge, and device environments.
  • Investigate production failure cases and turn them into better datasets, evaluations, and models.
  • Make practical decisions about what to train, what to adapt, and what to ship.

Core Stack

The work spans:

  • Python and deep learning frameworks for training and experimentation.
  • Embedding models, rerankers, contrastive learning, and semantic retrieval.
  • Multilingual and cross-lingual representation learning.
  • Agent evaluation, intent understanding, tool selection, and action prediction.
  • Dataset curation, hard-negative mining, synthetic data, and reproducible evaluation.
  • Model distillation, quantization, and portable inference.
  • Moss’s Rust runtime and SDKs across cloud and on-device environments.

You don’t need to have worked with every part of the stack. You do need to understand how model decisions affect the system running them and the user experience they create.

Your First 90 Days

From day one: Work directly with our models, evaluation pipelines, Founding Agent, and production use cases. Start contributing code and experiments immediately.

  • By 30 days: Understand the current quality and performance baselines. Own a concrete improvement to a model, dataset, evaluation pipeline, or Founding Agent capability, with evidence that it solves a real problem.

  • By 60 days: Take a model improvement through evaluation and deployment. This could mean improving multilingual retrieval, making the Founding Agent more effective, or advancing an Action Layer capability. Work with the engineering team to validate its behavior under realistic hardware and workload constraints.

  • By 90 days: Independently own a meaningful part of the ML roadmap. Identify the next bottleneck, define the experiments, and drive improvements into production without waiting for a tightly scoped task.

What We’re Looking For

  • Experience training or fine-tuning models and deploying them into production.
  • Strong foundations in representation learning, information retrieval, and model evaluation.
  • Strong Python skills and the ability to write maintainable code beyond a research notebook.
  • An understanding of how training data, objectives, and evaluation choices affect real-world model behavior.
  • Ability to reason about tradeoffs between quality, latency, memory, and compute.
  • Comfort working through ambiguous problems and owning the result.
  • Clear communication about what you tried, what worked, what failed, and what should happen next.

Nice to Have

  • Experience with embedding models, rerankers, or search relevance.
  • Experience building multilingual or cross-lingual models.
  • Experience evaluating or improving conversational agents.
  • Experience with tool selection, action prediction, or agentic systems.
  • Experience deploying models on CPU, ARM, mobile, or browser environments.
  • Work on distillation, quantization, or inference performance.
  • Familiarity with Rust or systems-level performance profiling.
  • Research or open-source contributions relevant to efficient ML, retrieval, or agents.

Who You’ll Work With

You’ll work directly with the founder and our ML, runtime, backend, product, and SDK engineers. You’ll also work with the team supporting customer deployments, so your priorities stay connected to how people actually use Moss.

Why Moss

Moss is a YC F25 company building infrastructure for AI applications that need relevant context and the ability to act on it in real time.

You’ll have ownership over core technology: the models we build, how we evaluate them, how they improve our Founding Agent, and how they power the Action Layer. There’s room to pursue new ideas, and a clear expectation that those ideas become useful, reliable software.

If you want to build models and own what happens after they leave the training environment, we’d like to talk.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
411,951 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$110k – $239k per year (Estimated) • Remote • Full-Time • 4+ years exp
Python
Python
FastAPI
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Anthropic
Claude
Claude Code
Copilot
Cursor
Deepgram
ElevenLabs
Embeddings
LangChain
LangGraph
LiveKit
LLM
LLM Guardrails
OpenAI
Promptfoo
RAG
Text-to-Speech
Mobile
Twilio
DevOps
AWS
CI/CD
Docker
Git
Terraform
Vector
WebRTC
Analytics
ETL/ELT
Apply
Sr AI Engineer 1 day ago
$108k – $234k per year (Estimated) • Remote • Full-Time • 4+ years exp
Python
Python
Django
FastAPI
Databases
Pinecone
Weaviate
AI/ML
AI Agents
AutoGPT
Claude
Edge AI
Giskard
LangChain
LLM
Mistral
OpenAI
Prompt Engineering
Promptfoo
RAG
Transformers
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Terraform
Vector
Apply
$57k – $172k per year (Estimated) • Remote • Full-Time • 2+ years exp
TypeScript
JavaScript
AI/ML
Agentic Workflows
AI Agents
Claude
Claude Code
Cursor
Frontend
React.js
Mobile
React Native
DevOps
CI/CD
Git
Apply
$100k – $227k per year (Estimated) • Remote • Full-Time • 2+ years exp
JavaScript
Node JS
AI/ML
Agentic Workflows
AI Agents
Claude
Claude Code
Cursor
Frontend
React.js
DevOps
CI/CD
Git
QA
Jest
Apply
$34k – $162k per year (Estimated) • Remote • Full-Time • 2+ years exp
TypeScript
JavaScript
AI/ML
Agentic Workflows
AI Agents
Claude
Claude Code
Cursor
Frontend
Less
React Query
React.js
Redux
Redux Toolkit
Sass
Zustand
Mobile
State Management
DevOps
CI/CD
Git
QA
Jest
Apply
$150k – $200k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp • San Francisco
Apply
$60k – $300k per year • Equity 0.1–0.5% • Remote • Full-Time • 6+ years exp • San Francisco
C++
Elixir
JavaScript
Kotlin
Node JS
Python
Rust
Swift
TypeScript
AI/ML
CUDA
CUDA Toolkit
Edge AI
Embeddings
Semantic Search
Semantic Search
Frontend
npm
WebAssembly
DevOps
Vector
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
411,951 more open roles from verified company boards, updated every day.