782,981open jobs
49,348companies
120,575added this week
Browse all
Salary
$120k – $250k per year
Location
Remote (United States)
Seniority
Junior · 1+ year exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
Cloudglue APIs turn videos into video context for AI — including speech, diarization, visual descriptions, sound, and more — so developers can search, chat, and build video-aware AI. The video understanding infrastructure for modern AI.

Cloudglue - Video Understanding Infrastructure

Cloudglue is a Y Combinator-backed startup building developer APIs that turn video and audio into structured, searchable data. We handle the hard infrastructure - transcription, visual analysis, search, extraction - so developers can build on top of video without managing ML pipelines themselves.

We process millions of minutes of video for customers building search, analytics, and automation products. The research problems are real: how do you retrieve the right 10 seconds from 10,000 hours of video? How do you extract structured facts from noisy, multimodal content? How do you reason across visual and spoken information at scale?

Our team has shipped large-scale systems at Snapchat and Amazon, with work presented at NeurIPS, ICCV, CVPR, KubeCon, and DEF CON. We’re a small, technical team where researchers ship code and engineers read papers.

The Role

We’re looking for a research engineer to work on the core multimodal retrieval and video reasoning systems that power Cloudglue. This is a 50/50 research and engineering role - you’ll design novel approaches to hard retrieval and understanding problems, and you’ll ship them into production where real customers depend on them.

You’ll work across:

  • Multimodal retrieval - finding relevant moments across visual, audio, and text signals in large video collections
  • Structured extraction - pulling entities, facts, and relationships from video content
  • Video reasoning - understanding temporal, causal, and semantic relationships across long-form content
  • Evaluation and benchmarking - designing metrics and datasets to measure real-world system quality

This is not a pure research role. You’ll be expected to take ideas from paper to prototype to production. But it’s also not a pure engineering role - we need someone with genuine research depth who can identify the right problems to work on and design novel solutions.

What You’ll Do

  • Multimodal retrieval: Design and improve retrieval systems that search across video, audio, and text - including embedding models, re-ranking, and hierarchical search strategies.

  • Video understanding: Build systems that extract structured information from video - temporal segmentation, entity extraction, scene understanding, and content summarization.

  • Model fine-tuning & integration: Fine-tune and adapt vision and language models (LoRA/PEFT, full fine-tuning) for production use cases. Evaluate open-source and proprietary models and orchestrate them in serving pipelines.

  • Experiment and ship: Run experiments, analyze results rigorously, and turn successful research into production systems that handle real-world video at scale.

  • Collaborate: Work directly with founders and infrastructure engineers. Short feedback loops, no layers of process.

What We’re Looking For

Required

  • MS or PhD in computer science, machine learning, or a related field
  • Research experience in one or more of: multimodal learning, information retrieval, computer vision, NLP, or video understanding
  • Strong implementation skills in Python and PyTorch (or equivalent)
  • Ability to independently drive research from idea to experiment to working system

Nice to Have

  • First-author publication at a top venue (NeurIPS, CVPR, ICCV, ECCV, ACL, EMNLP, SIGIR, ISMIR, ICASSP, or similar)
  • Experience with video or multimodal foundation models (CLIP, LLaVA, Qwen3-VL, etc.)
  • Experience with retrieval systems, embedding models, or ranking/re-ranking pipelines
  • Experience deploying ML systems in production
  • Familiarity with vector databases (Milvus, Weaviate) or search infrastructure
  • Experience with model fine-tuning techniques (LoRA, PEFT, QLoRA) and training infrastructure (Ray, Kubeflow, or similar)
  • Experience with ML inference serving (vLLM, TensorRT, Triton, or similar)

Why Cloudglue?

Video is the largest and most underutilized data source on the internet. Most software still can’t meaningfully search or reason over it. The research problems here - multimodal retrieval, temporal reasoning, structured extraction from noisy real-world content - are genuinely unsolved and directly tied to the product.

If you want to work on:

  • Research problems with immediate, measurable product impact
  • A domain where the state of the art is still being defined
  • A small team where your research directly shapes the product
  • Multimodal systems at real scale, not toy benchmarks

…this is that role.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
782,981 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
≈ $107k – $242k per year (Estimated) • In office • Full-Time • PhD • Seattle
Python
C++
AI/ML
LLM
NCCL
Apply
≈ $108k – $244k per year (Estimated) • In office • Full-Time • PhD • San Jose
Python
C
C++
C
FFmpeg
C++
TensorFlow C++
PyTorch C++
AI/ML
OpenCV
Transformers
TensorFlow
PyTorch
DevOps
Linux
Analytics
A/B Testing
Apply
≈ $102k – $230k per year (Estimated) • In office • Full-Time • PhD • San Diego
Python
C++
C++
PyTorch C++
AI/ML
PyTorch
Apply
≈ $110k – $248k per year (Estimated) • In office • Full-Time • Bachelor's Degree • San Jose
Python
Java
AI/ML
CUDA Toolkit
Fine-tuning
Computer Vision
NLP
TensorFlow
PyTorch
CUDA
Recommender Systems
Machine Learning
DevOps
Docker
Kubernetes
Linux
Apply
≈ $114k – $256k per year (Estimated) • In office • Full-Time • PhD • San Jose
Python
AI/ML
Multimodal AI
Computer Vision
TensorFlow
PyTorch
Machine Learning
Apply
AI Deployment Lead 7 hours ago
$110k – $150k per year • In office • Full-Time • 3+ years exp • New York
Python
SQL
AI/ML
AI Agents
LLM
Apply
$130k – $170k per year • In office • Full-Time • 3+ years exp • New York
Python
SQL
AI/ML
Copilot
AI Agents
Agentic Workflows
Analytics
ETL/ELT
Apply
$120k – $250k per year • Equity 0.5–2.5% • Remote (United States) • Full-Time • 3+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
Weaviate
Milvus
DevOps
GCP
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
Management
Stripe
Apply
$120k – $250k per year • Equity 0.5–2.5% • Remote (United States) • Full-Time • 3+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
SQL
Node JS
Databases
PostgreSQL
Weaviate
Milvus
Frontend
Next.js
React.js
DevOps
Rest API
AWS
Management
Stripe
Apply
$78k – $120k per year • Remote (EU, United States, United Kingdom) • Internship • PhD • London
Python
Databases
ElasticSearch
AI/ML
RLHF
AI Agents
PyTorch
DPO
GRPO
Apply
$120k – $250k per year • Equity 0.5–2.5% • Remote (United States) • Full-Time • 3+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
Weaviate
Milvus
DevOps
GCP
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
Management
Stripe
Apply
$120k – $250k per year • Equity 0.5–2.5% • Remote (United States) • Full-Time • 3+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
SQL
Node JS
Databases
PostgreSQL
Weaviate
Milvus
Frontend
Next.js
React.js
DevOps
Rest API
AWS
Management
Stripe
Apply
$125k – $175k per year • Equity 0.2–2% • In office • Full-Time • 3+ years exp • PhD • San Francisco
Python
AI/ML
Multimodal AI
Time Series Forecasting
Machine Learning
Robotics
Sensor Fusion
Apply
$180k – $250k per year • Remote (United States) • Full-Time • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
≈ $134k – $359k per year (Estimated) • In office • Internship • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
Founding Engineer 9 hours ago
$110k – $180k per year • Equity 0.1–1% • In office • Full-Time • San Francisco
Python
JavaScript
Node JS
AI/ML
Vertex AI
OpenAI
Anthropic
Frontend
Next.js
React.js
DevOps
Azure
Kubernetes
Apply
$60k – $84k per year • In office • Internship • San Francisco
Python
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
LLM
OpenAI
Anthropic
LLM Guardrails
Tool Use
Frontend
Next.js
React.js
DevOps
Azure
AWS
Kubernetes
Apply
See all jobs
This is one of many
782,981 more open roles from verified company boards, updated every day.