{"id":1245125,"url":"https://alion.io/job/cloudglue-research-engineer","title":"Research Engineer","company":{"id":3779092,"name":"Cloudglue","domain":"cloudglue.dev","url":"https://alion.io/company/cloudglue","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Work at a Startup","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"junior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":{"min":120000,"max":250000,"currency":"USD","period":"year","gross":null,"usd_annual":250000},"salary_estimate":null,"experience_years_min":1,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"Computer Vision","optional":false},{"name":"Fine-tuning","optional":false},{"name":"LoRA","optional":false},{"name":"Machine Learning","optional":false},{"name":"Multimodal AI","optional":false},{"name":"NLP","optional":false},{"name":"PEFT","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Reinforcement Learning","optional":false},{"name":"CLIP","optional":true},{"name":"Kubeflow","optional":true},{"name":"LLaVA","optional":true},{"name":"Milvus","optional":true},{"name":"QLoRA","optional":true},{"name":"Ray","optional":true},{"name":"TensorRT","optional":true},{"name":"Transformers","optional":true},{"name":"Triton","optional":true},{"name":"vLLM","optional":true},{"name":"Weaviate","optional":true}],"status":"live","first_seen_at":"2026-09-25T17:16:36Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-25T23:42:53Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Cloudglue - Video Understanding Infrastructure\nCloudglue is a Y Combinator-backed startup building developer APIs that turn video and audio into structured, searchable data. We handle the hard infrastructure - transcription, visual analysis, search, extraction - so developers can build on top of video without managing ML pipelines themselves.\nWe process millions of minutes of video for customers building search, analytics, and automation products. The research problems are real: how do you retrieve the right 10 seconds from 10,000 hours of video? How do you extract structured facts from noisy, multimodal content? How do you reason across visual and spoken information at scale?\nOur team has shipped large-scale systems at Snapchat and Amazon, with work presented at NeurIPS, ICCV, CVPR, KubeCon, and DEF CON. We’re a small, technical team where researchers ship code and engineers read papers.\nThe Role\nWe’re looking for a research engineer to work on the core multimodal retrieval and video reasoning systems that power Cloudglue. This is a 50/50 research and engineering role - you’ll design novel approaches to hard retrieval and understanding problems, and you’ll ship them into production where real customers depend on them.\nYou’ll work across:\nMultimodal retrieval - finding relevant moments across visual, audio, and text signals in large video collections\nStructured extraction - pulling entities, facts, and relationships from video content\nVideo reasoning - understanding temporal, causal, and semantic relationships across long-form content\nEvaluation and benchmarking - designing metrics and datasets to measure real-world system quality\nThis is not a pure research role. You’ll be expected to take ideas from paper to prototype to production. But it’s also not a pure engineering role - we need someone with genuine research depth who can identify the right problems to work on and design novel solutions.\nWhat You’ll Do\nMultimodal retrieval: Design and improve retrieval systems that search across video, audio, and text - including embedding models, re-ranking, and hierarchical search strategies.\n\nVideo understanding: Build systems that extract structured information from video - temporal segmentation, entity extraction, scene understanding, and content summarization.\n\nModel fine-tuning & integration: Fine-tune and adapt vision and language models (LoRA/PEFT, full fine-tuning) for production use cases. Evaluate open-source and proprietary models and orchestrate them in serving pipelines.\n\nExperiment and ship: Run experiments, analyze results rigorously, and turn successful research into production systems that handle real-world video at scale.\n\nCollaborate: Work directly with founders and infrastructure engineers. Short feedback loops, no layers of process.\n\nWhat We’re Looking For\nRequired\nMS or PhD in computer science, machine learning, or a related field\nResearch experience in one or more of: multimodal learning, information retrieval, computer vision, NLP, or video understanding\nStrong implementation skills in Python and PyTorch (or equivalent)\nAbility to independently drive research from idea to experiment to working system\nNice to Have\nFirst-author publication at a top venue (NeurIPS, CVPR, ICCV, ECCV, ACL, EMNLP, SIGIR, ISMIR, ICASSP, or similar)\nExperience with video or multimodal foundation models (CLIP, LLaVA, Qwen3-VL, etc.)\nExperience with retrieval systems, embedding models, or ranking/re-ranking pipelines\nExperience deploying ML systems in production\nFamiliarity with vector databases (Milvus, Weaviate) or search infrastructure\nExperience with model fine-tuning techniques (LoRA, PEFT, QLoRA) and training infrastructure (Ray, Kubeflow, or similar)\nExperience with ML inference serving (vLLM, TensorRT, Triton, or similar)\nWhy Cloudglue?\nVideo is the largest and most underutilized data source on the internet. Most software still can’t meaningfully search or reason over it. The research problems here - multimodal retrieval, temporal reasoning, structured extraction from noisy real-world content - are genuinely unsolved and directly tied to the product.\nIf you want to work on:\nResearch problems with immediate, measurable product impact\nA domain where the state of the art is still being defined\nA small team where your research directly shapes the product\nMultimodal systems at real scale, not toy benchmarks\n…this is that role.","description_format":"text","description_chars":4400,"description_truncated":false,"requirements":{"experience_years_min":1,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"master","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T17:16:36Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":19,"reasons":["conf:2","win:early","comp:junior"],"computed_at":"2026-09-26T02:36:20Z"},"pay":{"stated_usd_annual":250000,"is_top_pay":true},"html_url":"https://alion.io/job/cloudglue-research-engineer","json_url":"https://alion.io/job/cloudglue-research-engineer.json","meta":{"generated_at":"2026-09-26T02:36:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2718,"day_limit":5000,"remaining_today":2282,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}