{"id":1270990,"url":"https://alion.io/job/maxinsights-machine-learning-engineer-video-understanding-segmentation","title":"Machine Learning Engineer (Video Understanding & Segmentation)","company":{"id":2224833,"name":"Maxinsights","domain":"maxinsights.ai","url":"https://alion.io/company/maxinsights","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Santa Clara, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":136000,"max_usd":263000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":257},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"CLIP","optional":false},{"name":"Computer Vision","optional":false},{"name":"Embodied AI","optional":false},{"name":"FAISS","optional":false},{"name":"Fine-tuning","optional":false},{"name":"Function Calling","optional":false},{"name":"LangChain","optional":false},{"name":"LlamaIndex","optional":false},{"name":"LLM","optional":false},{"name":"Machine Learning","optional":false},{"name":"Milvus","optional":false},{"name":"Multimodal AI","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Tool Use","optional":false},{"name":"Vision-Language-Action","optional":true}],"status":"live","first_seen_at":"2026-09-25T20:40:01Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-27T02:48:44Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Job Description:\nWe are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale - turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.\nResponsibilities\nBuild and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.\n\nDevelop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.\n\nDesign and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.\n\nBuild automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.\n\nArchitect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.\n\nDevelop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.\n\nCollaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.\n\nEvaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.\n\nContribute to internal tooling, documentation, patents, and open-source initiatives where applicable.\n\nMentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.\n\nMinimum Qualifications\nMS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.\n\n3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.\n\nStrong proficiency in Python and PyTorch, with solid software engineering fundamentals.\n\nHands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.\n\nExperience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).\n\nFamiliarity with agentic system design - tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).\n\nExperience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).\n\nPreferred Qualifications\nPhD with a research focus in video understanding, multi-modal learning, or vision-language models.\n\nExperience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.\n\nExperience working with egocentric video datasets or head-mounted-device (HMD) captured data.\n\nTrack record of deploying production-scale video search or retrieval systems.\n\nExperience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.\n\nPublications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).\n\nExperience with humanoid robotics or embodied AI data pipelines is a plus.\n\nDefault Benefits:\nHealth insurance\n\nVision care\n\nDental coverage\n\n401(k)\n\nPaid holidays\n\nPTO (Paid Time Off)\n\nSick leave","description_format":"text","description_chars":4072,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"master","optional":false},"security_clearance":false,"languages":[]},"benefits":["Health insurance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Robotics AI","Data Providers & Datasets"],"lifecycle":[{"event":"open","at":"2026-09-26T00:09:15Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":17,"reasons":["conf:5","win:early"],"computed_at":"2026-09-26T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/maxinsights-machine-learning-engineer-video-understanding-segmentation","json_url":"https://alion.io/job/maxinsights-machine-learning-engineer-video-understanding-segmentation.json","meta":{"generated_at":"2026-09-27T04:01:59Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3762,"day_limit":5000,"remaining_today":1238,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}