{"id":1271014,"url":"https://alion.io/job/maxinsights-machine-learning-engineer-egocentric-3d-human-pose","title":"Machine Learning Engineer (Egocentric 3D Human Pose)","company":{"id":2224833,"name":"Maxinsights","domain":"maxinsights.ai","url":"https://alion.io/company/maxinsights","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Santa Clara, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":146000,"max_usd":278000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":334},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Computer Vision","optional":false},{"name":"Localization","optional":false},{"name":"Machine Learning","optional":false},{"name":"Multimodal AI","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"TensorFlow","optional":false},{"name":"Imitation Learning","optional":true},{"name":"Inverse Kinematics","optional":true},{"name":"Teleoperation","optional":true}],"status":"live","first_seen_at":"2026-07-13T21:13:51Z","employer_posted_date":"2026-07-13","last_verified_at":"2026-10-04T00:30:19Z","board_verified":true,"closed_at":null,"days_open":82,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":82},"description":"Job Description:\nWe are looking for a Machine Learning Engineer to join our core research and development team, focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.\nHuman demonstration data is the fuel for robot learning, and the quality of that data is bounded by how well we can reconstruct what the hands and body actually did. In this role, you will own models and pipelines that turn head-mounted and body-mounted camera streams - often wide-FOV, stereo, motion-blurred, and heavily self-occluded - into metrically accurate, temporally stable 3D pose that is directly usable for robot policy training and human-to-robot retargeting.\nYou will work across the full stack: capture rig and calibration, ground-truth annotation tooling, model training and evaluation, and production deployment at scale. This role suits engineers who are equally comfortable with multi-view geometry and modern deep learning, and who are motivated by hard, measurable accuracy problems on real-world data.\nResponsibilities\nBuild 3D body and hand pose estimation models for egocentric video, covering 2D/3D keypoints, parametric body and hand models (SMPL/SMPL-X, MANO), and full-sequence motion recovery from monocular and stereo first-person cameras.\n\nSolve the hard cases specific to the egocentric viewpoint - severe self-occlusion, truncated limbs, extreme perspective foreshortening, hand-object interaction, rapid head motion, and rolling-shutter and motion-blur artifacts.\n\nOwn camera geometry and calibration: fisheye and wide-FOV camera models (Kannala-Brandt, Double Sphere), intrinsic/extrinsic calibration, stereo triangulation, and head-to-body coordinate-frame alignment for metric-scale output.\n\nDrive temporal consistency and physical plausibility through robust estimation, smoothing and filtering, kinematic and anatomical constraints, contact and penetration reasoning, and multi-view or multi-modal fusion (e.g. IMU, exocentric cameras, marker-based mocap).\n\nBuild the ground-truth and evaluation loop: semi-automatic annotation and keypoint propagation tools, confidence-aware quality gating, and evaluation protocols that separate real accuracy gains from benchmark noise.\n\nShip end-to-end systems at scale - large-scale training, high-throughput video inference, and reliable production pipelines over high-bandwidth multi-camera data.\n\nTranslate reconstructed human motion into robot-usable data, collaborating with robotics and product teams on retargeting fidelity for dexterous hands and humanoid end-effectors.\n\nContribute to technical design, code quality, and best practices, and help shape the long-term direction of the company’s perception stack.\n\nRequired Qualifications\nBachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related technical field, or equivalent practical experience.\n\n3+ years of experience building and shipping machine learning systems.\n\nProven hands-on experience developing and deploying 3D human pose, hand pose, or human motion tracking models from video.\n\nWorking knowledge of multi-view geometry and camera models: projection, calibration, triangulation, rigid-body transforms, and coordinate-frame management.\n\nStrong proficiency in Python and at least one major deep learning framework (e.g. PyTorch, TensorFlow).\n\nSolid understanding of modern deep learning concepts, training workflows, model evaluation, and real-world, production-oriented ML pipelines.\n\nStrong problem-solving skills and the ability to work effectively in a fast-moving, collaborative environment.\n\nPreferred Qualifications\nDirect experience with egocentric or head-mounted perception (AR/VR headsets, smart glasses, chest- or head-mounted capture rigs), including fisheye and stereo pipelines.\n\nDeep expertise in human kinematics and parametric models - SMPL/SMPL-X, MANO, inverse kinematics, markerless motion capture, and hand-object pose estimation.\n\nFamiliarity with relevant egocentric vision datasets and benchmarks.\n\nFamiliarity with state-of-the-art architectures for video and 3D data (e.g. video transformers, diffusion-based motion priors, 3D CNNs, etc).\n\nExperience building or operating multi-camera capture systems, time synchronization, and calibration infrastructure.\n\nExperience with human-to-robot motion retargeting, teleoperation data, or imitation learning pipelines.\n\nExperience with annotation tooling, active learning, or data quality systems for large-scale video.\n\nPublications at leading venues (CVPR, ICCV, ECCV, NeurIPS, SIGGRAPH, 3DV), open-source contributions, or demonstrated impact in applied ML or AI systems.\n\nWhat We Offer\nCompetitive salary and options package.\n\nComprehensive health, dental, and vision insurance.\n\n401(k) plan.\n\nPaid time off.\n\nDirect collaboration with leading experts in the field of robotics and AI.\n\nMaxInsights is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.\nDefault Benefits:\nHealth insurance\n\nVision care\n\nDental coverage\n\n401(k)\n\nPaid holidays\n\nPTO (Paid Time Off)\n\nSick leave","description_format":"text","description_chars":5133,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Health insurance","Vision insurance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Robotics AI","Data Providers & Datasets"],"lifecycle":[{"event":"open","at":"2026-09-26T00:09:15Z"}],"visa":[{"country":"US","licensed_sponsor":true,"evidence":"H-1B filings in 12 months: 6","filings_12m":6,"filings_prev_12m":1,"green_card_filings_12m":0,"median_offered_wage_usd":162500,"route":null,"cap_exempt":false,"checked_at":"2026-10-03T21:08:04+00:00","sources":["US Department of Labor: LCA disclosure data (H-1B, H-1B1, E-3)"],"filings_for_role_12m":4}],"liveness":{"score":11,"band":"cold","label":"Long shot","p_open":1,"p_active":0.385,"p_room":0.28,"age_days":81,"expected_fill_days":19,"reasons":["conf:15","velocity","win:tail","crowd:"],"computed_at":"2026-10-03T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/maxinsights-machine-learning-engineer-egocentric-3d-human-pose","json_url":"https://alion.io/job/maxinsights-machine-learning-engineer-egocentric-3d-human-pose.json","meta":{"generated_at":"2026-10-04T01:18:42Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1713,"day_limit":5000,"remaining_today":3287,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}