{"id":1561352,"url":"https://alion.io/job/zoox-perception-deployment-engineer-model-deployment-optimization","title":"Perception Deployment Engineer - Model Deployment & Optimization","company":{"id":3626,"name":"Zoox","domain":"zoox.com","url":"https://alion.io/company/zoox","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":{"grade":"B","score":72,"open_postings":32,"ghost_share":0,"stale_share":0.906,"repost_share":0,"time_to_fill_p50_days":90,"computed_at":"2026-10-03T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Foster City, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":199000,"max":270000,"currency":"USD","period":"year","gross":null,"usd_annual":270000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"Computer Vision","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"KV Cache","optional":false},{"name":"Multimodal AI","optional":false},{"name":"ONNX","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"Quantization","optional":false},{"name":"Sensor Fusion","optional":false},{"name":"TensorRT","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"Vision-Language-Action","optional":false},{"name":"VLM","optional":false}],"status":"live","first_seen_at":"2026-09-30T21:31:00Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-04T00:32:46Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"The Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence.\nAs a Perception Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex computer vision or foundation models for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.\nIn this role, you will:\nDesign and develop production-level, low latency, and memory-safe C++ and CUDA code for real-time perception algorithms on vehicle systems.\n\nOptimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks.\n\nArchitect and implement model conversion and compilation pipelines using TensorRT for edge deployment.\n\nPerform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.\n\nDevelop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels to minimize latency and maximize memory bandwidth on AI accelerators.\n\nQualifications:\nProduction-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices.\n\nDeep expertise in model compression technologies (e.g., model quantization such as PTQ and QAT) and mixed-precision inference frameworks (INT8, FP8, BF16/FP16).\n\nProven experience optimizing large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs/VLAs) utilizing Efficient Attention mechanisms (e.g., FlashAttention, Linear Attention), KV-cache optimization (e.g., PagedAttention.\n\nExtensive experience with model conversion/compilation pipelines (e.g., ONNX, TensorRT, torch.compile) and performing rigorous latency benchmark and model quality parity valuation.\n\nProficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations.\n\nBonus Qualifications:\nFamiliarity with SOTA autonomous driving perception algorithms (temporal 3D object detection, BEV, 3D Occupancy Networks) and multi-modal sensor processing (Vision, LiDAR, Radar).\n\nExperience with end-to-end autonomous driving paradigms (VLM/VLA models, Foundation models) and edge deployment technologies (e.g., TensorRT-LLM).","description_format":"text","description_chars":2644,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Autonomous Driving","Smart Mobility","Automotive Manufacturing","Ride Hailing"],"lifecycle":[{"event":"open","at":"2026-10-01T04:09:15Z"}],"visa":[{"country":"US","licensed_sponsor":true,"evidence":"H-1B filings in 12 months: 373 · green card filings: 120","filings_12m":373,"filings_prev_12m":311,"green_card_filings_12m":120,"median_offered_wage_usd":181283,"route":null,"cap_exempt":false,"checked_at":"2026-10-03T21:08:04+00:00","sources":["US Department of Labor: LCA disclosure data (H-1B, H-1B1, E-3)","US Department of Labor: PERM disclosure data (green cards)"],"filings_for_role_12m":167}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":2,"expected_fill_days":90,"reasons":["conf:3","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-10-03T05:45:00Z"},"pay":{"stated_usd_annual":270000,"is_top_pay":true},"html_url":"https://alion.io/job/zoox-perception-deployment-engineer-model-deployment-optimization","json_url":"https://alion.io/job/zoox-perception-deployment-engineer-model-deployment-optimization.json","meta":{"generated_at":"2026-10-04T02:15:49Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3146,"day_limit":5000,"remaining_today":1854,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}