{"id":1387741,"url":"https://alion.io/job/tavus-multimodal-ai-model-optimization-research-engineer","title":"Multimodal AI Model Optimization Research Engineer","company":{"id":1185,"name":"Tavus","domain":"tavus.io","url":"https://alion.io/company/tavus","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"B","score":70,"open_postings":14,"ghost_share":0.5,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-28T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":142000,"max_usd":309000,"period":"year","method":"role_country_seniority_unknown","sample_n":2986},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Knowledge Distillation","optional":false},{"name":"Model Distillation","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Quantization","optional":false},{"name":"Apache TVM","optional":true},{"name":"CUDA","optional":true},{"name":"CUDA Toolkit","optional":true},{"name":"Diffusion Models","optional":true},{"name":"Multimodal AI","optional":true},{"name":"ONNX Runtime","optional":true},{"name":"TensorRT","optional":true},{"name":"Text-to-Speech","optional":true},{"name":"Triton","optional":true},{"name":"WebRTC","optional":true},{"name":"XLA","optional":true}],"status":"live","first_seen_at":"2026-09-28T09:10:31Z","employer_posted_date":"2026-09-28","last_verified_at":"2026-09-29T03:21:41Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"About Us\nAt Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable.\nWe achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.\nBy enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.\nWe are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.\nThe Role\nWe’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team.\nOur ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.\nYour Mission\nTake cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization\n\nOwn the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality\n\nPartner closely with researchers and engineers to turn new ideas into deployable systems\n\nRequirements\nStrong experience in deep learning using PyTorch\n\nHands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision\n\nUnderstanding of efficient architectures such as low-rank adapters\n\nStrong understanding of inference performance and GPU/accelerator fundamentals\n\nStrong Python coding skills and reliable research engineering practices\n\nExperience working with large models and datasets in cloud environments\n\nAbility to read ML papers, reproduce results, and adapt ideas\n\nClear communication and collaboration skills\n\nPreferred Experience\nOptimization of diffusion models, video/audio generative models, or large language models\n\nExperience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)\n\nFamiliarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA\n\nExperience writing custom Triton/CUDA kernels or low-level performance tuning\n\nExperience with experiment tracking, benchmarking, and profiling at scale\n\nPrior experience in research engineering or applied science roles\n\nLocation\nThis position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.\nBenefits\nWhen you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.\nCulture & Diversity\nWe are not looking for cultural fits - we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients..","description_format":"text","description_chars":3307,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Flexible schedule","Unlimited PTO"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":true,"industries":["Conversational AI","Generative Video","Multimodal AI"],"lifecycle":[{"event":"open","at":"2026-09-28T10:56:15Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":40,"reasons":["conf:1","win:early"],"computed_at":"2026-09-29T05:09:33Z"},"pay":null,"html_url":"https://alion.io/job/tavus-multimodal-ai-model-optimization-research-engineer","json_url":"https://alion.io/job/tavus-multimodal-ai-model-optimization-research-engineer.json","meta":{"generated_at":"2026-09-29T05:09:33Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4917,"day_limit":5000,"remaining_today":83,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}