Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 21, 2026. VinFast scores B on the Alion truth index.
Responsibilities: Research and build a unified, real-time Speech Understanding model - powered by an LLM - that lets the Voice Agent perceive speakers, their state, the environment, and audio content beyond the transcript; prototype agent-awareness behaviors using simulated agents and scripted scenarios. Build the Vietnamese data and benchmark foundations: data collection, annotation, and synthetic data, plus a shared evaluation framework that measures real impact on agent behavior, not only model metrics. Train, fine-tune, and post-train audio and audio-LLM models (pretraining, SFT, alignment) and optimize them for low-latency streaming inference. Work with the ASR, TTS, Foundation, and Agent teams to integrate models into conversational systems and voice assistants, including demanding in-vehicle conditions. Stay current with the research community and contribute new ideas and, where appropriate, publications. In this role, you'll join the Speech Understanding team and work across the full loop - data, models, evaluation, and integration into a production Voice Agent. We're looking for people with relevant experience, passion, and drive, ready to learn fast and thrive under pressure.
Requirements Bachelor's Degree or higher in Computer Science, Artificial Intelligence, or related fields. Minimum 2 years of experience in AI/ML, with a focus on speech and audio (ASR, audio understanding, or related areas) or on audio/multimodal LLMs. Proficient in Python and PyTorch, with a strong grasp of Transformers and modern speech/audio models (self-supervised encoders, Whisper-style models, audio-LLMs). Hands-on experience training or fine-tuning models - ideally LLMs (SFT, PEFT/LoRA) - and building or evaluating speech datasets. Comfortable with open-ended, research-driven problems where the problem definition is still evolving. Good English proficiency to read technical documentation and research papers. Preferred Experience with Vietnamese, multilingual, or code-switching speech. Experience with real-time streaming systems or voice agents, and interest in multimodal (voice and vision) agents. Publications in top-tier AI/speech conferences (e.g., Interspeech, ICASSP, NeurIPS).
Benefits Competitive salary Premium healthcare package, including PVI insurance & annual health check-ups 13th-month salary & performance bonuses to reward your contributions Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool... Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain Work Location: Technopark Tower, Vinhomes Ocean Park, Gia Lam, Hanoi, Vietnam With respect to all your personal data shared to VinFast in the application and the entire recruitment process of VinFast, by clicking “Apply”, submitting your resumé/CV and/or participating in VinFast's recruitment process, you agree that you have read VinFast's Personal Data Protection Policy ("Policy") posted at https://vinfastauto.com/vn_vi/dieu-khoan-phap-ly or https://vinfast.vn/privacy-policy/, you agree to the Policy and consent for VinFast to process your personal data in accordance with the Policy and the applicable regulations on personal data protection. To all recruitment agencies: VinFast does not accept agency resumes. Please do not forward resumes to our careers alias or other VinFast employees. VinFast is not responsible for any fees related to unsolicited resumes.

