{"id":1163370,"url":"https://alion.io/job/the-vanguard-group-principal-aiml-engineer-2","title":"Principal AI/ML Engineer","company":{"id":253,"name":"The Vanguard Group","domain":"vanguard.com","url":"https://alion.io/company/vanguard","size_band":"1001-5000","is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":82,"open_postings":53,"ghost_share":0.038,"stale_share":0.585,"repost_share":0.075,"time_to_fill_p50_days":22,"computed_at":"2026-09-24T05:45:00Z"}},"role":"AI/ML","role_family":"AI/ML","seniority":"lead","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Toronto, Canada"],"countries":["CA"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":118000,"max_usd":231000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":26},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Incident Management","optional":false},{"name":"LLMOps","optional":false},{"name":"Machine Learning","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Multi-Agent Systems","optional":true}],"status":"live","first_seen_at":"2026-09-23T00:00:00Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-25T01:15:34Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable value for our clients. You will lead the architecture, engineering, operationalization, and ongoing reliability of advanced AI systems, ensuring they can scale across enterprise environments while meeting rigorous standards for performance, security, resilience, and responsible AI.This role sits at the critical intersection of AI research, engineering, product development, and operations. You will partner closely with world-class AI researchers, product leaders, and engineering teams to accelerate the journey from prototype to production. Your work will span some of the most advanced areas of AI, including Large Language Models (LLMs), Trustworthy AI, agentic systems, and emerging AI technologies.\nYou will mentor engineers, shape architecture, guide production support strategy, and serve as a thought leader for scaling AI across the organization. In addition to building and scaling AI solutions, you will help establish an engineering culture that emphasizes operational excellence, ownership, reliability, and continuous improvement.\nKey Responsibilities:\nAI Architecture & Technical Leadership\nDefine and lead the technical architecture for enterprise-scale AI and ML platforms. \nDesign scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads. \nEstablish architectural standards, engineering patterns, and best practices for AI deployment and operations. \nDrive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure. \nProductize AI Research\nPartner closely with AI researchers to transform cutting-edge prototypes into production-grade solutions. \nLead efforts to operationalize advanced AI capabilities across areas such as: \nLarge Language Models (LLMs) \nTrustworthy and Responsible AI \nAgentic AI Systems \nEstablish repeatable pathways that accelerate innovation-to-production cycles. \nEnsure production solutions maintain scientific rigor while meeting enterprise engineering standards. \nBridge the gap between research breakthroughs and sustainable business value. \nEngineering Excellence & Scalability\nSolve the organization's most complex AI engineering and scalability challenges. \nDesign systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost. \nDrive adoption of MLOps, LLMOps, and AI platform engineering best practices. \nImprove the robustness, maintainability, observability, and operational readiness of our AI products. \nIdentify and eliminate architectural bottlenecks that impact scale, reliability, or client experience. \nRaise standards through coaching, architecture reviews, design guidance, and technical leadership. \nProduction Reliability & Operational Leadership\nOwn the operational excellence, reliability, performance and availability of our products. \nLead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages. \nServe as the senior technical escalation point for the team's most challenging production challenges. \nEstablish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs). \nMentor and lead junior engineers in troubleshooting, root cause analysis, operational decision-making, and incident response. \nDrive post-incident reviews focused on learning, continuous improvement, and long-term corrective actions. \nDevelop operational processes that ensure AI solutions remain secure, scalable, performant, and reliable for business-critical use cases. \nPartner with product, infrastructure, security, and support teams to proactively identify operational risks and continuously improve service reliability. \nMentorship & Thought Leadership\nMentor AI and ML engineers within the team. \nFoster a culture of technical excellence and operational ownership where engineers are accountable not only for building systems, but also for running and supporting them successfully in production. \nRepresent our team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices. \nRequired Qualifications\n10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines. \nDeep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments. \nDemonstrated success leading complex technical initiatives from concept through deployment and ongoing operations. \nStrong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies. \nProven ability to mentor engineers and lead teams through highly complex technical and operational challenges. \nPreferred Qualifications\nExperience with foundation models, Large Language Models, and agentic AI architectures. \nExperience deploying agentic AI systems and multi-agent workflows. \nExperience with Trustworthy AI, Responsible AI, AI governance, or model risk management frameworks. \nExperience optimizing large-scale inference systems and AI infrastructure. \nExperience working in highly regulated environments and mission-critical production systems. \nWhat You'll Gain\nThis role offers a unique opportunity to operate at the forefront of applied artificial intelligence and help bridge world-class research with real-world impact.\nYou will:\nWork directly with world-class AI researchers on breakthrough technologies and next-generation AI capabilities. \nOwn a critical position in the pipeline that transforms cutting-edge research into client value. \nTackle some of the most difficult AI engineering, scalability, and operational challenges in the industry. \nBuild AI capabilities that deliver meaningful business outcomes for clients. \nDevelop deep expertise in operating advanced AI systems at scale while collaborating with leaders across research, product, and engineering. \nHow We Work\nVanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.","description_format":"text","description_chars":6593,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Hybrid work"],"hiring_locations":[{"name":"Canada","iso":"CA","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Financial Services"],"lifecycle":[{"event":"open","at":"2026-09-24T00:58:59Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":1,"expected_fill_days":22,"reasons":["conf:1","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/the-vanguard-group-principal-aiml-engineer-2","json_url":"https://alion.io/job/the-vanguard-group-principal-aiml-engineer-2.json","meta":{"generated_at":"2026-09-25T02:56:24Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3198,"day_limit":5000,"remaining_today":1802,"minute_limit":60,"resets_at":"2026-09-26T00:00:00Z"}}}