746,581open jobs
44,818companies
107,195added this week
Browse all
Salary
≈ $118k – $231k per year (Estimated)
Location
Hybrid (Toronto, Canada)
Seniority
Principal · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 23, 2026. The Vanguard Group scores B on the Alion truth index.

Overview
Company
Impact
Profile match
The Vanguard Group is an American investment manager founded in 1975 that pioneered the retail index fund and is now one of the two largest asset managers in the world. It is owned by the funds it runs, and therefore by their shareholders, a mutual structure that lets it push management fees toward zero and has reshaped pricing across the whole industry. From Malvern in Pennsylvania it operates mutual funds, exchange traded funds, brokerage, retirement plan administration and advice services for tens of millions of individual and institutional clients.
As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable value for our clients. You will lead the architecture, engineering, operationalization, and ongoing reliability of advanced AI systems, ensuring they can scale across enterprise environments while meeting rigorous standards for performance, security, resilience, and responsible AI.

This role sits at the critical intersection of AI research, engineering, product development, and operations. You will partner closely with world-class AI researchers, product leaders, and engineering teams to accelerate the journey from prototype to production. Your work will span some of the most advanced areas of AI, including Large Language Models (LLMs), Trustworthy AI, agentic systems, and emerging AI technologies.

You will mentor engineers, shape architecture, guide production support strategy, and serve as a thought leader for scaling AI across the organization. In addition to building and scaling AI solutions, you will help establish an engineering culture that emphasizes operational excellence, ownership, reliability, and continuous improvement.

Key Responsibilities:

AI Architecture & Technical Leadership

  • Define and lead the technical architecture for enterprise-scale AI and ML platforms.
  • Design scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads.
  • Establish architectural standards, engineering patterns, and best practices for AI deployment and operations.
  • Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure.

Productize AI Research

  • Partner closely with AI researchers to transform cutting-edge prototypes into production-grade solutions.
  • Lead efforts to operationalize advanced AI capabilities across areas such as:
  • Large Language Models (LLMs)
  • Trustworthy and Responsible AI
  • Agentic AI Systems
  • Establish repeatable pathways that accelerate innovation-to-production cycles.
  • Ensure production solutions maintain scientific rigor while meeting enterprise engineering standards.
  • Bridge the gap between research breakthroughs and sustainable business value.

Engineering Excellence & Scalability

  • Solve the organization's most complex AI engineering and scalability challenges.
  • Design systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost.
  • Drive adoption of MLOps, LLMOps, and AI platform engineering best practices.
  • Improve the robustness, maintainability, observability, and operational readiness of our AI products.
  • Identify and eliminate architectural bottlenecks that impact scale, reliability, or client experience.
  • Raise standards through coaching, architecture reviews, design guidance, and technical leadership.

Production Reliability & Operational Leadership

  • Own the operational excellence, reliability, performance and availability of our products.
  • Lead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages.
  • Serve as the senior technical escalation point for the team's most challenging production challenges.
  • Establish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs).
  • Mentor and lead junior engineers in troubleshooting, root cause analysis, operational decision-making, and incident response.
  • Drive post-incident reviews focused on learning, continuous improvement, and long-term corrective actions.
  • Develop operational processes that ensure AI solutions remain secure, scalable, performant, and reliable for business-critical use cases.
  • Partner with product, infrastructure, security, and support teams to proactively identify operational risks and continuously improve service reliability.

Mentorship & Thought Leadership

  • Mentor AI and ML engineers within the team.
  • Foster a culture of technical excellence and operational ownership where engineers are accountable not only for building systems, but also for running and supporting them successfully in production.
  • Represent our team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices.

Required Qualifications

  • 10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines.
  • Deep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments.
  • Demonstrated success leading complex technical initiatives from concept through deployment and ongoing operations.
  • Strong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies.
  • Proven ability to mentor engineers and lead teams through highly complex technical and operational challenges.

Preferred Qualifications

  • Experience with foundation models, Large Language Models, and agentic AI architectures.
  • Experience deploying agentic AI systems and multi-agent workflows.
  • Experience with Trustworthy AI, Responsible AI, AI governance, or model risk management frameworks.
  • Experience optimizing large-scale inference systems and AI infrastructure.
  • Experience working in highly regulated environments and mission-critical production systems.

What You'll Gain

This role offers a unique opportunity to operate at the forefront of applied artificial intelligence and help bridge world-class research with real-world impact.

You will:

  • Work directly with world-class AI researchers on breakthrough technologies and next-generation AI capabilities.
  • Own a critical position in the pipeline that transforms cutting-edge research into client value.
  • Tackle some of the most difficult AI engineering, scalability, and operational challenges in the industry.
  • Build AI capabilities that deliver meaningful business outcomes for clients.
  • Develop deep expertise in operating advanced AI systems at scale while collaborating with leaders across research, product, and engineering.

How We Work

Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
746,581 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Toronto
≈ $132k – $259k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Montreal
AI/ML
AI Agents
NLP
Machine Learning
Management
ServiceNow
Apply
≈ $73k – $138k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Toronto
Python
Go
JavaScript
Rust
TypeScript
PowerShell
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
AI/ML
LangGraph
AutoGen
LangChain
Spark
Embeddings
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Semantic Search
LLMOps
Context Engineering
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Machine Learning
DevOps
GCP
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Spinnaker
Cybersecurity
SIEM
Apply
≈ $144k – $278k per year (Estimated) • Equity • Remote (Mexico) • 8+ years exp
AI/ML
Function Calling
AI Agents
LLM
RAG
Tool Use
Apply
$64k – $85k per year • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Toronto
Python
SQL
Python
pySpark
Databases
Google BigQuery
BigQuery
AI/ML
Spark
Airflow
XGBoost
Vertex AI
AI Agents
LightGBM
TensorFlow
PyTorch
LLM
Amazon SageMaker
Feature Store
Recommender Systems
Machine Learning
DevOps
Terraform
Azure
IAM
Apply
$158k – $239k per year • Equity • Remote (United States, Canada) • 8+ years exp
AI/ML
AI Agents
LLM
RAG
OpenAI
Anthropic
Agentic Workflows
DevOps
AWS
GitHub
Cybersecurity
Threat Modeling
Apply
Equity • Remote (Poland) • Full-Time • 5+ years exp
Apex
Apex
Lightning Web Components
AI/ML
AI Agents
RAG
Anthropic
Agentforce
LLM Guardrails
DevOps
GCP
Azure
AWS
Platform Engineering
Management
Agile
Apply
≈ $78k – $182k per year (Estimated) • Equity • Remote (Poland) • Full-Time • 3+ years exp
Apex
Apex
Lightning Web Components
AI/ML
AI Agents
RAG
Agentforce
LLM Guardrails
DevOps
GCP
Azure
AWS
Platform Engineering
Apply
$170k – $256k per year • Hybrid • Bachelor's Degree • Palo Alto
AI/ML
Fine-tuning
AI Agents
LLM Guardrails
DevOps
GCP
VMWare
Azure
AWS
KVM
Hyper-V
Apply
Data Scientist 1 day ago
$73k – $121k per year • Equity • Remote (United States) • 3+ years exp • Bachelor's Degree
Python
AI/ML
Scikit-learn
TensorFlow
Pandas
NumPy
Keras
PyTorch
Machine Learning
DevOps
GCP
Azure
AWS
Apply
≈ $111k – $232k per year (Estimated) • Equity • Remote (Poland) • Full-Time • 10+ years exp
ABAP
DevOps
GCP
Azure
AWS
Platform Engineering
Apply
AI/ML Engineer 11 days ago
≈ $100k – $227k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Toronto
Python
Databases
Apache Kafka
AI/ML
Fine-tuning
AI Agents
Flink
RAG
Semantic Search
Amazon SageMaker
Semantic Search
Knowledge Graph
Machine Learning
DevOps
CI/CD
AWS
Docker
Kubernetes
AWS Lambda
Amazon S3
Amazon ECS
Amazon Kinesis
Amazon EventBridge
AWS Step Functions
Analytics
ETL/ELT
Apply
≈ $104k – $235k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Master's Degree • Toronto
AI/ML
SHAP
Fine-tuning
Reinforcement Learning
Scikit-learn
Diffusion Models
AI Agents
Transformers
TensorFlow
NumPy
PyTorch
LLM
Knowledge Graph
Machine Learning
Apply
≈ $115k – $225k per year (Estimated) • Hybrid • Full-Time • 12+ years exp • Toronto
Databases
Databricks
AI/ML
LangGraph
LangChain
Model Context Protocol
AI Agents
AgentOps
AWS Bedrock
RAG
Amazon SageMaker
GraphRAG
DevOps
AWS
AWS Step Functions
Analytics
AWS Glue
Apply
≈ $153k – $324k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Malvern • Charlotte
AI/ML
AI Agents
LLMOps
Multi-Agent Systems
Machine Learning
DevOps
Platform Engineering
Incident Management
Apply
≈ $145k – $263k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Malvern • Dallas
Python
Java
AI/ML
AI Agents
LLM
DevOps
Azure
AWS
GitHub
Cybersecurity
SIEM
Apply
≈ $71k – $185k per year (Estimated) • In office • Full-Time • Master's Degree • Toronto
Python
SQL
R
AI/ML
XGBoost
Reinforcement Learning
Scikit-learn
PyTorch
Machine Learning
DevOps
GCP
Git
AWS
Docker
Apply
Immigration Lawyer 9 hours ago
$98k – $117k per year • Hybrid • Full-Time • 2+ years exp • Toronto
Apply
$113k – $135k per year • Hybrid • Full-Time • 2+ years exp • Toronto
AI/ML
LLM
Apply
≈ $47k – $117k per year (Estimated) • In office • Full-Time • 3+ years exp • Quebec • Toronto
Apply
≈ $32k – $62k per year (Estimated) • Hybrid • Full-Time • 2+ years exp • Toronto • Waterloo • Halifax • Montreal
Apply
See all jobs
This is one of many
746,581 more open roles from verified company boards, updated every day.