429,851open jobs
14,548companies
65,110added this week
Browse all
Salary
$93k – $206k per year (Estimated)
Location
Remote/Hybrid (Toronto, Canada)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Equitable Bank is a Canadian Schedule I bank and the country's seventh largest by assets, held as a wholly owned subsidiary of the listed parent group EQB. It specialises in residential and commercial real estate lending, reverse mortgages, and savings and investment products, and serves personal customers through its digital arm EQ Bank. Founded in 1970 as The Equitable Trust Company and headquartered in Toronto, it manages tens of billions of dollars in combined assets and hires credit, treasury, technology and operations staff in Toronto, Montreal and Vancouver.

Purpose of Job

The Senior AI Platform Operations Engineer is accountable for the reliability, operability, and controlled enablement of the organization’s AI platform.

This role ensures that AI Platform services and solutions are production-ready, secure, observable, and compliant by executing disciplined operational practices, across platform management, monitoring, incident coordination, and governance control enforcement.

The incumbent plays a key role in enabling the safe and scalable adoption of AI by ensuring that AI solutions are deployed, monitored, supported, and continuously improved in line with enterprise standards for reliability, security, and compliance

Main Activities

    AI Platform Reliability and Operations

    • Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure.
    • Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
    • Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience.
    • Track and report on service reliability indicators, incident trends, and operational performance.

    AI Platform Enablement & Production Readiness

    • Enable approved AI use cases into production by ensuring:

    o Environment readiness,

    o Dependency validation,

    o Completion of operational readiness checklists,

    o Structured service transition activities

    • Support platform lifecycle management through:

    o Release coordination,

    o Change readiness validation,

    o Maintenance and capacity planning.

    • Ensure AI platform changes meet defined operational and control readiness criteria prior to release

    Observability, Automation & AI Ops

    • Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
    • Analyze operational data to identify anomalies, recurring issues, root-cause patterns.
    • Implement AI Ops use cases such as:

    o Alert correlation,

    o Anomaly detection,

    o Root-cause support,

    o Forecasting and predictive insights,

    o Automation of repetitive operational tasks.

    • Continuously improve operational efficiency through targeted automation and process optimization.

    Governance, Risk, & Control Execution

    • Execute governance controls for AI solutions, including:

    o Usage and access controls,

    o Data privacy considerations,

    o Auditability and traceability,

    o Human oversight requirements

    • Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
    • Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
    • Identify control gaps and escalate risks appropriately to relevant governance and risk stakeholders.

    AI Asset Visibility & Operational Integrity

    • Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment.
    • Validate asset ownership, relationships, and lifecycle status in collaboration with application and platform owners.
    • Support ongoing audits to ensure AI assets and associated cost attribution remain accurate and current.

Knowledge/Skill Requirements

    • University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
    • 5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
    • Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.

    Technical Expertise:

    • Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
    • Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
    • Strong experience administering Microsoft Power Platform (Power Apps, Power Automate, Dataverse), including environment management, security roles, and solution deployments.
    • Hands-on administration of Copilot Studio, including agent lifecycle management, publishing, monitoring, analytics, knowledge sources, and governance controls. Proficiency with Microsoft 365 Admin Center, user and license management, RBAC, Entra ID groups, and tenant administration.
    • Experience implementing security, compliance, and DLP policies, following least-privilege access principles and operational governance standards.
    • Ability to monitor platform health, manage incidents, optimize licensing and Copilot credit consumption, and support production operations in an enterprise environment.
    • Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
    • Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
    • Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
    • Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.

    Additional Capabilities:

    • Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring.
    • Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
    • Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills.
    • Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
    • Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
    • Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.

    Job Complexities / Thinking Challenges

    • This role requires balancing platform reliability, operational efficiency, and governance discipline in a rapidly evolving AI environment.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
429,851 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Toronto
$23k – $56k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Chennai
Python
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
Databases
PostgreSQL
Oracle
AI/ML
LLM
OpenAI
Amazon SageMaker
Frontend
React.js
DevOps
Rest API
Terraform
Azure
CI/CD
AWS
Docker
Kubernetes
SLI/SLO/SLA
QA
Jest
Apply
$51k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • Buenos Aires
Python
DevOps
GCP
Azure
AWS
Kubernetes
Google GKE
IAM
Management
ServiceNow
Apply
$28k – $63k per year (Estimated) • In office • Full-Time • 7+ years exp • Chennai
Python
PowerShell
AI/ML
Anomaly Detection
DevOps
Platform Engineering
Cybersecurity
Microsoft Sentinel
ISO 27001
Microsoft Defender
NIST CSF
Apply
$23k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • Saint Petersburg
Python
SQL
Databases
PostgreSQL
SQLite
AI/ML
Cursor
Claude
Claude Code
OpenAI
OpenAI Codex
DevOps
Rest API
CI/CD
Git
Docker
GitHub
Management
n8n
Power Automate
Apply
$22k per year (gross) • Remote • Full-Time • Minsk
Go
JavaScript
PHP
TypeScript
SQL
PHP
Symfony
Databases
PostgreSQL
Redis
AI/ML
Copilot
Cursor
Claude
ChatGPT
Claude Code
Embeddings
Prompt Engineering
AI Agents
Gemini
LLM
RAG
OpenAI
Anthropic
Frontend
React.js
DevOps
Rest API
GCP
GitHub Actions
WebSockets
CI/CD
Docker
Kubernetes
Google GKE
GitHub
Apply
$60k – $128k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Master's Degree • Toronto
Python
Java
C#
C++
MATLAB
SAS
Apply
$60k – $129k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Master's Degree • Toronto
Python
SQL
Apply
$35k – $68k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Toronto
Python
SQL
Python
Flask
FastAPI
DevOps
Azure
Apply
$60k – $137k per year (Estimated) • Remote/Hybrid • Contractor • 6+ years exp • Bachelor's Degree • Toronto
Apply
$49k – $128k per year (Estimated) • Remote/Hybrid • Contractor • Bachelor's Degree • Toronto
SQL
SAS
Analytics
Microsoft Excel
Apply
$102k – $132k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Burnaby • Toronto
Python
Marketing
LinkedIn
Apply
$45k – $100k per year • In office • Internship • 3+ years exp • Toronto
Management
ServiceNow
Apply
$80k – $175k per year • In office • Full-Time • 5+ years exp • Master's Degree • Toronto
Python
Apply
$110k – $277k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Toronto
Management
Stripe
Apply
$52k per year • In office • Internship • Toronto
Python
SQL
AI/ML
Cursor
Claude
Claude Code
OpenAI Codex
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
429,851 more open roles from verified company boards, updated every day.