1,174,835open jobs
13,798companies
228,726added this week
Browse all
Salary
$97k – $205k per year (Estimated)
Location
Remote (Canada)
Seniority
Staff · 6+ years exp
Overview
Company
Impact
Profile match
Movable Ink is a New York marketing technology company founded in 2010 that generates personalised content at the moment a message is opened. Its technology assembles images and offers per recipient rather than sending pre-rendered creative. The company works with large retail, travel and financial brands.

Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.

As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross-team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.

Responsibilities:

  • Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents
  • Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives
  • Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization
  • Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios
  • Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery
  • Demonstrate a  high level of  autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

Qualifications:

  • Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy
  • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges
  • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution
  • Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP).
  • Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
  • Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation
  • Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
  • Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
  • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
  • Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting

The base pay range for this position is $154,000-$200,000 CAD/year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job-related knowledge, skills, and experience.

Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.

We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,174,835 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Ontario
.NET Developer 1 hour ago
In office • 3+ years exp
C#
SQL
C#
.NET
Databases
MS SQL
DevOps
AWS
Azure
Azure DevOps
CI/CD
Git
GitLab
Apply
$55k – $142k per year (Estimated) • Remote/Hybrid • Sheffield
DevOps
AWS
Cybersecurity
Microsoft Defender
Microsoft Defender for Cloud
Nessus
Qualys Cloud Platform
Apply
In office • 5+ years exp
DevOps
AWS
Azure
FinOps
GCP
Analytics
Power BI
Apply
In office • 10+ years exp • Bachelor's Degree
DevOps
AWS
Azure
SLI/SLO/SLA
Management
Jira
ServiceNow
Apply
$27k – $70k per year (Estimated) • Remote/Hybrid • 11+ years exp • Noida
JavaScript
Node JS
Python
SQL
TypeScript
Java
Java
Hibernate
Spring MVC
Databases
PostgreSQL
Frontend
JQuery
React.js
DevOps
Amazon CloudWatch
Amazon S3
AWS
AWS Lambda
CI/CD
Docker
IAM
Jenkins
Kubernetes
Platform Engineering
Rest API
Apply
Deal Desk Analyst 3 days ago
In office • 3+ years exp
Management
Google Sheets
Marketing
Salesforce
Apply
$88k – $115k per year • Remote • 5+ years exp • Toronto
Databases
Apache Kafka
Databricks
Kafka
Snowflake
Apply
Operations Analyst 4 days ago
In office • Contractor • 1+ year exp
AI/ML
Claude
OpenAI
Management
Asana
Slack
Apply
Web Manager 4 days ago
In office • 3+ years exp • San José
JavaScript
PHP
PHP
WordPress
Design
Adobe Photoshop
Figma
Webflow
Marketing
GA4
Marketo
Apply
$220k – $287k per year • Remote • 8+ years exp
Databases
Apache Kafka
Kafka
Snowflake
AI/ML
AI Agents
dbt
Model Context Protocol
Analytics
ETL/ELT
Apply
In office • Full-Time • Bachelor's Degree • Ontario
Python
AI/ML
Anthropic
Multimodal AI
Interpretability
Apply
In office • Full-Time • Bachelor's Degree • Ontario
Python
AI/ML
AI Agents
Claude
Fine-tuning
LLM
Multimodal AI
Reinforcement Learning
Synthetic Data
Anthropic
Red Teaming
Interpretability
Web3
Smart Contracts
Apply
In office • Full-Time • Bachelor's Degree • Ontario
Python
AI/ML
AI Agents
Anthropic
Claude
Fine-tuning
Multimodal AI
Reinforcement Learning
Synthetic Data
Interpretability
Web3
Smart Contracts
Apply
In office • Full-Time • 2+ years exp • High School Diploma • Ontario
Management
Outlook
Apply
In office • Full-Time • 2+ years exp • High School Diploma • Ontario
Design
Adobe After Effects
Adobe Photoshop
Adobe Premiere Pro
Apply
See all jobs
This is one of many
1,174,835 more open roles from verified company boards, updated every day.