368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$80k – $191k per year (Estimated)
Location
Remote/Hybrid (London, United Kingdom)
Seniority
Principal · 7+ years exp
Overview
Company
Impact
Profile match
commercetools is a Munich company founded in 2006 that builds a composable commerce platform delivered entirely through interfaces. Its approach lets retailers assemble checkout, catalogue and pricing services rather than buying a monolithic suite. The company works with global brands and helped define the composable architecture movement in retail technology.

About commercetools

Real innovation starts with a strong foundation, and at commercetools, that comes from the perfect balance of our product and our people.

Behind every leap forward is a collective of builders, explorers, doers, makers, and problem-solvers.The kind of people who not only pioneered a more flexible approach to commerce architecture but also shaped the culture of experimentation that approach unlocked. Together they are the engine of commerce innovation today.

At commercetools, we power the next era of autonomous commerce for our customers. Whether it’s AI-driven solutions that help enterprises make smarter business decisions, bridging digital and physical shopping experiences, or enabling entirely new ways for industries to connect with their customers, we help the world’s most ambitious companies experiment, scale, and grow without limits.

Here the best idea wins, not the loudest voice. You will have the tools, trust, and space to not only build the future of commerce, but to build your own.

Your Impact

As the Principal Engineer championing Resiliency, you'll be the driving force behind how commercetools prepares for, responds to, and learns from operational incidents at scale. Our customers rely on us for mission-critical commerce infrastructure, including during their highest-stakes moments of the year, like Black Friday. You'll own the discipline of resiliency end-to-end: mature incident management, strong operational visibility, data-driven process improvement, and organization-wide readiness for peak-traffic events.

  • Standardize Incident Management:  Build and champion intuitive, end-to-end processes for incident detection, response, communication, and postmortems across the company.
  • Enhance System Visibility:  Develop clear, real-time metrics, dashboards, and signals to track system health and incident trends.
  • Drive Data-Backed Improvements:  Use operational data to find process gaps, partnering with product engineering teams to fix them.
  • Own Peak-Event Readiness:  Scale and lead the organization-wide readiness program for massive traffic spikes like Black Friday.
  • Lead Cross-Team Initiatives:  Identify resilience gaps, collaborate with infrastructure/product teams on solutions, and turn ideas into concrete technical outcomes.
  • Cross-Functional Collaboration:  Partner closely with engineering leadership, Staff Engineers, and domain-specific Principal Engineers (Cloud, Security, API, Architecture, Performance).
  • Foster Knowledge Sharing:  Drive organizational communication, documentation, and training around resiliency and operational excellence.

This role is hybrid, with three days a week spent in our Berlin, London, Munich or Valencia office.

What Sets You Apart

You're a creative problem-solver who is wired to find solutions. You confidently dive into complex challenges and have a talent for making them simple for others. Your curiosity drives you to constantly grow and contribute to an environment of trust and teamwork. Great ideas come from many paths, and your unique perspective matters more than checking every box. What matters most is the mindset you bring to the work.

You bring:

  • Experience:  7+ years  driving incident management/operational excellence; 5+ years  leading org-wide resiliency and reliability initiatives.
  • Peak Load Execution:  Proven track record of managing and scaling systems for high-stakes, multi-team operational events (e.g., Black Friday, major launches).
  • Data Literacy:  Strong ability to analyze metrics to diagnose technical issues and measure process improvements.
  • Leadership & Influence:  Ability to evaluate both technical bugs and the organizational/human dynamics behind them.
  • Project Management:  Demonstrated success managing large-scale initiatives that span multiple engineering teams in an Agile  environment.
  • Communication:  Fluent English with exceptional written and verbal communication skills; experience running technical training or onboarding is a plus.
  • Soft Skills:  High self-awareness, a strong customer focus, and a passion for mentoring others and learning new technologies.

Our Benefits

Because work and life are connected, our benefits are too. We’ve designed them to give you the security, flexibility, and opportunities you need to focus on what matters most.

 Comprehensive health benefits  for you and your dependents, including access to OpenUp for personalized mental health support

 Learning and development  opportunities including an annual learning budget, access to self-paced learning platforms and language training, personalized coaching, mentorship, and leadership programs

 Family Leave Plus  gives you additional fully paid weeks of parental leave on top of government-provided leave, so you can spend more time with your new addition

 Our equity participation program  allows you to share in our success 

For more information on our benefits, visit this page.

Come as you are. Build with us.

Your unique perspective is essential to our success. We are committed to building a team that reflects the world around us because we know it’s the only way to build the future. We celebrate our differences and have created a hiring process that’s fair, inclusive, and designed to let your talent shine.

We proudly welcome applicants of every race, color, religion, gender identity, sexual orientation, age, and any other part of your identity that makes you who you are. As an equal opportunity employer, we believe that our strength lies in our diversity, and we invite you to be a part of our global community. 

For more information on our diversity, equity, inclusion, and belonging practices,  visit this page

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$80k – $169k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$112k – $217k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$42k – $106k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$175k – $195k per year • Remote • 8+ years exp
Python
Python
pySpark
AI/ML
Spark
Edge AI
DevOps
AWS
CI/CD
Incident Management
Platform Engineering
Amazon S3
Analytics
ETL/ELT
Apply
$69k – $163k per year (Estimated) • Remote/Hybrid • 7+ years exp • Munich
DevOps
Incident Management
Apply
$69k – $163k per year (Estimated) • Remote/Hybrid • 7+ years exp • Berlin
DevOps
Incident Management
Apply
$52k – $123k per year (Estimated) • Remote/Hybrid • 7+ years exp • Valencia
DevOps
Incident Management
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
$87k – $159k per year (Estimated) • In office • Contractor • 5+ years exp • London • Stockholm
Design
Figma
Marketing
Zendesk
Apply
$105k – $204k per year (Estimated) • In office • Full-Time • London
Python
SQL
Databases
Snowflake
AI/ML
Dagster
dbt
Analytics
A/B Testing
Apply
$27k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • London
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
$34k – $85k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
AI/ML
AI Agents
Edge AI
QA
Appium
Cucumber
Cypress
Selenium
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.