We are looking for a Senior DevOps Engineer to join our team!
You will own and evolve
Mailtrap's production infrastructure - a multi-region AWS platform powering high-volume email sending and testing. You will keep it reliable, secure, and cost-efficient day-to-day.
You will also lead a key strategic initiative: planning and executing a partial migration of selected workloads from AWS to rented bare-metal / colocation infrastructure to cut costs - a hybrid-by-design effort, not a full cloud exit. You will decide what moves and what stays, design the target platform (likely Kubernetes or a similar orchestrator on bare metal), and own the migration end-to-end, from business case to cutover.
Our engineering team thrives in an agile, continuously improving, and automation-oriented environment. We value ongoing evolution, objective evaluation of our processes, and taking action to make things better.
What are we looking for?
Must-have skills:
5+ years DevOps/platform (senior), production ownership
Deep AWS (VPC, ECS, IAM, RDS/ElastiCache or equivalents, networking)
Terraform at scale (modules, remote state / Terraform Cloud)
CI/CD with GitHub Actions, including automated deploys to production (e.g. blue/green deployments)
Containers (Docker); comfortable operating services on ECS or K8s
Strong Linux networking (DNS, TLS, load balancing, firewalls, VPN/hybrid)
Proven experience planning and executing migration to on-prem, colo, or private cloud (partial/hybrid OK - full exit not required)
Comfortable owning a multi-quarter infra initiative: TCO, design, vendors, cutover, rollback
Comfortable automating operational tasks with shell and/or Python
Fluent English (both spoken and written)
Strongly preferred:
Colo/bare-metal ops: IPAM (NetBox), Ansible, image-based provisioning (Packer or equivalent), HAProxy/Nginx
Replacing managed AWS services with self-hosted (Postgres HA, Redis/Valkey, Kafka, OpenSearch)
Email infrastructure (MTA, SMTP, IP reputation, DKIM/SPF/DMARC) - Halon or similar
Cloudflare (DNS/WAF/Access)
Observability beyond CloudWatch (Prometheus/Grafana/Loki or equivalent)
Prior work with multi-region SaaS or EU data residency
Cost-driven architecture / FinOps mindset
Would be a plus:
GCP (BigQuery/certificates)
Comfortable reading Ruby or Go (used in our services and tooling)
AWS Certificates
Responsibilities
Day-to-day:
Operate and evolve AWS multi-account / multi-region infra
Terraform modules/workspaces
Ensure safe infrastructure changes across network, storage, and services, with zero-downtime deployments.
ECS services, blue/green deploys, Docker image pipelines
Reliability: CloudWatch/PagerDuty/Sentry, capacity, cost tags
Security baseline: IAM, secrets (SSM), Cloudflare edge rules
Partner with engineers on release automation and production readiness
Maintain the hybrid estate (AWS and rented bare metal) as one operable platform
Migration leadership:
Build the business and technical case for what moves off AWS vs what stays
Design the rented bare-metal / colo landing zone (compute, network, storage, observability, secrets)
Produce migration waves, dependency maps, cutover/rollback plans
Stand up hybrid connectivity and dual-run periods; shift traffic safely (e.g. via Cloudflare)
Replace or re-home managed services where it pays off (compute, queues, search, cache, MTA nodes)
Coordinate colo/vendors, timelines, and eng teams; report progress and risk

