Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Jun 29, 2026. MangoApps scores B on the Alion truth index.
MangoApps runs an enterprise SaaS platform that thousands of customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform in production, primarily on AWS.
This is a deep individual-contributor role, not a management track. You'll spend your time in the systems: tuning infrastructure, building observability, automating away toil, and leading the technical response when production is on the line. You'll influence how the rest of engineering builds and operates reliable services through your work and your judgment, not through a reporting line.
If you get satisfaction from understanding a production system end to end, finding the real root cause instead of the convenient one, and making the next incident less likely, this role is built for you.
Who you are:
- 5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (not pre-production or internal-only systems).
- You operate AWS at production scale today and can speak in specifics about AWS compute, networking, IAM, EKS/Kubernetes, and the operational realities of running real workloads there. This is a hard requirement.
- A second cloud (Google Cloud or Azure) is a plus, not a substitute. We value it, but AWS depth is what the role is focused on.
- You debug Linux at the level of "why is this latency spike happening," not just "restart the service."
- You reach for automation by reflex. Manual operational work bothers you, and you've built the tooling to remove it.
What will you own:
- Reliability & incident response - Keep the platform available and performant. Define and continuously sharpen monitoring, alerting, and observability. Lead production troubleshooting, drive root-cause analysis, and run post-incident reviews that actually change the system afterward. Participate in the on-call rotation for the services you own.
- AWS infrastructure & operations - Design, deploy, and optimize our AWS infrastructure - compute, storage, networking, DNS, load balancing, and security services. Drive architecture improvements for reliability, scalability, performance, and cost. Own disaster recovery and business-continuity processes, and prove they work before you need them.
- Containers & orchestration - Build and operate containerized workloads on Docker with a focus on security, performance, and predictable scaling across environments.
- Automation & Infrastructure as Code Provision - Build the scripts and tooling that make operations boring and repeatable.
- CI/CD & release engineering - Maintain CI/CD pipelines and deployment automation. Partner with engineering to make releases safer, faster, and easier to roll back.
- Security & compliance - Apply cloud security practices across IAM, network security, secrets management, and vulnerability remediation. Keep infrastructure aligned to our internal security standards and compliance obligations.
Why Join MangoApps?
- You'll have a real impact. We have strong product-market fit, solving communication and collaboration challenges for organizations of every size around the world.
- Our work is recognized. Leading analysts like IDC, Forrester, and Gartner, and independent review sites like Capterra, consistently rate our products highly.
- You'll keep growing. MangoApps is a genuinely collaborative place, and careers here come with steady learning and growth opportunities.
- You'll get to solve hard problems. If complex cloud infrastructure and reliability challenges energize you, you'll feel at home.
- We're flat. We keep the structure lean and treat everyone as an equal.

