We're looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premises environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You'll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security. This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Responsibilities:
- Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, and security groups).
- Administer and harden AlmaLinux VMs across production, staging, and dev environments.
- Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling).
- Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch.
- Lead incident response diagnosis, RCA, and post-incident documentation with no dedicated ops team to escalate to.
- Make and document build-vs-buy and architecture decisions as the product and team scale.
- Work directly with founders/engineering to translate ambiguous asks into scoped technical plans.
- Accelerate engineering velocity through scalable developer platforms and automation.
- Improve deployment reliability, platform uptime, and operational efficiency.
- Enable secure and scalable AI-driven cybersecurity workloads.
- Reduce operational overhead through infrastructure automation and self-service systems.
- Help establish enterprise-grade cloud and on-premises deployment capabilities.
- Enhance product resiliency, observability, and operational excellence.
- Shape the long-term platform architecture powering next-generation cybersecurity products.
- Enable rapid and secure delivery of critical security innovations to customers.
Requirements:
- 4+ years of hands-on Linux administration (RHEL-family strongly preferred: AlmaLinux, CentOS, RHEL).
- Deep Linux internals: systemd, networking, storage/LVM, process/resource management, and kernel-level troubleshooting.
- Real AWS architecture experience, not just operating existing infra, but designing it (VPC, EC2 IAM, security groups, networking).
- Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to.
- Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts built into tooling that runs unattended.
- Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one).
- Clear, proactive communicator who documents decisions and explains reasoning without being asked.
- Strong Linux system administration and troubleshooting skills.
- Red Hat certifications.
- Strong understanding of networking fundamentals, security, and distributed systems.
- Proficiency with Docker and container orchestration.
- Experience with Terraform, Ansible, or similar infrastructure automation tools.
- Strong scripting or programming skills in Python, Bash, or Go.
- Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry.
- Understanding of platform security best practices and secure infrastructure design.
- Familiarity with virtualization technologies and hybrid infrastructure environments.
- Strong problem-solving and debugging abilities.
- Excellent communication and collaboration skills.
- Ability to thrive in fast-paced startup environments.
Nice to Have:
- Configuration management/automation at scale (Ansible, AWX, Terraform, or similar).
- Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog).
- Container experience (Docker; Kubernetes is a plus but not core to this VM-based stack).
- Experience in a startup or small-team environment where infrastructure was built from scratch.
- Security/compliance exposure (vulnerability remediation, hardening, SSO/access control).
The Mindset:
- Problem Solver: You thrive on complex, ambiguous challenges and engineer elegant solutions.
- Ownership-Driven: You take initiative, move fast, and deliver outcomes without hand-holding.
- Continuous Learner: You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
- Startup DNA: You excel in fast-moving environments where priorities evolve, and impact is immediate.

