We are seeking a versatile Lead DevOps and Platform Engineer with 8-12 years of overall experience who brings a strong blend of core application engineering and cloud/infrastructure platform leadership.
This role requires someone who has built and shipped software as a core application developer (in runtimes like Python, Java, Go, or Node.js ) and later transitioned into or spearheaded a DevOps/Infrastructure charter. You will bridge the gap between application development and cloud operations, ensuring our AI agents and SaaS platform are architected for extreme scale, resiliency, and speed of delivery. If you are a hands-on leader who can "lead by example, " write clean code, and design high-availability infrastructure from the ground up, we'd love to hear from you.
The candidate will have responsibilities across the following functions:
Cloud Infrastructure and Platform Architecture:
- Design, implement, and manage scalable cloud infrastructure primarily on AWS (with exposure to GCP/Azure).
- Build ground-up multi-region network architectures for enterprise workloads, securing networks into layered trust zones (subnets, VPC peering/privatelinks, NAT gateways at scale).
- Architect high-availability (HA) and disaster recovery (DR) setups aligned with strict RPOs and RTOs.
- Design compute, storage, and networking layers for complex, high-throughput systems (microservices, data pipelines, LLM inference endpoints, and vector databases).
- Automate infrastructure management end-to-end using Infrastructure-as-Code (IaC) tools (e. g., Terraform, CloudFormation).
Containerization and Orchestration:
- Set up, monitor, and manage large-scale Kubernetes (EKS/GKE) clusters for seamless deployment of microservices and AI workloads.
- Manage persistent volume configurations for dynamic pods at scale.
- Implement best practices for networking policies, service meshes, and ingress control within Kubernetes.
Engineering and Technical Leadership:
- Leverage your application engineering background (Python, Java, Node.js, Go, etc. ) to review developer architectures, optimise backend performance, and build internal platform tooling.
- Act as a bridge between core product engineering and infrastructure, ensuring applications are designed natively for auto-scaling and cloud reliability.
- Provide technical mentorship to engineers on deployment strategies, distributed debugging, and cloud-native development practices.
Performance Monitoring and Troubleshooting:
- Implement robust observability and telemetry frameworks (APM, distributed tracing, metrics, logging) to track system health and AI model runtime performance.
- Lead root-cause analysis (RCA) for critical incidents and drive proactive system fixes to eliminate downtime.
CI/CD and Security:
- Oversee and optimise existing CI/CD automation and deployment workflows to maintain high developer velocity.
- Integrate basic security scanning (SAST/DAST) and ensure general alignment with security and compliance standards.
Requirements:
- Experience: 8-12 years total experience in software engineering and infrastructure management, ideally within early-stage to growth-stage startups or AI/SaaS companies.
- Dual Engineering Background: Proven track record having spent time as a Lead Backend/Software Engineer (Python, Java, Node.js, Go) before leading a DevOps / Infrastructure / Platform Engineering charter.
- Cloud and K8S Mastery: Heavy hands-on experience designing HA/DR cloud architectures on AWS/GCP and orchestrating Kubernetes clusters at scale (CKA certification is a plus).
- Startup Mindset: Comfortable wearing multiple hats, working in fast-paced environments, and driving systems from 0 to 1 and 1 to 10
- Leadership: Strong hands-on technical leader who "leads by example"capable of diving deep into code, network configs, and architectural whiteboarding.

