We are seeking a highly experienced and motivated Lead DevOps-AI Platform Engineer to join our technology organisation. This role is responsible for ensuring the reliability, security, scalability, and operational excellence of our AWS-based infrastructure while also supporting next-generation agentic AI platform capabilities. The ideal candidate will have deep expertise in Linux, AWS, automation, CI/CD, troubleshooting, and production operations, along with hands-on experience in multi-agent frameworks, retrieval-augmented generation, observability, evaluation, and AI infrastructure tooling. This position requires a strong blend of infrastructure engineering, DevOps leadership, platform automation, and familiarity with modern AI agent orchestration and governance technologies.
Responsibilities:
- Ensure the organisation's AWS infrastructure is available, secure, and operating reliably 24x7
- Design, build, maintain, and optimise cloud infrastructure across load balancers, containers, databases, and related platform services.
- Lead infrastructure migration, deployment, and launch activities, ensuring smooth transitions and minimal disruption to business operations.
- Partner closely with DevOps, Security, and Engineering teams to maintain application integrity, infrastructure security, and operational resilience.
- Build and maintain deployment tooling, operational services, and microservices that improve efficiency, reduce manual effort, and enhance customer service standards.
- Troubleshoot and resolve incidents across development, testing, staging, and production environments in a timely and proactive manner.
- Validate system integrity, infrastructure designs, application deployments, and operational processes, recommending improvements where appropriate.
- Develop, update, and document operational procedures, technical workflows, and platform standards.
- Automate operational and infrastructure processes with an emphasis on accuracy, repeatability, compliance, and security.
- Specify, document, and support the development of new platform features, scripts, and operational capabilities.
- Manage code deployments, releases, fixes, updates, and related operational processes.
- Work with open-source technologies, cloud services, and modern CI/CD tooling to support reliable delivery pipelines.
- Collaborate effectively with teams using Git, Agile methodologies, and workflow management tools such as Jira, Workfront, Scrum, Kanban, or SAFe.
- Support the design, implementation, and operationalisation of multi-agent AI systems and related enterprise workflows.
- Build integrations for autonomous agents that interact with external APIs, databases, and enterprise systems through function calling and secure tool usage.
- Contribute to the orchestration of multi-agent workflows using frameworks such as CrewAI, LlamaIndex, LangGraph, AutoGen, Semantic Kernel, Haystack, and DSPy.
- Support retrieval-augmented generation (RAG) systems and scalable vector database architectures, including Qdrant, Chroma, pgvector, Pinecone, and Weaviate.
- Implement observability, tracing, and evaluation for AI and agentic systems using tools such as LangSmith, Langfuse, Arize AI, Phoenix by Arize, Helicone, and MLflow.
- Support Model Context Protocol (MCP) integrations to enable secure and standardised agent connectivity.
- Define and enforce agent permissions, guardrails, and governance controls for autonomous systems.
- Support AI gateway and routing solutions such as Portkey, LiteLLM, and Envoy AI Gateway.
- Contribute to GPU scheduling and infrastructure planning for model-serving and agentic workloads.
- Promote operational best practices for autonomous systems management, reliability, evaluation, and continuous improvement.
Requirements:
- 10+ years of experience as a DevOps Engineer, Platform Engineer, Infrastructure Engineer, or in a similar role, with additional software or infrastructure development experience preferred.
- Strong experience with Linux-based infrastructures, Linux/Unix administration, and AWS.
- Strong experience with databases and data platforms such as SQL, MySQL, NoSQL, Elasticsearch, Redis, and MongoDB.
- Proficiency in one or more scripting or programming languages such as Java, JavaScript, Perl, Ruby, Python, Golang, PHP, Groovy, or Bash.
- Experience with CI/CD tools, Git, deployment automation, and release engineering practices.
- Experience with open-source technologies and cloud services.
- Experience with configuration management and automation tools such as Puppet or Chef.
- Familiarity with Agile delivery methodologies and workflow tools such as Jira, Workfront, Scrum, Kanban, or SAFe.
- Strong communication skills, with the ability to explain technical protocols, processes, and decisions to both technical and non-technical stakeholders.
- Strong troubleshooting and analytical skills, with the ability to identify issues before they become problems.
- Demonstrated ability to stay current with industry trends, IT operations best practices, and emerging platform technologies.
- Experience with agentic AI platforms, multi-agent frameworks, observability, RAG, and related infrastructure is strongly preferred.
- A bachelor's or master's degree in computer science, engineering, software engineering, or a related field.
Preferred Qualifications:
- Experience with multi-agent orchestration frameworks such as CrewAI, LangGraph, AutoGen, or Semantic Kernel.
- Experience with AI observability and evaluation platforms such as LangSmith, Langfuse, Arize AI, Phoenix by Arize, Helicone, or MLflow.
- Experience with Model Context Protocol (MCP) and secure agent-tool integration patterns.
- Experience with vector databases and retrieval architectures used in RAG-based applications.
- Experience with AI gateways, policy enforcement, and operational guardrails for agentic systems.
- Exposure to GPU scheduling, distributed inference, or AI platform operations.

