Responsibilities Build and maintain data pipelines for model training, validation, and continuous retraining Instrument, monitor, and manage the operational health of the AI/ML capabilities in production including performance, drift, latency, and data quality. Build CI/CD pipelines, model versioning, rollback procedures, and A/B testing infrastructure Manage the container orchestration layer and server-side configurations to ensure high availability for internal tools and product operations suites Lead the feasibility assessment and the implementation of on-premise AI deployment Own the qualification and governance documentation process across data provenance, model architecture, explainability, and human oversight procedures, robustness testing with applicable regulatory frameworks Establish and operate a recurring governance review process across all deployed capabilities
Requirements Required Qualifications 4+ years MLOps or ML infrastructure engineering with production systems experience CI/CD design and implementation for ML systems (MLflow, DVC, Weights & Biases, or equivalent) Model monitoring: drift detection, performance degradation alerting, data quality checks Deep expertise in Kubernetes cluster management, service mesh, and cloud/on-prem hybrid infrastructure Proven experience in setting up CI/CD pipelines for both non-ML software and ML models Containerised ML deployment (Docker, Kubernetes)
Preferred
Qualifications Experience deploying ML systems in Safety-critical or regulated domain background where AI output quality must be explainable Familiarity with change management processes in regulatory environments RAG system infrastructure experience at scale Cloud compute job queue management for computationally intensive workloads On-premise ML/AI deployment experience: quantisation, inference optimisation, GPU cluster management

