Responsibilities Conduct systematic data audits of existing simulation data including schema assessment, volume, cleanliness, and gaps; define supplementary data generation requirements Build and maintain data pipelines for model training, validation, and continuous retraining Build multi-domain model pipelines that chain individual surrogate models without manual handoff Develop training pipelines, architecture, and prototyping for ML algorithms Work on productising research prototypes Conduct experiments to benchmark new techniques and evaluate model behavior Develop systematic evaluation methodology: test sets, accuracy metrics, citation quality scoring, false positive/negative analysis Deploy AI tools to engineering teams with structured pilots, baseline measurement, and documented adoption outcomes
Requirements Required Qualifications 3+ years ML engineering with a focus on deep learning for scientific or engineering applications Experience training regression/emulation models on physics or simulation data (surrogate modelling or reduced order modelling) Strong ML stack: PyTorch or TensorFlow, Pandas, NumPy, SciPy Surrogate modeling via Neural Networks or Gaussian Processes for use as fast-running model proxies. Proven understanding of fundamental data structures and the ability to apply them to solve complex problems. Development experience with retrieval pipeline skills and relational databases
Preferred
Qualifications Understanding and deployment of Reinforcement Learning based tools Understanding of mathematics, particularly linear algebra and probability theory Experience with physics-informed neural networks (PiNNs) or hybrid physics-ML models Experience with multi-fidelity modelling or chained model pipelines Modeling complex multi-physics systems of ODEs and DAEs Gradient-based optimization Automatic differentiation tools and development

