DexCare is looking for a Site Reliability Engineer II who will report to the Head of Software Reliability Engineering. We are looking for an engineer who is passionate about the foundation blocks for a data science team. Candidates must be able to demonstrate how they have enabled data science and engineering teams to quickly experiment with data and how their work influenced the ability of their broader organisation to iterate quickly in exploring machine learning models. As a site reliability engineer (data and infrastructure) within DexCare, you will be responsible for partner changes and the data tasks and business integrations to deliver best-in-class tooling for exploring data. With DexCare being an early-stage startup, candidates must be willing to quickly adapt and learn about software changes and customer asks and be curious to take on new challenges as they present themselves.
Requirements:
- Candidates must be able to demonstrate how they have enabled data science and engineering teams to quickly experiment with data and how their work influenced the ability of their broader organisation to iterate quickly in exploring machine learning models.
- DexCare is looking for a candidate that has 5/6+ years of experience in computer science or a related field/senior SRE or DevOps role supporting production cloud infrastructure at scale.
- Candidates need to have working knowledge of Python, TypeScript, Node.js, or JavaScript or another modern programming language with a keen interest to quickly come up to speed on TypeScript.
- Candidates must have experience in administering, deploying and maintaining SQL services and their databases.
- You will need to have a demonstrated ability to deliver results as a site reliability engineer (data and infrastructure).
- Candidates will have an understanding of how software systems are built.
- Deep experience with AWS (IAM, EKS, VPC, EC2 Secrets Manager, Serverless) and RBAC.
- Hands-on proficiency with Terraform, Terragrunt, Helm, and container orchestration.
- Proven experience building and maintaining GitHub Actions for CI/CD, including GitHub Advanced Security features like secret scanning and code policy enforcement.
- Strong Datadog experience building dashboards, tuning alerts, setting up monitors, and interpreting telemetry.
- Solid Python scripting experience for automation and internal tools.
- Comfortable working in Agile/Scrum environments with well-tracked Jira workflows.
- Practical experience with resource analysis and infrastructure optimisation.

