861,412open jobs
54,280companies
144,969added this week
Browse all
Salary
≈ $82k – $205k per year (Estimated)
Location
Hybrid (Melbourne, Australia)
Seniority
Senior · 1+ year exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 28, 2026. First seen by Alion on Sep 14, 2026. Heidi scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Heidi is an Australian health technology company based in Collingwood, Melbourne, that makes an AI care partner for clinicians, automating clinical documentation, form filling and task management in general practice, emergency departments and specialist clinics. Founded in 2022 as Heidi Health, it reports millions of patient sessions a week in 110 languages, has raised 96.6 million US dollars from investors including Point72, Blackbird and Headline, and supplies NHS trusts and GP practices in the UK. Heidi hires product managers, backend and platform engineers, forward deployed engineers, and customer success, sales and partnerships staff in Melbourne, Sydney, London and New York.

We’re Heidi.

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.

Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart.

We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.

If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi.

The role

This role sits in the model team, the researchers and engineers who train, deploy, and own the AI models behind every Heidi product. You’ll build and operate the infrastructure that makes those models fast, reliable, and cost-effective at scale.

Your work will span production model serving, GPU cluster management, and the infrastructure supporting training and evaluation. You’ll decide how workloads share compute, diagnose performance bottlenecks, and build the deployment and observability tools that help the team move quickly with confidence.

We’re looking for a hands-on engineer who has deployed and operated models in production, understands the demands of GPU workloads, and can take a system from initial design through rollout, incidents, and ongoing improvement. You’ll partner closely with researchers and our platform engineers, with ownership of the systems you build.

What you’ll do

  • Build and own model-serving infrastructure. Take models from checkpoint to production, with repeatable deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts and rollbacks across regions.

  • Manage GPU clusters and workload scheduling. Improve resource allocation across inference, training, and evaluation. Build scheduling policies around workload priority, quotas, hardware topology, and recovery requirements so online services stay responsive while other workloads make productive use of capacity.

  • Improve inference performance. Profile real workloads and improve latency, throughput, and memory efficiency. Evaluate batching, KV-cache management, quantization, speculative decoding, and parallelism strategies against production traffic and quality requirements.

  • Support distributed training and model iteration. Give researchers reliable ways to launch fine-tuning and training jobs, manage model artifacts, save and restore checkpoints, and move validated models into serving. Reduce time lost to failed jobs, slow data loading, and manual setup.

  • Make deployments observable and incidents traceable. Connect application requests and sessions to the exact model, deployment configuration, and worker that served them. Build dashboards and alerts covering model latency, queueing, errors, GPU health, memory pressure, and workload performance.

  • Own production reliability. Define service objectives, investigate incidents across the application, inference engine, GPU, and network layers, and build recovery procedures that work. Turn recurring failures into fixes, automated checks, and useful runbooks.

  • Make compute costs actionable. Track GPU usage, idle capacity, and inference cost by model and workload. Use capacity forecasts and measured performance to guide deployment choices and improve cost per successful request without sacrificing quality or reliability.

  • Build a platform the model team can use independently. Automate provisioning, configuration, benchmarking, and releases. Partner with the engineers behind ASR, note generation, Evidence, and Dictate so new models can be deployed and evaluated through consistent, well-supported workflows.

What you'll need

  • A strong engineering foundation. Hands-on AI infrastructure experience. At least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training. You can design the system, write the code, and own it in production.

  • Production model deployment experience. You’ve deployed and maintained LLMs or other demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems. You understand the work between loading a model and running a reliable service.

  • GPU cluster and orchestration experience. You’ve managed GPU workloads using Kubernetes, Slurm, or an equivalent platform, with practical experience in scheduling, resource allocation, capacity planning, and failure recovery.

  • Performance debugging skills. You can use traces, metrics, and profiling tools to distinguish compute, memory, communication, and scheduling bottlenecks. You understand how batch size, context length, precision, and multi-GPU execution affect performance and cost.

  • Strong software and systems skills. You’re proficient in Python and comfortable with backend or systems development in Go, C++, Rust, or a comparable language. You have practical experience with Linux, containers, deployment automation, and distributed services.

  • Operational ownership. You’ve owned production incidents, built useful monitoring, and made releases recoverable. You can explain the trade-offs behind a design and work effectively with researchers, product engineers, and infrastructure partners.

Nice to have

  • Experience with distributed training frameworks such as PyTorch FSDP or Megatron, or infrastructure for reinforcement learning and rollout generation.

  • Experience tuning inference engines, serving MoE models, or implementing quantization, speculative decoding, and prefill/decode disaggregation.

  • Familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, or diagnosing multi-node communication problems.

  • CUDA or Triton kernel development, contributions to AI infrastructure projects, or experience building cluster operators and scheduling integrations.

  • Experience operating infrastructure across multiple regions or providers, particularly for healthcare or other sensitive production workloads.

How we show up

  • Build for the next decade, not next quarter. Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.

  • Lead, don't wait. We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.

  • Follow the evidence. Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.

  • Own the outcome. Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.

  • Ship, measure, go again. A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.

  • Live in clinicians' reality. Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.

Why Heidi?

You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here.

The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally.

Building what we’re building isn’t always easy. But we didn’t choose easy, we chose to build something that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognise that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it. The work is hard, but you will trust and admire the people you work beside, and rest easy knowing you’re doing the defining work of your career.

We take care of you.

We offer a $1,000 annual learning and development budget, a $150/month health and wellness allowance, a $500 home office budget, 26 weeks paid primary parental leave and 18 weeks paid secondary parental leave, fertility support up to $10,000, four weeks of work from anywhere per year, and serious equity.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
861,412 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Melbourne
≈ $74k – $187k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Sydney
Python
DevOps
Terraform
GCP
CloudFormation
GitLab CI
CI/CD
Jenkins
AWS
Docker
Kubernetes
Platform Engineering
TeamCity
FinOps
Management
Jira
Apply
≈ $70k – $174k per year (Estimated) • In office • Full-Time • 4+ years exp • Sydney
Java
PowerShell
C#
Java
Maven
C#
.NET
DevOps
Terraform
Puppet
Azure DevOps
VMWare
Azure
CI/CD
Windows Server
Kubernetes
SRE
Platform Engineering
Configuration Management
Windows
Management
Agile
Scrum
Kanban
Apply
Production Engineer 2 hours ago
$9.4k – $15k per year (gross) • In office • Full-Time • Kuala Lumpur
Apply
In office
DevOps
Linux
Windows
Apply
In office
DevOps
Azure
Windows Server
Windows
Cybersecurity
Crowdstrike
Apply
≈ $25k – $53k per year (Estimated) • In office • Full-Time • 13+ years exp • Bachelor's Degree • Pune
Python
Java
C++
Perl
DevOps
Linux
Apply
AI Engineer 2 days ago
≈ $52k – $140k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Barcelona
Python
SQL
Databases
Pinecone
AI/ML
LangGraph
Weights & Biases
LangChain
Prompt Engineering
AI Agents
AWS Bedrock
LLM
RAG
OpenAI
Multi-Agent Systems
DevOps
CI/CD
Git
AWS
Docker
Kubernetes
Apply
≈ $70k – $144k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Frankfurt am Main
Python
Java
TypeScript
C#
Java
Spring Boot
C#
.NET
DevOps
GCP
IAM
Apply
Hybrid • Full-Time • London
Go
Go
Chi
AI/ML
Claude
ChatGPT
Marketing
Salesforce
YouTube
LinkedIn
Apply
Founding Engineer 2 days ago
$110k – $180k per year • Equity 0.1–1% • In office • Full-Time • San Francisco
Python
JavaScript
Node JS
AI/ML
Vertex AI
OpenAI
Anthropic
Frontend
Next.js
React.js
DevOps
Azure
Kubernetes
Apply
≈ $96k – $229k per year (Estimated) • Hybrid • Full-Time • Melbourne • Sydney
Python
DevOps
Terraform
Datadog
Prometheus
AWS
Kubernetes
Apply
≈ $96k – $187k per year (Estimated) • In office • Full-Time • London
Python
DevOps
Terraform
Datadog
Prometheus
AWS
Kubernetes
Apply
≈ $102k – $216k per year (Estimated) • Remote (Australia) • Full-Time • Sydney
Apply
≈ $101k – $212k per year (Estimated) • Hybrid • Full-Time • Melbourne
Apply
$180k – $250k per year • In office • Full-Time • 8+ years exp • San Francisco • New York
Apply
In office • Full-Time • 1+ year exp • Melbourne
DevOps
AWS
Linux
Apply
In office • Full-Time • 4+ years exp • Bachelor's Degree • Melbourne
DevOps
AWS
Incident Management
Apply
$113k – $189k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Melbourne
Python
Databases
PostgreSQL
DevOps
Helm
Istio
FluxCD
Azure
GitOps
ArgoCD
AWS
Docker
Kubernetes
Service Mesh
JFrog Artifactory
Incident Management
Management
Jira
ServiceNow
ITIL
Apply
≈ $83k – $209k per year (Estimated) • In office • Full-Time • 5+ years exp • Melbourne
Python
SQL
AI/ML
Amazon SageMaker
Machine Learning
DevOps
Azure
AWS
Apply
≈ $72k – $172k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Melbourne
Python
Java
Ruby
C#
C++
C#
.NET
AI/ML
AI Agents
LLM
DevOps
AWS
TCP/IP
DNS
Cybersecurity
Threat Modeling
Apply
See all jobs
This is one of many
861,412 more open roles from verified company boards, updated every day.