701,621open jobs
41,200companies
103,639added this week
Browse all
Location
Remote (India)
Employment
Full-Time
Overview
Company
Impact
Profile match
Top of page About Us. At Saaf Finance, we are redefining the future of mortgage lending with AI-driven automation .

MLOps Engineer - AI Platform & Infrastructure

AHL + Saaf AI - Mortgage Lending, Reimagined

Saaf AI is building the future of mortgage lending by combining cutting-edge AI with proven lending operations. Saaf AI is part of American Heritage Lending, a top-10 private lender processing billions in loan volume, with 15+ years of growth. We are backed by some of the largest asset managers and funds and are growing fast.

We’re not experimenting with AI. We’re deploying it. From underwriting to document processing to borrower experience, all are shaped by AI. If you want to work somewhere that uses AI as a core building material - you’re in the right place.

The Role

We’re hiring an MLOps Engineer to own the infrastructure our AI runs on.

Our AI systems are already in production and already touching real loans. The next chapter is making them scale, deploy safely, and change quickly - infrastructure that absorbs real production load, a promotion path that makes shipping boring, and internal tooling that lets the team improve AI behaviour without an infrastructure engineer in the loop for every change.

This is not a role where you inherit a mature platform and keep the lights on. You’ll take AI services that were built to prove the idea and turn them into systems that hold up under load, deploy predictably across environments, and give the rest of engineering a safe, fast path to production. You own the arc from “this works on one machine” to “this is a platform the whole company builds on.”

This role is right for you if:

  • You’ve taken AI or ML systems from a single deployment to infrastructure that scales - and you’ve been on call for the result

  • You treat infrastructure as a product with users, not a ticket queue

  • You think deploys should be boring, reversible, and frequent - and you’ve built the systems that make that true

  • You’d rather remove yourself from the critical path than be the person everyone has to ask

What You’ll Build

Scalable AI Serving Infrastructure

Our AI workloads are moving from early-stage deployment to production scale. You’ll design what they run on:

  • Re-architect how AI services are deployed and run - from single-host setups to horizontally scalable, orchestrated infrastructure that absorbs traffic spikes without degrading

  • Design for the specific realities of LLM-backed workloads: long-running requests, streaming responses, bursty concurrency, expensive downstream calls, and upstream rate limits

  • Own capacity, autoscaling, and unit economics - you should be able to say what we spend per unit of work, and why

Deployment & Environment Promotion

Shipping AI changes should be a routine, reviewable event - not a coordinated risk:

  • Build the promotion path from development through pre-production to production, with environments that are consistent and reproducible rather than each one a special case

  • Make everything version-controlled and reviewable - infrastructure, configuration, and application logic on the same rails

  • Ship CI/CD that gives engineers fast deploys with real rollback, staged rollout, and change history you can audit

Internal Platform & Self-Serve Tooling

The highest-leverage thing you’ll build is the thing that lets other people ship without you:

  • Build internal tooling that lets engineers - and technically-minded teammates outside engineering - define, modify, and test AI workflow logic without touching deployment plumbing

  • Design the guardrails that make that safe: validation, versioning, review, staged promotion, and a clear path back when something is wrong

  • Treat internal users as customers. Success is measured in how many changes ship correctly without you being involved

Reliability, Observability & Operations

You’ll care about whether the system actually holds - not just whether it deployed:

  • Instrument the AI stack end to end: latency, throughput, failure modes, cost, and output-quality signals

  • Build alerting that catches real degradation rather than noise, and incident practice that changes the system instead of assigning blame

  • Own secrets, access control, and environment isolation in a regulated industry where those things carry real consequences

What We’re Looking For

Must-Have

  • Production infrastructure ownership: You’ve deployed and operated containerized services in production under real traffic. You’ve dealt with rollouts, autoscaling, failure isolation, resource limits, and rollback - not just a first deploy.

  • Infrastructure as code: You build environments declaratively and reproducibly. Manually-configured infrastructure makes you uncomfortable, and you know how to migrate away from it without a big-bang rewrite.

  • CI/CD depth: You’ve built pipelines that others depend on daily - automated testing, promotion between environments, safe rollout, and fast recovery.

  • Strong Python: Our AI stack is Python. You should be able to read and change application code, profile it, and fix it - not just package and deploy it.

  • Operational judgment: You can debug a production problem across application, network, and infrastructure boundaries, and you know the difference between a fix and a workaround.

  • Ownership without guardrails: You drive work end to end - design, implementation, deployment, monitoring, and the follow-up when it doesn’t behave.

  • Fast iteration: You ship incrementally and improve. You’re impatient with process that doesn’t reduce risk.

Strong Preferences

  • Experience operating LLM or ML workloads in production - where cost, latency, and non-deterministic output are all live concerns

  • Cloud deployment experience (AWS preferred) - you can containerize, deploy, scale, and operate the systems you own

  • Experience in fintech, lending, insurance, or another regulated industry - secrets management, access control, audit trails, PII handling

  • Built internal developer platforms or self-serve tooling that non-infrastructure engineers actually used

  • Full-stack comfort - you can build a lightweight UI or internal tool when the platform needs one, rather than waiting for someone else to

Nice-to-Have

  • Experience designing multi-environment promotion pipelines where configuration and logic move together with code

  • Familiarity with agent orchestration frameworks and what it takes to run them reliably

  • Observability for non-deterministic systems - tracing, evaluation signals, quality monitoring alongside standard telemetry

  • Inference optimization: model serving, batching, caching, and cost reduction strategies

  • Exposure to workflow automation tooling and low-code builders

Why Saaf

Mortgage AI is still wide open. Unlike ad tech or e-commerce, where AI optimization is a rounding error, in mortgage lending an AI system that works changes whether a family gets their home. The problems are complex, the data is rich, and the solutions don’t exist yet.

The infrastructure problems here are real ones. Non-deterministic workloads, strict regulatory constraints, real financial stakes, and a company shipping AI changes faster than most infrastructure can absorb. This is not maintaining someone else’s platform - it’s designing the one everything else will be built on, at the moment when that decision matters most.

AI-first means AI-first. Every engineer here uses AI to build AI. Claude Code, agentic workflows, AI-assisted code review - we don’t add AI to our process, AI is our process. The team you’ll join is already operating at the frontier.

Direct leverage on everything we ship. Every AI feature this company builds runs on what you build. When deploys get faster and safer, every engineer here gets faster. When the platform scales, every loan we process moves quicker.

Compensation & Benefits

  • Competitive compensation

  • Unlimited PTO

  • Remote-first with flexible hours

  • $2,000/year professional development budget

  • Home office setup stipend

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
701,621 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
SR. AI Engineer 2 days ago
$59k – $147k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Taichung
Python
Go
JavaScript
Rust
SQL
Databases
Snowflake
MS SQL
AI/ML
Model Context Protocol
AI Agents
LLM
Agentic Workflows
Machine Learning
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Apply
$51k – $124k per year (Estimated) • In office • Full-Time • Master's Degree • Taichung
Python
JavaScript
SQL
Python
pySpark
Databases
Snowflake
AI/ML
Spark
AI Agents
LLM
Streamlit
Machine Learning
DevOps
GCP
Apply
Software Engineer 3 days ago
$95k – $183k per year (Estimated) • Remote/Hybrid • Full-Time • Sydney
Python
Go
JavaScript
Java
C#
Node JS
Databases
PostgreSQL
Redis
DynamoDB
MS SQL
AI/ML
Copilot
Claude Code
RAG
DevOps
GCP
Azure
AWS
Apply
In office • 5+ years exp
JavaScript
TypeScript
SQL
C#
C#
.NET
DevOps
GCP
Azure
CI/CD
Git
AWS
Management
Power Automate
Power Apps
Agile
Apply
In office • 15+ years exp • Bachelor's Degree
JavaScript
Java
TypeScript
Java
Maven
Databases
ElasticSearch
Frontend
Angular
JQuery
DevOps
GCP
Kibana
Logstash
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Nginx
Cybersecurity
SonarQube
Fortify
Management
Agile
Scrum
QA
Selenium
Cypress
Apply
Junior Data Engineer 12 days ago
$20k – $55k per year (Estimated) • Remote • Full-Time • 1+ year exp • Bachelor's Degree
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Spark
dbt
DevOps
GCP
Azure
Git
AWS
Docker
Bitbucket
GitHub
Analytics
Tableau
Power BI
ETL/ELT
Fivetran
Looker
Management
Agile
Apply
$28k – $62k per year (Estimated) • Remote • Full-Time • 15+ years exp
Analytics
Microsoft Excel
Management
Microsoft Office
Apply
Remote • Full-Time
Python
AI/ML
Claude Code
Fine-tuning
Prompt Engineering
Function Calling
AI Agents
LLM
RAG
Hallucination
OCR
Human-in-the-Loop
Structured Outputs
LLM Guardrails
Edge AI
Agentic Workflows
Tool Use
DevOps
AWS
Management
n8n
Apply
Senior Product QA 5 months ago
$24k – $56k per year (Estimated) • Remote • Full-Time • 5+ years exp
AI/ML
Copilot
Cursor
Claude
Management
Linear
Trello
Apply
$37k – $97k per year (Estimated) • Remote • Full-Time • 15+ years exp
JavaScript
TypeScript
SQL
Node JS
Databases
MySQL
PostgreSQL
AI/ML
Copilot
Cursor
Claude Code
Prompt Engineering
Frontend
GraphQL
Angular
React.js
DevOps
Terraform
CloudFormation
AWS
AWS Lambda
Analytics
ETL/ELT
Management
Agile
Apply
See all jobs
This is one of many
701,621 more open roles from verified company boards, updated every day.