607,540open jobs
29,092companies
85,853added this week
Browse all
Salary
$45k – $122k per year (Estimated)
Location
Remote (Singapore)
Employment
Full-Time
Overview
Company
Impact
Profile match
Braintrust is a San Francisco company founded in 2023 that provides evaluation and observability tooling for teams shipping large language model products. Its platform stores prompts and datasets, scores model outputs against them and tracks regressions as models and prompts change, turning AI development into a measurable engineering loop. It is used by software companies that need to compare providers and validate agents before release.

About the company

Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.

Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.

About the role

Our largest customers don't just use Braintrust - they run it. They deploy our stack inside their own AWS, Azure, and GCP accounts, behind their own VPCs, under their own compliance requirements, at their own scale. When a hybrid deployment stalls, when ingest backs up, when a query that was fast last week isn't, they come to us. Platform Support is the team that owns that. We're the technical front line for infrastructure, performance, and reliability.

We're hiring Platform Support Engineers at both mid and senior levels to join a small, high-ownership team. You'll work shoulder to shoulder with our Cloud Infrastructure and Engineering teams, and alongside our Developer Support Engineers, who own the SDK and API side of the customer experience. If you like hard infrastructure problems, and you like them more when a real customer is on the other end, this is the role.

What you'll do

  • Own customer-facing support for hybrid and self-hosted Braintrust deployments across AWS, Azure, and GCP - from first install through steady-state operation.
  • Debug real infrastructure problems: Kubernetes workloads, Terraform state, networking and VPC configuration, IAM and permissions, TLS, and cloud-provider quirks.
  • Diagnose performance and reliability issues in the backend - ingest throughput, query latency, database and object-store behavior - using logs, metrics, and traces to get to cause rather than symptom.
  • Lead incident response for customer-impacting issues: triage, communicate clearly while it's still on fire, and drive it to resolution.
  • Ship fixes. Submit PRs to our backend services, Terraform modules, and deployment tooling rather than handing every problem to Engineering.
  • Build the tooling that makes the next one easier - diagnostics, health checks, preflight validation, and self-service paths that let customers unblock themselves.
  • Write and maintain the runbooks and deployment documentation that turn one hard-won answer into a permanent one.
  • Feed patterns back to Engineering and Product, so the recurring failure modes stop recurring.
  • Participate in an on-call rotation for critical customer issues.

What we're looking for

  • Experience in a customer-facing technical role - Support Engineering, SRE, DevOps, Solutions Architecture, or Infrastructure Engineering - or backend/infra engineering experience with real appetite for customer work.
  • Strong Kubernetes fundamentals: you can deploy, debug, and scale actual workloads, and read a failing pod's story from its events and logs.
  • Hands-on Terraform, and depth in at least one major cloud (AWS strongly preferred).
  • Comfort in a backend codebase - Python, TypeScript, or Go - enough to reproduce a bug, trace it to its source, and fix it.
  • Fluency with observability tooling, and the instinct to reach for data before opinion.
  • Clear, calm, direct communication under pressure, especially when the customer is technical, blocked, and losing time.
  • Ownership. You take a problem personally and follow it until the customer is running again.

Bonus points for

  • Supporting self-hosted or on-prem enterprise software, especially in regulated environments.
  • Multi-cloud experience, particularly Azure or GCP alongside AWS.
  • Database and data-infrastructure depth - Postgres, ClickHouse, or similar analytical stores.
  • Experience with observability, ML infrastructure, or developer platforms.
  • Familiarity with LLM APIs and how teams are building and evaluating agents in production.
  • Having built support or diagnostic tooling that measurably reduced ticket volume.

Why join Braintrust

  • Work on genuinely hard infrastructure problems, at the scale and pace of the teams building the best AI products in the world.
  • Join a team early enough to shape how it operates - its standards, its tooling, and its bar.
  • Sit close to both the customer and the code, with the mandate to fix things in either direction.

Benefits include

  • Medical, dental, and vision insurance
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity
  • Wifi & cellphone stipend

Equal opportunity

Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
607,540 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
Remote/Hybrid • 7+ years exp
Python
Java
SQL
AI/ML
Copilot
Claude Code
Model Context Protocol
AI Agents
LLM Guardrails
Agentic Workflows
DevOps
Terraform
GCP
Management
Google Drive
Agile
Apply
$191k – $262k per year • Remote/Hybrid • Bachelor's Degree • Kitchener
Python
JavaScript
Java
Kotlin
TypeScript
Java
Hibernate
Databases
MySQL
CockroachDB
AI/ML
Cursor
Claude
LLM
Frontend
React.js
Mobile
JUnit
DevOps
CI/CD
AWS
GitHub
Management
Linear
Apply
In office
Python
Bash
Databases
Redis
DevOps
Terraform
CloudFormation
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Bitbucket
Amazon EKS
AWS Lambda
Amazon EC2
Amazon S3
IAM
Amazon ECS
Cybersecurity
SOC 2
Management
Confluence
Jira
Apply
In office • 8+ years exp
Python
JavaScript
TypeScript
Python
FastAPI
AI/ML
Model Context Protocol
Prompt Engineering
Multimodal AI
RAG
OpenAI
Agentic Workflows
Frontend
React.js
DevOps
Rest API
Azure
CI/CD
Apply
$196k per year • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
Python
SQL
DevOps
Docker
Apply
$43k – $100k per year (Estimated) • Remote • Full-Time • London
Python
TypeScript
AI/ML
Prompt Engineering
LLM
Braintrust
OpenAI
Anthropic
DevOps
Terraform
Kubernetes
Cloudflare
Management
Slack
Stripe
Apply
Remote • Full-Time
Python
Go
TypeScript
Databases
PostgreSQL
ClickHouse
AI/ML
LLM
Braintrust
OpenAI
DevOps
Terraform
GCP
Azure
AWS
Kubernetes
Cloudflare
IAM
Management
Stripe
Apply
$140k – $190k per year • Remote • Full-Time • Singapore
Python
TypeScript
AI/ML
Prompt Engineering
LLM
Braintrust
OpenAI
Anthropic
DevOps
Terraform
Kubernetes
Cloudflare
Management
Slack
Apply
$68k – $154k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
Apply
GTM Engineer 1 day ago
$100k – $150k per year • In office • Full-Time • 2+ years exp • Singapore
AI/ML
Reinforcement Learning
Post-training
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
AI Agents
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
LLM
Post-training
DevOps
Docker
Apply
See all jobs
This is one of many
607,540 more open roles from verified company boards, updated every day.