368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$140k – $306k per year (Estimated)
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
LangChain is an AI technology company that provides open-source frameworks and developer tools for building applications powered by large language models and AI agents. Its ecosystem includes LangChain for agent development, LangGraph for workflow orchestration, and LangSmith for testing, monitoring, evaluation, and deployment. The platform helps developers and enterprises create reliable generative AI applications by connecting models, tools, data sources, and business workflows.

About Us

At LangChain, our mission is to make intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to production-ready AI agents that teams can rely on. We began as widely adopted open-source tools and have grown to also offer a platform for building, evaluating, deploying, and operating agents at scale.

With $125M raised at Series B from IVP, Sequoia, Benchmark, CapitalG, and Sapphire Ventures, we’re at a stage where we’re continuing to develop new products, growth is accelerating, and all team members have meaningful impact on what we build and how we work together. LangChain is a place where your contributions can shape how this technology shows up in the real world.

Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. We have 100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of the Fortune 500 overall), including teams at Klarna, Clay, Coinbase, Workday, Lyft, Cloudflare, Harvey, Rippling, Vanta, LinkedIn, Monday.com, Nvidia, and Bridgewater.

About the team

SmithDB is LangChain's internal database team. We're building a storage and query layer purpose-built for AI observability and evaluation. Within six months we went from idea to a production system that offers industry leading performance and scalability for agent observability data. We're a small, fast team of systems engineers tackling genuinely hard problems: storage layout, query execution, compaction, and scaling toward trillions of agent traces. We develop in Rust, run on Kubernetes, and integrate tightly with S3/GCS/Azure Blob. There are no legacy constraints; this is a greenfield system with real production load and ambitious engineering goals.

About the role

We're building a database specifically designed for AI observability and evaluation, and we need someone to own the infrastructure layer that keeps it running reliably at scale. As a DatabaseInfra Engineer on the SmithDB team, you won't be designing the storage engine - you'll be making sure the engine never goes down, scales seamlessly as our customer base grows, and is operationally excellent across cloud environments.

What you'll do

  • Own the deployment and operations of SmithDB across cloud environments - including cluster lifecycle management, blue/green and rolling upgrades, and automated failover

  • Build and maintain the infrastructure tooling (Terraform, Kubernetes, Helm, or equivalent) that provisions, configures, and scales SmithDB nodes

  • Own the Kubernetes infrastructure that runs our distributed database services (multi-tenant, high throughput, low latency)

  • Build and improve deployment pipelines, rollout strategies, and infrastructure-as-code for the storage layer

  • Drive reliability engineering efforts: incident response, postmortems, SLOs, and disaster recovery for a system operating at massive scale

  • Manage capacity planning and cost efficiency - model growth, rightsize resources, and ensure SmithDB can absorb traffic spikes from our largest customers without manual intervention

  • Build the CI/CD pipeline for database infrastructure changes - safe, tested, and fast promotion from dev through staging to production

  • Collaborate closely with SmithDB internals engineers to translate new engine features into production-ready infrastructure changes and ensure safe, low-risk rollouts

What you'll bring

  • 5+ years of experience in infrastructure, platform engineering, or SRE with hands-on

  • Strong hands-on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure)

  • Solid scripting/systems programming ability (Go, Python, or similar);

  • Experience with infrastructure-as-code and CI/CD tooling (Terraform, Helm, ArgoCD, or similar)

  • Deep familiarity with at least one major cloud provider (AWS, GCP, or Azure) and the primitives used to run stateful workloads reliably - persistent volumes, managed node groups, cloud storage, etc.

  • Infrastructure-as-code fluency - you write Terraform (or Pulumi/CDK) as your primary language, not an afterthought

  • Strong operational instincts - you've been on-call for high-traffic data systems, you know how to triage under pressure, and you write runbooks that actually get used

  • Experience with container orchestration (Kubernetes) and deploying stateful workloads in production

  • A bias for automation - if you've done something manual twice, you're already thinking about how to make it never happen again

  • Strong written and oral communication skills, with the ability to translate infrastructure health into language product and business stakeholders understand

  • The DNA to thrive in a fast-moving, high-autonomy environment - you see gaps as opportunities and own them end to end

Nice to Have

  • Ownership of production database systems (Postgres, ClickHouse, Redis, or similar)

  • Comfort reading and reasoning about Rust is a plus, as it's the language our database is written in

  • Understanding of database reliability concepts - replication, backups, point-in-time recovery, connection pooling, and graceful degradation under load

Compensation

Salary Range: $180,000-$230,000 USD

Compensation Philosophy:

We offer competitive compensation that includes base salary, variable compensation for relevant roles, meaningful equity, benefits, and perks. Actual compensation and offerings will vary based on role, level, and location. Team members in the EU, UK, and APAC receive locally competitive benefits aligned with regional norms and regulations.

Benefits

Benefits include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$230k – $260k per year • Equity • Remote • Internship • Bachelor's Degree
Python
DevOps
Amazon EKS
AWS
Azure
CI/CD
GCP
Helm
Kubernetes
Terraform
Cybersecurity
FedRAMP
Orca Security
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $200k per year • Equity 0.5–5% • In office • Full-Time • 1+ year exp • New York
Python
TypeScript
JavaScript
Python
FastAPI
Databases
DynamoDB
PostgreSQL
AI/ML
Claude
LLM
OpenAI
AI Agents
Frontend
Next.js
Tailwind CSS
React.js
DevOps
AWS
Docker
Vercel
GitHub
Management
Slack
Apply
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
In office • Full-Time • 7+ years exp • Singapore
Python
TypeScript
AI/ML
AI Agents
LangChain
LangGraph
LangSmith
LLM
Prompt Engineering
RAG
Edge AI
DevOps
Amazon EKS
AWS
Azure
Azure AKS
CI/CD
Cloudflare
Datadog
GCP
GitOps
Google GKE
Grafana
Helm
Kubernetes
Platform Engineering
Prometheus
Terraform
Vector
Analytics
A/B Testing
Management
Monday.com
Apply
In office • Full-Time • 7+ years exp • Amsterdam
Python
TypeScript
AI/ML
AI Agents
LangChain
LangGraph
LangSmith
LLM
Prompt Engineering
RAG
Edge AI
DevOps
Amazon EKS
AWS
Azure
Azure AKS
CI/CD
Cloudflare
Datadog
GCP
GitOps
Google GKE
Grafana
Helm
Kubernetes
Platform Engineering
Prometheus
Terraform
Vector
Analytics
A/B Testing
Management
Monday.com
Apply
In office • Full-Time • 7+ years exp • London
Python
TypeScript
AI/ML
AI Agents
LangChain
LangGraph
LangSmith
LLM
Prompt Engineering
RAG
Edge AI
DevOps
Amazon EKS
AWS
Azure
Azure AKS
CI/CD
Cloudflare
Datadog
GCP
GitOps
Google GKE
Grafana
Helm
Kubernetes
Platform Engineering
Prometheus
Terraform
Vector
Analytics
A/B Testing
Management
Monday.com
Apply
$122k – $244k per year (Estimated) • In office • Full-Time • 4+ years exp • Atlanta • Austin • Dallas • Las Vegas • Nashville
JavaScript
Python
TypeScript
AI/ML
AI Agents
Axolotl
Fine-tuning
LangChain
LangGraph
LangSmith
RLHF
TRL
Unsloth
Transformers
DPO
Hugging Face
Post-training
SFT
DevOps
Cloudflare
Management
Monday.com
Apply
$140k – $279k per year (Estimated) • In office • Full-Time • 4+ years exp • New York
JavaScript
Python
TypeScript
AI/ML
AI Agents
Axolotl
Fine-tuning
LangChain
LangGraph
LangSmith
RLHF
TRL
Unsloth
Transformers
DPO
Hugging Face
Post-training
SFT
DevOps
Cloudflare
Management
Monday.com
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.