746,022open jobs
44,798companies
108,231added this week
Browse all
Salary
≈ $106k – $198k per year (Estimated)
Location
Hybrid (London, United Kingdom)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 22, 2026.

Overview
Company
Impact
Profile match
Writer is an enterprise generative artificial intelligence company headquartered in San Francisco, California, and founded in 2020. The company provides a full-stack platform that includes proprietary large language models, an integrated graph-based retrieval-augmented generation system, and tools for building autonomous AI agents. It serves Fortune 500 companies across sectors such as financial services, healthcare, and retail, focusing on delivering secure and compliant AI applications for marketing, sales, and operations.

About WRITER

WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible - through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI.

Founded in 2020 with office hubs in San Francisco, New York City, Seattle, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI.

About the role

At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering.

What you'll do

Technical

  • Breadth across disciplines. Bring deep focus to one problem at a time, with the breadth to move between SRE, DevOps, Infrastructure, and Platform work over a quarter or two as the leverage shifts. This is not a thrash-every-week role - most of the time you're heads-down on one substantial initiative (the on-call posture, the release pipeline, the multi-region Terraform layout, the internal platform surface). Cross-layer fluency is what lets you pick the right next initiative; it isn't a weekly context-switch.

  • Simplicity / via negativa. Challenge the status quo and remove toil before adding features - automate operational tasks and infrastructure management with Python or Go, reject tools that don't fit the problem, and treat manual on-call work as a defect to be designed out, not a status quo to be staffed up.

  • Breadth across the stack. Design scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently across Kubernetes, Helm, Terraform, and the supporting cloud and AI tooling that backs WRITER's high-traffic platform.

  • AI in workflow. Run agents in your daily loop - Claude Code, Droid, Codex, internal skills - to investigate incidents, draft Terraform / Helm changes, write runbooks, scaffold tooling, and review PRs. Build the agentic setup as a collective surface: humans and digital teammates working as one team, with shared skills, shared context, and shared on-call workflows. Encode recurring infra tasks as internal skills any teammate (human or agent) can pick up and run, so the team's throughput compounds - not just your own.

  • Debugging fluency. Lead incident response, post-mortems, and root-cause analyses - trace failures to the underlying problem (never the symptom), apply the learning back into the architecture, and prevent the same incident from happening twice.

Non-technical

  • End-to-end ownership. Own the reliability, performance, and efficiency of WRITER's core services end-to-end - define and uphold the SLOs and error budgets, carry the on-call pager, and stand behind the outcome metric, not just the system you shipped.

  • Strategic vs. tactical balance. Balance this week's critical work with the 6-12-month platform direction - ship the on-call-driving fix today while shaping the multi-year observability, cost, and reliability investments that move WRITER's enterprise customers.

  • Cross-functional collaboration. Operate at the seams with product, security, and engineering peers - provide expert guidance on system design for reliability, performance, and scalability from conception through launch, Connect the infra agenda to product and revenue context, and disagree with evidence, not volume.

⭐ What you need

Technical

  • Track record. 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company.

  • Breadth. Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or Pulumi on at least one major cloud (AWS preferred), plus good proficiency in Python or Go for automation and tooling.

  • AI in workflow. AI is part of how you ship, not a thing you've read about - agentic tooling (Claude Code, Droid, Codex, internal skills) is in your daily loop, you've built or adopted AI-assisted workflows others now use, and you have strong opinions on where it's unreliable. This is a hard requirement, not a bonus. Candidates whose actual daily workflow does not already include AI tooling will not be advanced.

  • First-principles + decision-making. Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems - reason from constraints and failure modes (not analogy or vendor defaults), name the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off), and reject the "best practices" answer when it doesn't fit the problem.

  • Reversibility & blast-radius. Make reversible calls by default - write the rollback before you touch production, work fluently with monitoring and logging stacks (Prometheus, Grafana, ELK or equivalent), and stress the system in safe places so it comes back stronger.

Non-technical

  • Cross-functional collaboration. Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams - surface non-goals before anyone asks, and partner with product, security, and platform peers as one delivery surface.

  • Autonomy & end-to-end ownership. A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability. At least one 0-to-1 infrastructure build you owned end-to-end, with the outcome metric attached.

Bonus if you have

  • Software-engineering depth. A software-engineering background, not only config and scripting - you've designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or a comparable language, you can read and modify the codebases your infrastructure runs, and you move between infra automation and feature engineering without changing brains.

Benefits & perks (UK full-time employees):

  • Generous PTO, plus company holidays

  • Comprehensive medical and dental insurance

  • Paid parental leave for all parents (16 weeks)

  • Fertility and family planning support

  • Early-detection cancer testing through Galleri

  • Competitive pension scheme and company contribution

  • Annual work-life stipends for:

    • Wellness stipend for gym, massage/chiropractor, personal training, etc.

    • Learning and development stipend

  • Company-wide off-sites and team off-sites

  • Competitive compensation and company stock options

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
746,022 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
London
$160k – $200k per year • Remote (United Kingdom, Ireland) • Full-Time • 8+ years exp • Dublin
Python
Go
Java
Kotlin
DevOps
Terraform
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
FinOps
Apply
≈ $94k – $210k per year (Estimated) • Remote (United Kingdom) • Full-Time • 8+ years exp • Edinburgh
Python
Go
Java
AI/ML
Copilot
Cursor
Claude Code
Prompt Engineering
DevOps
Terraform
CloudFormation
Pulumi
CI/CD
AWS
Kubernetes
Platform Engineering
Incident Management
Apply
≈ $90k – $168k per year (Estimated) • Remote (United Kingdom) • Full-Time • 5+ years exp • Edinburgh
Python
Go
Java
AI/ML
Copilot
Cursor
Claude Code
Prompt Engineering
DevOps
Terraform
CloudFormation
Datadog
Prometheus
Pulumi
CI/CD
AWS
Kubernetes
Grafana
Platform Engineering
Apply
≈ $97k – $217k per year (Estimated) • In office • London
Python
Go
Java
DevOps
CI/CD
Kubernetes
Platform Engineering
FinOps
Incident Management
Apply
≈ $83k – $155k per year (Estimated) • In office • Full-Time • Aberdeen
Python
PowerShell
DevOps
Terraform
GCP
Azure
Management
ITIL
Apply
$305k – $457k per year • In office • Full-Time • Yokohama
Python
JavaScript
Java
PHP
Node JS
Databases
MySQL
PostgreSQL
ElasticSearch
DevOps
Terraform
GCP
CloudFormation
AWS
Docker
Nginx
Apache HTTP Server
Apply
$94k – $142k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Mississauga
Python
JavaScript
TypeScript
SQL
AI/ML
Copilot
Edge AI
Machine Learning
DevOps
GitHub Actions
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Shift-Left
Self-Healing
Pipeline as Code
Tekton
Cybersecurity
Shift-Left Security
Management
Agile
Scrum
QA
Selenium
Apply
≈ $138k – $269k per year (Estimated) • Hybrid • Suitland
Java
Java
Spring Boot
Databases
Apache Kafka
DevOps
Rest API
GCP
Azure
AWS
Docker
Kubernetes
Amazon EKS
SOAP
Apply
$107k – $157k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Toronto
JavaScript
Node JS
Databases
MySQL
Redis
DynamoDB
Frontend
GraphQL
React.js
DevOps
Rest API
Splunk
Terraform
Dynatrace
CI/CD
Jenkins
AWS
Docker
Amazon EC2
Amazon CloudWatch
Management
Agile
Scrum
Apply
$83k – $149k per year • Equity • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Python
SQL
AI/ML
Physical AI
Analytics
Tableau
Power BI
Apply
$155k – $240k per year • Equity • Hybrid • Full-Time • 5+ years exp • New York • Seattle • San Francisco
Python
AI/ML
Claude Code
AI Agents
OpenAI Codex
DevOps
Terraform
GCP
Helm
Prometheus
Pulumi
Azure
AWS
Kubernetes
Grafana
Apply
$170k – $240k per year • Equity • Hybrid • Full-Time • 6+ years exp • San Francisco • New York
AI/ML
AI Agents
Apply
$183k – $240k per year • Equity • Hybrid • Full-Time • 4+ years exp • New York • Seattle • San Francisco
Python
Go
JavaScript
TypeScript
AI/ML
AI Agents
LLM
DevOps
CI/CD
Cybersecurity
Threat Modeling
Apply
≈ $78k – $144k per year (Estimated) • Equity • Hybrid • Full-Time • 5+ years exp • London
Python
TypeScript
AI/ML
Model Context Protocol
AI Agents
LLM
Machine Learning
DevOps
GCP
GitHub Actions
Azure
CI/CD
AWS
Docker
Kubernetes
QA
Playwright
Apply
AI engineer (UK) 2 days ago
≈ $117k – $237k per year (Estimated) • Equity • Hybrid • Full-Time • 5+ years exp • London
Python
AI/ML
JAX
AI Agents
TensorFlow
PyTorch
Machine Learning
DevOps
GCP
Azure
AWS
Apply
Hybrid • Full-Time • London
Python
JavaScript
PowerShell
Bash
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Linux
Management
Agile
Apply
Security Engineer 8 hours ago
≈ $75k – $162k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
DevOps
Incident Management
Apply
≈ $68k – $111k per year (Estimated) • Hybrid • Full-Time • 1+ year exp • PhD • London
Management
Agile
Apply
≈ $30k – $47k per year (Estimated) • In office • Internship • London
Apply
≈ $92k – $218k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • London
AI/ML
AI Agents
Apply
See all jobs
This is one of many
746,022 more open roles from verified company boards, updated every day.