818,177open jobs
52,655companies
132,591added this week
Browse all
Salary
≈ $169k – $312k per year (Estimated)
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 24, 2026. Arena scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Arena is a technology company that develops artificial intelligence based software to enhance customer engagement and real time data experiences for online businesses. Its products are used to personalise user interactions, increase retention, and improve analytics across digital platforms.

About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.

Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.

We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is looking for a Software Engineer - Platform to build the core infrastructure that sits beneath our online evaluation systems - the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.

This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.

Our AI gateway is currently in private, gated launch with individual developers as our primary users today; enterprise use cases will follow later. Next month we're shipping direct model access and new infrastructure features, and Arena may begin collecting usage traces and building leaderboards shared with lab partners - so there's a lot of near-term, zero-to-one work ahead.

You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.

This is a hands-on individual contributor role - we're not hiring for a tech lead or SRE function right now; everyone on the team is heads-down building.

What You'll Do

  • Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.

  • Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.

  • Ship enterprise-grade infrastructure. Build the systems enterprise customers will eventually expect - rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance - as we grow beyond our current individual-developer user base.

  • Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.

  • Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products. (This role does not involve data labeling or third-party data-verification work.)

  • Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.

You’ll have

  • ~5+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.

    • Senior: proven delivery on meaningful backend work.

    • Staff: extensive, deep experience owning backend systems end to end.

  • Strong proficiency in Go - this is our primary backend language and a must-have for the role.

  • Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.

  • A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."

  • Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.

Nice to Have

  • Cloud infrastructure experience (AWS, GCP, or Azure), Kubernetes, Terraform, and database systems like Postgres and Redis - helpful, but not a hard filter.

  • Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).

  • Background in AI/ML infrastructure, model serving, inference, or evaluation frameworks.

  • Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.

  • Experience building billing infrastructure around systems like Stripe, Metronome, and Orb.

  • Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).

Location

This role is based in San Francisco, with a minimum of 3 days/week onsite. Fully remote candidates will only be considered with a very strong endorsement.

What we offer

  • We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.

  • The opportunity to work on cutting-edge AI with a small, mission-driven team

  • A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
818,177 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
≈ $100k – $195k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Morrisville
JavaScript
Node JS
Databases
MySQL
DynamoDB
Frontend
React.js
DevOps
Rest API
GCP
Azure
AWS
IAM
Cybersecurity
HIPAA
Least Privilege
Management
Agile
Apply
$230k – $262k per year • In office • Full-Time • 9+ years exp • Bachelor's Degree • New York • McLean • Richmond
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Apply
≈ $146k – $285k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Python
JavaScript
TypeScript
AI/ML
Reinforcement Learning
Frontend
Next.js
React.js
DevOps
Terraform
CI/CD
AWS
Docker
Kubernetes
Grafana
Apply
Founding Engineer 1 hour ago
$150k – $200k per year • In office • Full-Time • 3+ years exp • San Francisco
TypeScript
AI/ML
AI Agents
Apply
≈ $134k – $247k per year (Estimated) • Hybrid • 5+ years exp • Bachelor's Degree • Los Angeles
Python
C++
Apply
$152k – $274k per year • Remote (United States) • 15+ years exp • Bachelor's Degree
AI/ML
Model Context Protocol
Vertex AI
AI Agents
LLM
OpenAI
DevOps
Terraform
GCP
Azure
AWS
KVM
FinOps
AIOps
Cybersecurity
Okta
Zero Trust
Management
ServiceNow
Apply
≈ $111k – $250k per year (Estimated) • Remote (location not specified)
AI/ML
Anomaly Detection
DevOps
IAM
Cybersecurity
SOC 2
CVE
Least Privilege
Apply
Software Engineer 3 hours ago
$142k – $236k per year • In office • TS/SCI • 7+ years exp • Bachelor's Degree • Annapolis Junction
Python
Go
JavaScript
TypeScript
C++
Frontend
Vue.js
Angular
React.js
Apply
AI Engineer 3 hours ago
≈ $40k – $89k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • Master's Degree • Gdańsk
SQL
Databases
Snowflake
MS SQL
Microsoft Fabric
AI/ML
Copilot
AI Agents
Copilot Studio
Analytics
Power BI
Management
Agile
Scrum
Apply
$74k – $84k per year • Hybrid • Full-Time • Gdańsk
JavaScript
TypeScript
SQL
C#
C#
.NET
Databases
MS SQL
AI/ML
Copilot
Prompt Engineering
OpenAI
OCR
Copilot Studio
DevOps
Azure DevOps
Azure
CI/CD
GitHub
Analytics
Power BI
Management
Power Automate
Power Apps
Apply
≈ $139k – $270k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
JavaScript
TypeScript
Databases
PostgreSQL
Supabase
AI/ML
AI Agents
Edge AI
Vercel AI SDK
Frontend
Tailwind CSS
Next.js
React.js
Radix UI
shadcn/ui
DevOps
Vercel
QA
Vitest
Apply
≈ $161k – $297k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
JavaScript
TypeScript
AI/ML
AI Agents
Edge AI
Frontend
GSAP
Tailwind CSS
Next.js
React.js
Radix UI
shadcn/ui
Apply
$180k – $300k per year • In office • Full-Time • 6+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
SQL
Node JS
AI/ML
AI Agents
LLM
Edge AI
Tool Use
Apply
$150k – $300k per year • In office • Full-Time • 4+ years exp • San Francisco
JavaScript
TypeScript
Databases
PostgreSQL
Supabase
AI/ML
AI Agents
Edge AI
Vercel AI SDK
Frontend
Tailwind CSS
Next.js
React.js
Radix UI
shadcn/ui
DevOps
Vercel
QA
Vitest
Apply
$200k – $350k per year • In office • Full-Time • 8+ years exp • San Francisco
JavaScript
TypeScript
Databases
PostgreSQL
Supabase
AI/ML
AI Agents
LLM
Edge AI
Vercel AI SDK
Frontend
Tailwind CSS
Next.js
React.js
Radix UI
shadcn/ui
DevOps
Vercel
QA
Vitest
Apply
$130k – $175k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Apply
≈ $105k – $207k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$140k – $295k per year • In office • 5+ years exp • San Francisco
Apply
$100k – $125k per year • In office • 2+ years exp • San Francisco
Python
Apply
$145k – $235k per year • In office • 3+ years exp • San Francisco
Apply
See all jobs
This is one of many
818,177 more open roles from verified company boards, updated every day.