368,611open jobs
9,439companies
50,719added this week
Browse all
Location
In office (Trondheim)
Seniority
Principal · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
SHIFTER - teknologi, innovasjon, business. Portrettintervju – Det er gründerblod i meg Les også.

Does it sound interesting to work on an open source platform managing the data and real-time search and inference for some of the largest companies in the world? Would you thrive on keeping large, globally distributed systems reliable, fast, and observable - and on building the practices and tooling that let a small team operate at massive scale? If so, we want you to join our team at Vespa.ai as a Principal Site Reliability Engineer!

About Vespa.ai:

Vespa.ai is a team of passionate builders. We maintain and develop the Apache 2.0 licensed open-source AI search platform Vespa.

Vespa is a fully featured search engine and vector database. It supports vector search (ANN), lexical search, and structured data search, all in a single query. Integrated machine-learning model inference enables the application of AI to make sense of data in real time. Together with Vespa’s proven scalability and high availability, this empowers to create production-ready search applications at any scale and with any combination of features. Our users and customers are #1 in e-commerce, content, and financial services globally, and are used by companies such as Perplexity, Spotify, Yahoo, Wix, and many more.

In addition to our open-source platform, Vespa.ai develops and runs Vespa Cloud, a robust SaaS offering that allows businesses to harness the power of our technology with ease.

At Vespa.ai, we are extremely focused on automating everything we do to grow fast and maintain high quality. In all roles, we scale through technology, not simply by adding larger teams. We take pride in being small, nimble, and the most productive.

Position overview

At Vespa.ai, we embrace DevOps as a company culture, seeking to solve technical problems with automation and code rather than repetitive manual effort. For our Vespa Cloud production systems, we have had this mindset from day one.

We are seeking a Principal Site Reliability Engineer to join our team and help keep Vespa Cloud reliable, fast, and observable at global scale. This is a senior individual contributor role on the team that operates and improves our production systems. You will also help shape and develop our approach to SRE and DevOps as we grow. We are looking for a strong engineer who earns influence through contributions and has the ambition to take on greater responsibility over time. You will also participate in our 24x7 on-call rotation, approximately every third to fourth week.

At our Trondheim office, we work office-first: you will be based on-site most of the time, with the flexibility to work from home/remotely when needed, as agreed with your manager.

Responsibilities

  • Help ensure the reliability, availability, and performance of Vespa Cloud production systems running globally at scale.
  • Participate in a 24x7 on-call rotation (approximately every 3rd-4th week), lead incident response, and drive blameless postmortems through to durable fixes.
  • Help define and track SLOs/SLIs, and build proactive alerting, capacity planning, and remediation strategies.
  • Design and improve observability - metrics, logging, and tracing - across a large fleet.
  • Eliminate operational toil by solving problems with automation and code rather than manual effort.
  • Contribute to, and help shape, our SRE and DevOps practices and culture as the organization grows, sharing knowledge and mentoring across the team.
  • Work with the rest of the Vespa.ai developing team on reliability, scalability, and architecture.

Qualifications

  • 5-10 years building and operating large-scale production systems, with deep SRE/DevOps experience.
  • Solid programming skills in Java, Python, Go, or similar languages.
  • Good understanding of sound software engineering principles and practices.
  • Experience with cloud platforms (AWS, Azure, or GCP).
  • Solid understanding of networking, operating systems, distributed systems, and security principles.
  • Proven incident management and on-call experience.
  • A track record of influencing technical direction and improving how teams work - not just executing tickets.
  • Excellent problem-solving and analytical skills, and the ability to lead through influence as well as work independently.

Desired Skills

  • Experience with Infrastructure as Code tools such as Terraform, Tofu, Spacelift, etc.
  • Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK).
  • Experience with CI/CD tooling such as GitHub Actions, Buildkite, etc.
  • Experience operating data-intensive or stateful systems at scale.
  • Experience defining SLOs and establishing reliability programs.
  • Ambitions beyond pure SRE - an interest in growing, over time, into a technical leadership role.

Some of Our Tools and Services

  • JumpCloud, Google Workspace, and Slack
  • GitHub Enterprise Cloud (including GitHub Actions)
  • Jira Cloud and Jira Service Desk
  • StrongDM, Grafana, Spacelift, and Buildkite
  • AWS, GCP, and Azure

Why Join Us:

  • Opportunities for professional growth and development as part of one of Europe’s most exciting start-ups!
  • Be part of a cutting-edge team working on innovative search and recommendation technology.
  • Work on a team where we don’t believe in silos between engineers; there aren’t “developers”, “ops people”, and “sysadmins”. We’re all engineers solving problems the smart way together!
  • Competitive salary and benefits.

Note: Vespa.ai is an equal-opportunity employer. We are committed to creating an inclusive environment for all employees. We believe in fostering a collaborative and inclusive environment where every team member has the opportunity to make a significant impact.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Trondheim
$19k – $53k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Bash
JavaScript
Python
TypeScript
Frontend
Angular
React.js
DevOps
AWS
Azure
Datadog
Docker
GCP
Grafana
Kubernetes
Prometheus
Splunk
IAM
Cybersecurity
Keycloak
Apply
SDET 1 day ago
$12k – $39k per year (Estimated) • In office • 4+ years exp • Gurgaon
Java
DevOps
AWS
Azure
CI/CD
Jenkins
QA
Appium
JMeter
Playwright
Rest-Assured
Selenium
Apply
$67k – $173k per year (Estimated) • In office • Contractor • 3+ years exp • Bachelor's Degree • Singapore
JavaScript
Python
DevOps
AWS
Azure
GCP
Cybersecurity
ISO 27001
OWASP Top 10
Apply
$131k – $281k per year (Estimated) • In office • Contractor • 15+ years exp • Singapore
Java
SQL
Java
Maven
Databases
MS SQL
Oracle
DevOps
AWS
CI/CD
Incident Management
Kubernetes
OpenShift
OpenStack
Apply
$20k – $45k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Python
pySpark
Databases
Databricks
Microsoft Fabric
AI/ML
Hadoop
Spark
DevOps
AWS
Azure
Analytics
Power BI
Tableau
Apply
Software Developer 2 years ago
In office • Full-Time • Master's Degree • Trondheim
Bash
C++
Go
JavaScript
Python
TypeScript
Databases
Vespa
AI/ML
LangChain
ONNX
Hugging Face
Recommender Systems
Frontend
Mantine
React.js
DevOps
Ansible
AWS
Azure
Docker
GCP
Grafana
Kubernetes
Opsgenie
Podman
Prometheus
Terraform
Apply
In office • Full-Time • Trondheim
DevOps
SLI/SLO/SLA
Apply
In office • Full-Time • Bachelor's Degree • Trondheim
Java
SQL
Java
Spring Boot
Databases
Amazon Aurora
Amazon Redshift
ClickHouse
MySQL
AI/ML
Cursor
DevOps
Amazon EKS
AWS
Kubernetes
Amazon S3
GitLab
Apply
Data Engineer 2 months ago
In office • Full-Time • Trondheim
C#
Java
Kotlin
Python
Databases
Apache Kafka
Databricks
Delta Lake
AI/ML
MLFlow
DevOps
AWS
Azure
CI/CD
GCP
Terraform
Apply
In office • Full-Time • Trondheim
JavaScript
TypeScript
Frontend
Angular
React.js
Apply
In office • Full-Time • Trondheim
Bash
Python
Databases
ElasticSearch
OpenSearch
DevOps
Azure
Bicep
CI/CD
Docker
GitHub Actions
GitLab CI
Grafana
Helm
Kubernetes
Terraform
GitHub
GitLab
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.