738,400open jobs
44,340companies
105,262added this week
Browse all
Location
In office
Seniority
Senior · 6+ years exp

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026.

Overview
Company
Impact
Profile match
Aliz is an IT consultancy and cloud solutions provider specializing in data engineering, artificial intelligence, and Google Cloud Platform (GCP) integration. Headquartered in Budapest, Hungary - with regional hubs across Europe, Asia-Pacific, and North America - the company operates as a certified Google Cloud Partner with specializations in infrastructure, data analytics, machine learning, and generative AI. Its core offerings include enterprise cloud migration, BigQuery data architecture, custom application development, and production-ready AI platform deployment.

The Role

Help Rabbit build and ship faster with AI - safely, securely and reliably.

Rabbit runs automated cost optimization across enterprise Google Cloud environments. When we change a customer's BigQuery reservations or rightsize their GKE clusters, those changes need to be correct and dependable. Reliability is central to the trust customers place in our product.

Our foundation is already in place: logging, alerting, automated deployment and Terraform-managed infrastructure. Your mission is to evolve that foundation for an AI-accelerated engineering team: turn faster implementation into faster, dependable delivery through automated validation, safe releases and rapid feedback.

You'll apply proven SRE practices - SLOs, observability, incident response and deployment safety - to AI-assisted development and agent-driven workflows. The goal is to increase how quickly the team can deliver verified improvements, while controlling production risk and reducing manual operational work.

What You'll Do

  • Make AI-assisted delivery faster and safer.Build automated validation, progressive rollout and recovery mechanisms that let engineers and agents move quickly with clear checks before and after changes reach production.
  • Make reliability measurable.Define and operationalize SLOs, SLIs and error budgets, and use them to guide practical decisions about delivery speed, stability and reliability work.
  • Automate workflows with AI agents.Identify repetitive operational work and build reusable agent-driven workflows for alert triage, incident investigation, routine maintenance and reporting. Add verification and human approval where needed, and measure the reduction in manual effort.
  • Improve observability and feedback.Evolve logging, metrics, tracing and alerting so failures are detected early and changes can be traced, investigated and verified.
  • Turn incidents into lasting improvements.Improve runbooks, investigation and blameless postmortems, and translate recurring problems into tests, safeguards and automation.
  • Keep infrastructure reproducible.Extend our Terraform and delivery tooling so environments remain consistent and changes stay reviewable as the platform grows.
  • Improve GCP reliability and efficiency.Strengthen our cloud infrastructure, networking, access controls and capacity management, balancing performance, reliability and cost.
  • Use AI to accelerate reliability engineering itself.Build maintainable tooling in Go, Python or a comparable language, and use agents to accelerate investigation, implementation, testing and documentation while verifying their outputs.

How We Work - AI-First, Agentic by Default

AI-assisted engineering is an expectation of this role, not an optional experiment. Claude Code, Cursor and agent-driven workflows are part of how we work, including infrastructure and reliability engineering.

We want someone who actively looks for ways to increase engineering speed with AI and makes those improvements safe to repeat. That means shorter feedback loops, automated checks, traceable changes and recovery paths - not simply generating more code.

You remain accountable for engineering judgment: what to automate, how to verify it, when human approval is needed and when a change should be stopped or rolled back. Success means faster delivery of reliable improvements, less repetitive work and a platform the team can trust.

What You'll Bring

Must-have

  • 6+ years in SRE, production engineering or infrastructure-heavy backend roles, with hands-on ownership of production systems.
  • Strong production GCP experience, including Cloud Run, networking and IAM. Hands-on Google Cloud experience is required and will be assessed during the interview process.
  • Infrastructure-as-code fluency with Terraform, plus solid experience in CI/CD and deployment safety.
  • Strong observability and troubleshooting skills: you can make systems debuggable, identify root causes and verify that a fix works.
  • Coding ability in Go, Python or a comparable language, with experience building maintainable operational tooling.
  • Experience leading production incident investigation and driving follow-up improvements that prevent recurrence.
  • Practical fluency with AI coding agents and the ability to critically review, test and validate their work. You are motivated to make AI-assisted engineering faster and more dependable.
  • Strong written English and the discipline to collaborate asynchronously with a distributed team.

Nice to have

  • Kubernetes / GKE experience, including deploying, operating, debugging and scaling containerized services.
  • Hands-on Datadog experience, including dashboards, monitors, logs, APM and distributed tracing.
  • Experience improving delivery speed and safety through progressive delivery, policy-as-code and automated rollback.
  • GCP cost management or FinOps experience.
  • Security experience is a plus.

Why This Role

  • Shape how an AI-first engineering team scales.Build the reliability practices and automation that let Rabbit turn faster development into dependable customer outcomes.
  • Work on systems with real customer impact.Rabbit operates inside enterprise GCP environments, where reliability and security directly affect customer trust.
  • Own meaningful improvements.Work in a small team with short decision paths and end-to-end ownership, supported by appropriate review and production safeguards.
  • Use AI as an engineering multiplier.Apply agentic tools to infrastructure, delivery and operations, and help define how we measure and improve their impact.

About Aliz

Aliz was founded by a group of driven software engineers with one vision in mind: Build a human-centered company to solve exciting business challenges with cutting-edge technologies all around the globe.

Our team includes engineers, industry experts, and digital professionals who ensure our clients are ready to face the challenges of the ever-changing digital world.

Aliz is a proud Google Cloud Partner with specializations in Infrastructure, Data Analytics, Cloud Migration, and Machine Learning. We deliver data analytics, machine learning, infrastructure, and application development solutions, off the shelf, or custom-built on GCP using an agile, holistic approach. Recently we've launched our very own cloud cost optimization tool Rabbit.

Life has been pretty exciting at Aliz recently (here's proof). After we became Google Cloud’s Breakthrough Partner of the Year in APAC (2019), we built a team in Singapore and Berlin, opened an engineering hub in Jakarta, launched a business in Italy and Qatar, and started to support IT education.

You couldn’t pick a better time to join this world-class team of cloud experts.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
738,400 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
$29k – $41k per year • In office • Full-Time • Osaka
Python
JavaScript
Databases
Redis
Frontend
React.js
DevOps
Jenkins
Git
AWS
Nginx
GitHub
Linux
Management
Slack
Trello
Apply
$35k – $98k per year (Estimated) • In office • Full-Time • Master's Degree • Tokyo
DevOps
AWS
Apply
Equity • Remote/Hybrid • Full-Time • 4+ years exp • Master's Degree
Python
JavaScript
PHP
SQL
Apex
AI/ML
Model Context Protocol
AI Agents
Frontend
JQuery
DevOps
GCP
Azure
AWS
Management
Slack
Apply
$36k – $99k per year (Estimated) • In office • Full-Time • Master's Degree • Tokyo
DevOps
GCP
AWS
Platform Engineering
Management
Slack
Google Workspace
Zapier
Apply
$36k – $100k per year (Estimated) • In office • Full-Time • Master's Degree • Tokyo
Chips/EDA
PoC Library
Management
Slack
Google Workspace
Apply
$34k – $87k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Python
Java
AI/ML
Fine-tuning
Embeddings
Quantization
Function Calling
AI Agents
LLM
RAG
Hallucination
Anomaly Detection
Google AI Studio
Feature Store
Knowledge Graph
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GCP
Azure
AWS
Kubernetes
Platform Engineering
Apply
$41k – $107k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Hsinchu
Python
MATLAB
DevOps
Windows
Management
Microsoft Teams
Apply
$72k – $153k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
Python
Stata
Apply
$20k per year (net) • In office • 1+ year exp • Bachelor's Degree • Moscow
Python
PowerShell
DevOps
Windows
Cybersecurity
SIEM
Apply
Junior DevOps 9 hours ago
$12k – $24k per year (Estimated) • Remote/Hybrid • 3+ years exp • Moscow
Python
Bash
Databases
ElasticSearch
Apache Kafka
OpenSearch
DevOps
Ansible
Kibana
Kong
Logstash
Prometheus
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
Alertmanager
GitLab
API Gateway
Linux
TCP/IP
DNS
Cybersecurity
HashiCorp Vault
Apply
Remote/Hybrid • 4+ years exp
AI/ML
AI Agents
Machine Learning
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Kubernetes
OpenStack
Linux
Management
Agile
Apply
Remote/Hybrid • 7+ years exp
Python
Java
SQL
AI/ML
Copilot
Claude Code
Model Context Protocol
AI Agents
LLM Guardrails
Agentic Workflows
Machine Learning
DevOps
Terraform
GCP
Management
Google Drive
Agile
Apply
In office • 4+ years exp
SQL
Databases
Google BigQuery
BigQuery
AI/ML
AI Agents
Machine Learning
DevOps
GCP
FinOps
Chips/EDA
PoC Library
Management
Agile
Apply
Remote (likely Hungary)
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Copilot
Cursor
Claude Code
AI Agents
Machine Learning
DevOps
Terraform
GCP
FinOps
Management
Agile
Apply
In office • 3+ years exp
SQL
AI/ML
AI Agents
DevOps
FinOps
Chips/EDA
PoC Library
Apply
See all jobs
This is one of many
738,400 more open roles from verified company boards, updated every day.