697,746open jobs
41,033companies
105,083added this week
Browse all
Salary
$68k – $104k per year (Estimated)
Location
Remote (Poland)
Seniority
Senior · 3+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match

Location: Poland only, fully remote

Job Type: B2B, full time

Overview

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We care about each customer's interaction, experience, behaviour, and insight and strive to ensure we’re always acting authentically.

Rooted in the kindred spirits of the Seminole Tribe of Florida, the new Hard Rock Digital taps a brand known all over the world as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space.

What’s the position?

We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents.

You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights-reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs.

Key Responsibilities

Application Reliability & Performance

  • Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment.

  • Troubleshoot and resolve complex issues across production and non-production environments.

  • Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance.

  • Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling.

Monitoring, Observability & AIOps

  • Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting.

  • Implement and refine observability strategies that enhance visibility into application and infrastructure health.

  • Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring.

  • Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction.

AI & Agentic Workflow Engineering

  • Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization.

  • Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval.

  • Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents.

  • Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems.

  • Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving.

  • Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization.

Incident Management & Root Cause Analysis

  • Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence.

  • Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents.

  • Document and share lessons learned, contributing to a culture of continuous improvement.

Automation & Toil Reduction

  • Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements.

  • Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language.

  • Measure and report on toil reduction metrics to quantify the impact of automation initiatives.

Collaboration & Cross-functional Support

  • Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities.

  • Collaborate with DevOps and NOC teams to support the application platform.

  • Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders.

  • Provide feedback on application performance, potential improvements, and observability metrics.

Why This Role Is Different

This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies-on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems.

Requirements

What are we looking for?

Core SRE & Infrastructure (Required)

  • Degree in Computer Science or a related field, or equivalent professional experience.

  • 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems.

  • 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security.

  • Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management.

  • Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting.

  • Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection.

  • Proficiency in PromQL and experience with Loki for log aggregation and analysis.

  • Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization.

  • Cloud platform expertise (AWS preferred; GCP or Azure also valued).

  • Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible.

  • ArgoCD proficiency for GitOps workflows and continuous deployment.

  • Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation.

  • Proven track record with on-call rotations, incident response, and root cause analysis.

AI, Automation & Agentic Systems (Required)

  • 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context.

  • Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks.

  • Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines.

  • Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent).

  • Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples.

  • Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows.

  • Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents.

  • Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems.

Preferred / Bonus

  • Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval.

  • Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production.

  • Contributions to open-source AI/ML or SRE tooling projects.

  • Background in data engineering or ML pipelines that complements SRE responsibilities.

Soft Skills

  • Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences.

  • Proactive problem-solver with a bias toward automation and continuous improvement.

  • Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches.

  • Positive attitude and openness to constructive feedback.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
697,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
AI Engineer 1 hour ago
$22k – $56k per year (Estimated) • In office • Full-Time • 4+ years exp
Python
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
FAISS
Apache Kafka
AI/ML
Copilot
Cursor
LangChain
Ray Serve
Qwen
Spark
Claude Code
LlamaIndex
LoRA
vLLM
Fine-tuning
Embeddings
Quantization
Knowledge Distillation
Computer Vision
NLP
AWQ
GPTQ
ONNX
PEFT
QLoRA
TGI
Llama
Mistral
TensorFlow
PyTorch
LLM
RAG
Ray
Hallucination
Triton
OpenAI
Anthropic
Feature Store
LLM Guardrails
Recommender Systems
Agentic Workflows
Model Distillation
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$145k – $218k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • Mississauga
Python
Java
Kotlin
SQL
Databases
Weaviate
Pinecone
Apache Kafka
AI/ML
Copilot
AutoGen
LangChain
Fine-tuning
Prompt Engineering
AI Agents
Flink
LLM
RAG
Anomaly Detection
Feature Store
DevOps
CI/CD
AWS
Docker
Kubernetes
Trunk-Based Development
Chaos Engineering
Progressive Delivery
AIOps
SLI/SLO/SLA
Cybersecurity
Threat Modeling
Management
Agile
Apply
Full Stack Engineer 1 hour ago
$22k – $55k per year (Estimated) • In office • Full-Time • 5+ years exp
Python
JavaScript
TypeScript
Python
FastAPI
Django
Databases
MySQL
PostgreSQL
AI/ML
Copilot
Cursor
Claude Code
LLM
RAG
Frontend
React.js
DevOps
Rest API
Terraform
GCP
Datadog
Azure
CI/CD
AWS
Docker
Grafana
QA
Sentry
Apply
DevOps Engineer 1 hour ago
$37k – $39k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree
Python
JavaScript
Node JS
Bash
Databases
PostgreSQL
TimescaleDB
DevOps
Terraform
Ansible
GCP
GitLab CI
Azure
CI/CD
Jenkins
AWS
Hetzner
Linux
Apply
$21k – $52k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Pretoria
Python
SQL
Databases
ClickHouse
AI/ML
Hadoop
Spark
Airflow
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Talend
Apply
$61k – $119k per year (Estimated) • Remote/Hybrid • 1+ year exp • Hollywood
Apply
$35k – $72k per year (Estimated) • In office • 2+ years exp • Las Vegas
Management
Outlook
Agile
Microsoft Office
Apply
$139k – $270k per year (Estimated) • Remote/Hybrid • 10+ years exp
Java
AI/ML
Copilot
Cursor
Claude Code
Apply
Remote/Hybrid • Hollywood
Marketing
YouTube
Apply
$42k – $74k per year (Estimated) • Remote/Hybrid • 1+ year exp
SQL
Mobile
AppsFlyer SDK
Adjust
Analytics
Tableau
Power BI
A/B Testing
Apply
See all jobs
This is one of many
697,746 more open roles from verified company boards, updated every day.