Viva.com is the first Tech Bank in Europe for businesses, serving 31 countries through its pan-European critical infrastructure for payments & banking services.
Viva.com’s offering spans omnichannel payment acceptance, deposit accounts, card issuing, and a broad suite of financing solutions, providing businesses of all sizes and sectors with the tools they need to operate efficiently, both domestically and internationally, through a single integrated platform.
Viva.com pioneered and leads in Tap on Any Device technology, enabling acceptance on any Android device or iPhone, while connecting directly to all major international and domestic card schemes, local payment systems, and alternative payment methods across Europe.
The Role
We are seeking a Site Reliability Engineer (SRE) to ensure the availability, performance, and scalability of our production systems. You will work with developers, system administrators, and support teams to improve system architecture and operational processes, with a strong focus on automation, observability, and incident response.
Key Responsibilities
- Ensure the reliability and uptime of critical production services and infrastructure;
- Contribute to the design of scalable monitoring, alerting, and observability systems;
- Develop tools and automation to eliminate manual and repetitive work;
- Lead and contribute to postmortems and root cause analysis, and see remediation actions through to completion;
- Maintain runbooks, escalation paths, and on-call documentation;
- Define and track SLIs and SLOs together with engineering teams;
- Collaborate with software and system engineers to improve system design for reliability;
- Identify and fix system weaknesses, bottlenecks, and single points of failure.
At Viva.com, we believe in amplifying human potential through AI. When you join us, you're stepping into a future-forward environment where technology empowers every role, driving smarter, more efficient.
Requirements
- More than 3 years of experience in a Site Reliability Engineer, DevOps, or similar role;
- Hands-on experience operating production systems in a live on-call capacity;
- Proficiency in scripting languages (Python, Bash, Go, or similar);
- Experience with cloud platforms (Azure preferred) and container orchestration (Kubernetes, Docker);
- Practical experience applying AI to operations, including agentic and managed AI tooling for automation, diagnostics, or incident triage, with sound judgement on where human oversight is required;
- Strong understanding of Linux systems, networking, and troubleshooting;
- Familiarity with infrastructure as code tools (Terraform, Ansible, etc.);
- Familiarity with observability stacks (Datadog, Prometheus, Grafana, ELK, etc.);
- Familiarity with CI/CD pipelines and automated rollback;
- Experience with incident management and on-call tooling such as PagerDuty or incident.io;
- Clear communication under pressure, including with non-technical stakeholders;
- Strong problem-solving skills and a passion for reliability and performance.
Benefits
Competitive compensation package;
Annual bonus based on your performance and targets’ achievement;
Private health insurance for you and your family;
Top of the Line tools and equipment;
Career development and regular feedback to develop your skills;
Employee Wellness Program like Daily group sessions led by professional coaches.

