Link Group
Site Reliability Engineering (Primary Focus)
Monitor, support, and manage production incidents for trading and market platforms, including root cause analysis.
Troubleshoot production issues and implement permanent fixes to prevent recurrence.
Perform capacity planning and support the stability of on-premises platforms.
Work in a highly critical production environment with a high degree of autonomy (without a local team lead on site).
AI & Automation (Growth Opportunity)
Build AI agents to automate monitoring and system health checks, including:
Daily and weekly platform health checks.
Pre-trade, post-trade, and settlement monitoring.
Order flow monitoring.
Work on the organization's central observability platform with dedicated training and support for developing AI agents.
Leverage AI tools such as GitHub Copilot and LLMs to support troubleshooting and operational decision-making.
Requirements
Must Have
Strong SRE fundamentals, including incident management, production support, and capacity planning.
Proven experience troubleshooting and resolving production issues in highly available environments.
Proficiency in Python (preferred), along with knowledge of Java or shell scripting.
Hands-on experience with Prometheus, Grafana, and Elasticsearch.
Experience working with on-premises environments (limited exposure to public cloud is sufficient).
Excellent English communication skills.
Ability to work independently with a high level of ownership and autonomy.
Nice to Have
Experience with Azure DevOps (ADO).
Experience building or working with AI agents and prompt engineering.
Experience in banking or similarly complex, mission-critical environments such as fintech, telecommunications, or large-scale e-commerce.
