First seen by Alion on Sep 30, 2026.
As a recruitment company, DCG understands that every business is powered by experienced professionals. Our management style and partnership approach enable us to meet your needs and provide continuous support. Due to our ongoing growth and the large number of recruitment projects we undertake for our partners, we are currently looking for:
Production Support & Operations Engineer
Responsibilities:
Own RCA activities for production incidents, including diagnosis, resolution, and preventive actions
Monitor production services, identify anomalies, and proactively address issues before they become incidents
Design, implement, maintain, and improve CI/CD pipelines using GitHub Actions and related tooling
Automate build, test, security scanning, and deployment processes
Oversee day-to-day stability and alignment of Pre-Production and Production environments
Maintain operational documentation, runbooks, known issues, and resolution procedures
Collaborate with development and platform teams to troubleshoot issues and clarify operational requirements
Identify recurring operational pain points and propose automation or tooling improvements
Enhance observability through dashboards, alerts, and log queries
Support service continuity initiatives and participate in disaster recovery exercises
Requirements:
Minimum 5 years of experience in IT operations, application support (2nd/3rd line), or a similar production-facing role
Proven experience managing incidents end-to-end, from alerting and troubleshooting through RCA and prevention
Minimum 2 years of experience working within ITIL processes including incident, problem, and change management
Experience working in Agile delivery environments alongside development teams
Excellent English communication skills (C1)
Strong knowledge of Jenkins for building, maintaining, and troubleshooting deployment pipelines
Hands-on experience with CI/CD pipelines, automation, and continuous improvement initiatives
Excellent troubleshooting and problem-solving skills for complex production issues
Proficiency with Splunk and Sysdig for log analysis and alerting
Strong understanding of Prometheus and Grafana for monitoring, dashboards, and alert tuning
Practical experience operating services running on Kubernetes, including pod health checks, log analysis, and service restarts
Expertise in Ansible for controlled configuration changes in operational environments
Strong knowledge of Docker and Docker Compose
Basic scripting skills in Bash and Python for operational automation and data reconciliation
Nice to have:
Experience with IBM Datastage operations
Awareness of or willingness to learn Pega and Airflow
Experience with Oracle and DB2, including querying, execution plan interpretation, and data incident analysis
Understanding of ETL application behavior and REST API communication
Experience supporting distributed systems and Kafka-based message flows
Java or development background supporting understanding of solutions and integrations
Offer:
Private medical care
Co-financing for the sports card
Constant support of dedicated consultant
Employee referral program

