{"id":45071,"url":"https://alion.io/job/dataart-senior-site-reliability-engineer","title":"Senior Site Reliability Engineer","company":{"id":46984,"name":"DataArt","domain":"dataart.com","url":"https://alion.io/company/dataart-com","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":{"label":"EST","utc_offset_min":-5,"utc_offset_max":-5},"hiring_geo_confidence":"structured","locations":["Wrocław, Poland"],"countries":["PL"],"hiring_countries":["PL"],"hiring_countries_total":1,"salary":{"min":20000,"max":23500,"currency":"PLN","period":"month","gross":null,"usd_annual":75384},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"C++","optional":false},{"name":"CI/CD","optional":false},{"name":"Claude","optional":false},{"name":"Datadog","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"Java","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Terraform","optional":false},{"name":"Chaos Engineering","optional":true}],"status":"live","first_seen_at":"2026-08-07T07:49:50Z","employer_posted_date":null,"last_verified_at":"2026-08-07T07:49:50Z","board_verified":false,"closed_at":null,"days_open":58,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":58},"description":"Client\nOur client is developing a reliable, scalable, and user-friendly ticketing and streaming platform for high school sports. Their goal is to create a solution that allows parents, students, and fans to purchase tickets and stream live events effortlessly, ensuring accessibility from any device, anywhere, at any time.\nWorking Schedule: 8:00-17:00 Eastern Standard Time (EST). A 4-hour overlap with EST is required for effective collaboration.\nProject overview\nThe project focuses on building and maintaining a highly available, scalable cloud platform that supports business-critical services. Engineering teams continuously improve system reliability, performance, and operational excellence through automation, observability, and modern software delivery practices.\nPosition overview\nAs a Senior Site Reliability Engineer, you will work at the intersection of software engineering and operations, partnering with application, DevOps, and QA teams. You'll enhance observability, automate operational processes, improve CI/CD workflows, and drive reliability initiatives that enable teams to deliver resilient, high-performing software at scale.\nResponsibilities\nEnhance platform observability by designing and maintaining metrics, alerts, dashboards, and monitoring capabilities that improve system visibility and reduce incident resolution time.\n\nBuild and maintain automation, operational tooling, and monitoring solutions that increase service reliability and uptime.\n\nWork closely with software development and QA teams to embed reliability best practices into software delivery, release processes, and testing strategies.\n\nPromote operational excellence by driving preventive measures, facilitating blameless post-incident reviews, and supporting capacity and scalability planning.\n\nTake part in an on-call rotation, ensuring timely investigation and resolution of production incidents affecting critical services.\n\nRequirements\nStrong hands-on experience with Python, particularly for scripting, automation, and operational tooling.\n\nProficiency in at least one of the following programming languages: Java, C++, or Go.\n\nSolid knowledge of Linux environments, cloud platforms (AWS, GCP, or Azure), and containerized infrastructure using technologies such as Docker, Kubernetes, and Terraform.\n\nExperience designing and maintaining CI/CD pipelines, working with version control systems, and implementing automated testing practices.\n\nPractical experience with observability and monitoring platforms (such as Prometheus, Grafana, ELK, Datadog, or similar), including troubleshooting through log and metric analysis.\n\nExperience identifying and documenting Critical User Journeys and translating them into measurable SLA/SLO objectives that support automation and operational excellence.\n\nStrong collaboration and communication skills, with the ability to work effectively across multidisciplinary engineering teams, especially during critical production events.\n\nA reliability-first mindset with the belief that system stability is a shared responsibility across engineering teams.\n\nFamiliarity with AI-assisted engineering tools (such as Claude and Codex) and their use within modern software development workflows.\n\nNice to have\nExperience developing or maintaining end-to-end and integration tests for distributed or microservices-based systems.\n\nKnowledge of performance optimization, capacity management, or chaos engineering practices.\n\nExperience contributing to internal developer platforms, automation tools, or reliability engineering initiatives.\n\nUnderstanding of production security, compliance requirements, or change management processes.\n\nRelevant industry certifications.\n\nWhat We Offer:\nVacation days: Up to 26 business days per year.\n\n10 illness/special days\noff per year (fully paid, no medical papers needed) for all contract types\n\nHealth and life insurance (Luxmed)\n\nMyBenefit platform with Multisport option\n\nInternal psychological support service\n\nEnglish language classes from the first working day\n\nAccess to external learning platforms: O'Reilly, LinkedIn Learning, Udemy, and a wide catalog of diverse internal training\n\nFlexible workplace: work from the office, from home, or choose a hybrid option\n\nTech Skills Mentoring Program\n\nOpportunities to develop as a public speaker, mentor, or technical interviewer\n\nFully paid idle (bench) when not involved in a project\n\nCertification reimbursement (AWS, GCP, Microsoft, etc.)","description_format":"text","description_chars":4463,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Certification reimbursement","Flexible schedule","Life insurance"],"hiring_locations":[{"name":"Poland","iso":"PL","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["IT Outsourcing & Dedicated Teams","Custom Software Development"],"lifecycle":[{"event":"open","at":"2026-08-07T07:49:50Z"}],"visa":[],"liveness":{"score":19,"band":"cold","label":"Long shot","p_open":0.4,"p_active":0.841,"p_room":0.55,"age_days":57,"expected_fill_days":42,"reasons":["seen:57","velocity","win:tail"],"computed_at":"2026-10-04T05:45:00Z"},"pay":{"stated_usd_annual":75384,"is_top_pay":false},"html_url":"https://alion.io/job/dataart-senior-site-reliability-engineer","json_url":"https://alion.io/job/dataart-senior-site-reliability-engineer.json","meta":{"generated_at":"2026-10-05T00:37:34Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":809,"day_limit":5000,"remaining_today":4191,"minute_limit":60,"resets_at":"2026-10-06T00:00:00Z"}}}