First seen by Alion on Oct 4, 2026.
Craftware is a technology company of over 500 experts, empowering large organizations to solve complex business challenges with modern IT solutions - from sales systems and automation to data platforms and AI. We operate where technology must be reliable, secure, and scalable. We deliver end-to-end projects: from analysis and architecture through implementation to development and maintenance. We are a trusted partner of industry leaders such as Salesforce, Veeva, UiPath, and Databricks.
Model: Remote
Engagement: Full-time (B2B)
We are looking for an experienced Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational continuity of a complex enterprise ecosystem consisting of multiple highly integrated systems. The SRE will be responsible for the technical stability of individual applications as well as the end-to-end reliability of critical business processes spanning multiple platforms and services, working closely with DevOps teams, application owners, architects, support teams, and Single Points of Contact (SPOCs) responsible for integrated and dependent systems mostly based on Salesforce platform with Sales Cloud, Service Cloud and Experience Cloud.
Responsibilities
Ensure high availability, reliability, performance, and operational continuity of business-critical systems and services
Monitor and analyze end-to-end business processes across multiple applications, APIs, middleware components, messaging platforms, and external systems
Identify system dependencies and assess their impact on business service availability
Establish and maintain monitoring, observability, alerting, and dashboards covering both technical and business-process metrics
Define and track SLIs, SLOs, availability targets, and other reliability metrics
Coordinate major incidents involving multiple systems and technical teams
Lead root cause analysis for production incidents, integration failures, performance degradation, and service interruptions
Coordinate troubleshooting and communication with SPOCs, system owners, vendors, infrastructure teams, and external providers
Manage and prioritize the work of the DevOps team responsible for deployment, monitoring, automation, infrastructure, and operational support
Drive automation of operational activities, deployments, health checks, recovery procedures, and system maintenance
Maintain operational runbooks, troubleshooting guides, escalation paths, and recovery procedures
Ensure appropriate backup, disaster recovery, failover, and business continuity mechanisms are implemented and validated
Support release planning, production readiness, risk assessment, dependency analysis, and rollback strategies
Proactively identify reliability risks, performance bottlenecks, single points of failure, and architectural weaknesses
Work with development and architecture teams to improve resilience, scalability, fault tolerance, retry mechanisms, and graceful degradation
Lead post-incident reviews and ensure corrective and preventive actions are implemented
Requirements
A minimum of intermediate proficiency with the Salesforce platform is required.
Strong experience in Site Reliability Engineering, DevOps, Production Engineering, Application Operations, or a similar role
Experience working with complex, highly integrated enterprise architectures
Strong understanding of end-to-end business process monitoring and dependency management
Experience with incident management, root cause analysis, problem management, and service restoration
Practical knowledge of monitoring, logging, alerting, and observability platforms
Good understanding of SLI, SLO, SLA, availability, latency, throughput, and reliability concepts
Experience with CI/CD, release management, infrastructure automation, and deployment processes
Understanding of high availability, disaster recovery, failover, and resilience patterns
Ability to coordinate technical activities across multiple teams and system owners
Experience managing or coordinating a DevOps or operations-focused engineering team
Strong analytical, troubleshooting, and communication skills
Nice to have
Experience with cloud platforms (AWS, Azure, or GCP) in a production operations context
Experience with containerization and orchestration (Docker, Kubernetes)
Scripting/automation skills (Python, Bash, or similar)
Familiarity with APM/observability tooling (e.g., Datadog, Dynatrace, New Relic, Grafana, Prometheus)
Experience with messaging/integration middleware (e.g., Kafka, MQ, ESB platforms)
ITIL or similar IT service management framework knowledge
We offer
B2B contract (rate up to 190 PLN net/h + VAT)
Fully remote service delivery
Broad range of projects (internal, international) - genuine variety of clients and tasks
Budget for skills development and certifications as part of the collaboration
Regular collaboration reviews and discussion of project scope
Additional benefits available as part of the collaboration
Networking and team-building events for project teams
Tech stack
- Incident management
- SRE
- Observability
- CI/CD
- Team Leadership

