667,602open jobs
39,042companies
100,809added this week
Browse all
Salary
$122k – $217k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
Redwood Software is an enterprise automation and orchestration provider headquartered in Houten, Netherlands, and founded in 1993. The company develops a cloud-native platform for workload automation, job scheduling, and managed file transfer, featuring products like RunMyJobs and specialized SAP finance automation tools. It serves over 3,000 global organizations, including half of the Fortune 50, and focuses on enabling autonomous enterprise operations through AI-powered orchestration.

OUR MISSION

At Redwood, we empower our customers with lights-out automation for their mission-critical business processes.

ABOUT US

Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric-with AI embedded across the automation lifecycle-Redwood accelerates the path to autonomous operations. Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next.

CORE VALUES

  • One Team. One Redwood
  • Make Your Own Weather
  • Obsess over Customer Success
  • Work the Problem
  • Be Curious
  • Own the Outcome
  • Respect Each Other

YOUR IMPACT

As a Site Reliability Engineer (SRE) / DevOps Engineer you will be responsible for ensuring the stability, performance, scalability, and reliability of Redwood’s mission-critical SaaS platform. You will apply engineering principles to operational challenges, automate repetitive work, strengthen observability, and collaborate across teams to deliver a resilient and high-performing customer experience.

  • Provide day-to-day management of system alerts, monitor system health, and escalate issues as necessary to maintain high availability.
  • Participate in a 24x7, team-shared on-call rotation for critical SaaS platform incidents and provide support during emergencies.
  • Lead incident response efforts to ensure fast and effective mitigation and resolution of production issues.
  • Perform thorough Root Cause Analysis (RCA) and lead blameless post-mortems to identify systemic weaknesses and establish corrective actions that prevent recurrence.
  • Collaborate with engineering teams to establish and enforce error budgets derived from Service Level Objectives (SLOs), balancing development velocity with system stability.
  • Automate routine operational tasks to reduce manual effort and toil while increasing team efficiency.
  • Design, deploy, and maintain cloud infrastructure using Infrastructure as Code (IaC), leveraging Terraform and Helm for deployment to EKS/Kubernetes clusters.
  • Design, secure, and troubleshoot AWS cloud network architecture, including VPCs, subnetting, routing, security groups/NACLs, load balancers (ALB/NLB), and VPN/Transit Gateway connectivity across a multi-region, multi-account environment supporting EKS and Docker Swarm on EC2.
  • Improve infrastructure health by developing and implementing checks, scripts, and automated remediation to proactively address known issues and enable platform self-healing.
  • Maintain, develop, and evolve Continuous Integration/Continuous Delivery (CI/CD) deployment code and pipelines.
  • Maintain existing infrastructure running on Docker and Docker Swarm while contributing to migration strategies toward EKS/Kubernetes.
  • Implement and integrate new technologies and services into Redwood’s cloud infrastructure to enhance platform capabilities and resilience.
  • Design and implement comprehensive observability strategies across metrics, logs, and traces.
  • Create and refine robust monitoring and alerting configurations within the EKS/Kubernetes ecosystem.
  • Utilize and maintain Datadog to gather performance data, create synthetic tests, configure monitoring, and visualize system health through dashboards.
  • Leverage existing monitoring solutions, including Grafana and Prometheus, while supporting the migration or integration of data into a unified observability platform.
  • Document issues, remediation steps, system architecture, and runbooks to support knowledge transfer and rapid incident response.
  • Collaborate closely with Support, Customer Success, Migration, and Professional Services teams to deliver a high level of SaaS service and minimize customer impact during changes.
  • Maintain a strong customer focus when planning deployments and updates, considering the impact on end users before implementing changes.

YOUR EXPERIENCE

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, ideally supporting a production SaaS platform.
  • Hands-on AWS Cloud Engineering experience with strong working knowledge of the AWS ecosystem, including AWS IAM roles and policies.
  • Strong knowledge of AWS cloud networking, including VPC design and peering, subnetting, security groups/NACLs, route tables, ALB/NLB load balancers, Transit Gateway/VPN connectivity, and Route 53 DNS.
  • Ability to design and troubleshoot connectivity across multi-region, multi-account AWS environments.
  • Proficiency with EKS/Kubernetes (K8s) and container orchestration technologies.
  • Demonstrated experience with Infrastructure as Code (IaC), specifically Terraform and Helm.
  • Working experience with Docker and maintaining systems using Docker Swarm.
  • Expertise in implementing and managing logging, monitoring, and observability solutions.
  • Direct experience with Datadog, including APM, infrastructure monitoring, synthetic testing, and custom dashboards.
  • Experience with Grafana and/or Prometheus is a plus.
  • Proficiency working in Linux environments with strong Bash and/or Python scripting skills for automation and troubleshooting.
  • Operational experience with relational databases in production environments, such as AWS RDS/PostgreSQL, including performance tuning, backup and restore, and troubleshooting under load.
  • Experience providing product or application support for high-availability SaaS products.
  • Experience designing, implementing, and operating within a DevSecOps environment.
  • Excellent written and verbal communication skills, with the ability to clearly explain complex technical issues and Root Cause Analyses to both technical and customer-facing audiences.

If you like growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

THE LEGAL BIT

Redwood is an equal opportunity employer. Redwood prohibits unlawful discrimination based on race, colour, religion, sex, gender identity, marital or veteran status, age, national origin, ancestry, citizenship, physical or mental disability, medical condition, genetic information or characteristics (or those of a family member), sexual orientation, pregnancy or any other consideration made unlawful by regional or local laws. We also prohibit discrimination based on a perception that anyone has any of those characteristics or is associated with a person who has or is perceived as having any of those characteristics. All such discrimination is unlawful and will have a zero tolerance policy applied to it.

Redwood will comply with all local data protection laws, including GDPR when it comes to the handling and processing of personal data. Should you wish for us to remove your personal data from our recruitment database, please email us directly at [email protected]

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
667,602 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Lead Engineer - India 3 hours ago
$32k – $73k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree
Python
SQL
AI/ML
LangGraph
LangChain
AI Agents
CrewAI
OpenAI
Hugging Face
Multi-Agent Systems
DevOps
Terraform
Datadog
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
Apply
Software Engineer 2 3 hours ago
$24k – $54k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
SQL
Python
Django
Django REST Framework
Databases
MySQL
PostgreSQL
OpenSearch
AI/ML
LangGraph
LangChain
AI Agents
NLP
RAG
Semantic Search
Semantic Search
Frontend
Next.js
React.js
DevOps
Terraform
Azure DevOps
GitHub Actions
Datadog
Azure
CI/CD
AWS
Docker
Vector
API Gateway
Management
Agile
Scrum
Apply
$136k – $200k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Milpitas
Python
C
C
MPI
Databases
ElasticSearch
AI/ML
CUDA Toolkit
OpenCL
CUDA
InfiniBand
DevOps
Puppet
Ansible
Red Hat
Chef
Prometheus
Jenkins
Git
Docker
Ubuntu
Grafana
Apply
$95k – $159k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • United States
JavaScript
Java
Kotlin
TypeScript
SQL
Java
Spring Framework
Maven
Spring Boot
Gradle
Micronaut
Kotlin
Mockito
Databases
PostgreSQL
DynamoDB
Frontend
Angular
npm
Mobile
JUnit
DevOps
Terraform
Azure DevOps
GitHub Actions
Azure
CI/CD
Git
AWS
Docker
Nginx
Amazon Kinesis
Amazon EventBridge
AWS Step Functions
Cybersecurity
SonarQube
Management
Agile
QA
Cypress
Apply
$66k – $135k per year (Estimated) • Remote • Full-Time • Hungary
JavaScript
Java
AI/ML
LLM
LLM Evaluation
Frontend
React.js
Lit
DevOps
Terraform
AWS
Apply
$58k – $132k per year (Estimated) • Remote • Bachelor's Degree
Cybersecurity
GDPR
Apply
$109k – $227k per year (Estimated) • Remote • 7+ years exp
AI/ML
AI Agents
DevOps
AWS
Cybersecurity
GDPR
Apply
$75k – $175k per year (Estimated) • Remote • 3+ years exp • Bachelor's Degree
AI/ML
Copilot
AI Agents
Cybersecurity
GDPR
Marketing
Salesforce
Apply
$82k – $154k per year (Estimated) • Remote • 3+ years exp
Cybersecurity
GDPR
Apply
HR Manager 12 days ago
$67k – $84k per year • In office • 5+ years exp • Bachelor's Degree
AI/ML
Claude
ChatGPT
Cybersecurity
GDPR
Apply
See all jobs
This is one of many
667,602 more open roles from verified company boards, updated every day.