368,530open jobs
9,432companies
50,439added this week
Browse all
Location
Remote (Costa Rica)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
CSC Generation is a retail technology holding company headquartered in Merrillville, Indiana, and founded in 2016. The company acquires legacy and underperforming retail brands and integrates them into its proprietary Genesis platform, which uses artificial intelligence and automation to modernize operations. Its portfolio includes iconic brands such as Sur La Table, One Kings Lane, and Backcountry, operating as a multi-brand platform across North America.

Adventure is our Culture. Join a team that celebrates a lifestyle as bold as the terrain we love. At Backcountry, we are rooted in adventure, recognition, and wellbeing on and off the mountain. We spotlight employee stories, celebrate milestones, and offer exclusive outdoor perks. Whether you are at HQ, in a retail store, or remote, you will be part of a team that thrives on energy, exploration, and connection.

Reports to: Gustavo Arguedas (Site Reliability Manager)

Location: Remote - Costa Rica

About the Role

    Backcountry's online platform serves as the backbone of our customer experience, and this role exists to ensure its reliability, performance, and scalability. As Site Reliability Engineer, you will partner with software engineering, DevOps, and IT operations teams to optimize systems and applications across a multi-cloud stack.

    Within 6-12 months, you will have contributed meaningful improvements to service resiliency and observability, reduced operational toil through automation, and established yourself as a trusted partner to development and infrastructure teams.

    This is a lean team. You will own a lot, move fast, and make decisions with full end-to-end responsibility.

What You'll Do

  • Work on service resiliency, performance tuning, and system design across Backcountry's platform
  • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems
  • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories
  • Reduce toil by designing and implementing automation
  • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads
  • Monitor system health and capacity, taking proactive action to fix problems before they occur
  • Collaborate with engineering teams to build, deploy, and support features
  • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services
  • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy
  • Participate in the on-call support rotation within the SRE team

Required Qualifications

  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written

Preferred Qualifications

  • Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements
  • Familiarity with PCI-scoped or other regulated environments
  • Previous experience working in ecommerce environments
  • Professional certifications: GCP, CKA, or AWS

Why Join

    The people who do best here are builders. They take ownership, move fast, and want to see the direct impact of their work.

  • Cross-Functional Impact: Your work directly affects platform reliability for every customer and every team that depends on Backcountry's systems.
  • Modern Tech Stack: Work across a multi-cloud environment (GCP and AWS) with modern observability tooling, AI-assisted engineering, and GitOps workflows.
  • End-to-End Ownership: Own projects from investigation through implementation - you will ship automation, improve resiliency, and see the results in production.
  • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

Interview Process

    • Recruiter Screen - A 30-minute conversation with our recruiting team to align on the role, your background, and what you are looking for.
    • Hiring Manager Interview - Conversation with the Site Reliability Manager focused on your SRE experience, approach to incident management, and team fit.
    • Technical/Case Discussion - A deeper dive into infrastructure, observability, and problem-solving scenarios relevant to the role.
    • Reference Checks - Conducted in parallel with the final stages where possible.
    • Offer - We move quickly for the right candidate.

    Interview process is subject to change. Any updates will be communicated promptly and clearly.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$36k – $78k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Pune
PowerShell
Python
SQL
DevOps
AppDynamics
AWS
Azure
GCP
Grafana
Prometheus
Splunk
Apply
$37k – $80k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Pune
PowerShell
Python
AI/ML
Anomaly Detection
DevOps
AppDynamics
AWS
Azure
CI/CD
GCP
Grafana
Incident Management
Prometheus
Splunk
Apply
$23k – $57k per year (Estimated) • Remote • Full-Time • 6+ years exp
Python
Databases
OpenSearch
DevOps
AIOps
ArgoCD
AWS
Azure
CI/CD
Datadog
Docker
Dynatrace
GCP
GitHub Actions
Grafana
Incident Management
Jaeger
Jenkins
Kubernetes
New Relic
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
Splunk
Terraform
GitHub
Apply
$39k – $87k per year (Estimated) • In office • 9+ years exp • Bengaluru
Java
Python
Java
Spring Boot
Databases
ElasticSearch
Google BigQuery
Neo4j
PostgreSQL
AI/ML
Claude
Claude Code
Copilot
Cursor
LLM
AI Agents
Devin
Model Context Protocol
OpenAI Codex
DevOps
AWS
Datadog
GCP
Grafana
New Relic
Prometheus
Management
Jira
Apply
ML Engineer 1 day ago
Remote/Hybrid • 3+ years exp
Python
SQL
AI/ML
AI Agents
Computer Vision
MLFlow
PyTorch
RAG
A2A
Model Context Protocol
DevOps
AWS
Azure
CI/CD
Docker
Docker Compose
GCP
Kubernetes
Analytics
ETL/ELT
Apply
In office • Full-Time • San Jose
JavaScript
Node JS
Python
TypeScript
Python
Django
Databases
DynamoDB
Frontend
React.js
DevOps
AWS
AWS Lambda
Amazon CloudWatch
Amazon S3
API Gateway
IAM
Cybersecurity
Least Privilege
Apply
$98k – $205k per year (Estimated) • In office • Full-Time • Toronto
Python
SQL
AI/ML
Reinforcement Learning
LLM Guardrails
Apply
$143k – $276k per year (Estimated) • In office • Full-Time • Austin
Python
SQL
AI/ML
Reinforcement Learning
LLM Guardrails
Apply
$143k – $256k per year (Estimated) • Remote • Full-Time • United States
Python
SQL
AI/ML
Reinforcement Learning
LLM Guardrails
Apply
$173k – $328k per year (Estimated) • Remote/Hybrid • Full-Time • Austin
DevOps
FinOps
Management
Slack
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.