{"id":1356809,"url":"https://alion.io/job/levi-strauss-site-reliability-engineer","title":"Site Reliability Engineer","company":{"id":1780263,"name":"Levi Strauss & Co.","domain":"levistrauss.com","url":"https://alion.io/company/levistrauss","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Mexico"],"countries":["MX"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":41000,"max_usd":106000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":2013},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"ArgoCD","optional":false},{"name":"Azure","optional":false},{"name":"BigQuery","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"GCP","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Google BigQuery","optional":false},{"name":"Google Cloud Run","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"Kubernetes","optional":false},{"name":"LLM","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Terraform","optional":false},{"name":"FinOps","optional":true}],"status":"live","first_seen_at":"2026-09-21T00:00:00Z","employer_posted_date":"2026-09-21","last_verified_at":"2026-09-29T23:46:08Z","board_verified":true,"closed_at":null,"days_open":9,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":9},"description":"Job Location:Mexico City, Mexico\nCalling all originals: At Levi Strauss & Co., you can be yourself - and be part of something bigger.We’rea company of people who like to forge our own path and leave the world better than we found it. Whobelievethat what makes us different makes usstronger.Soadd your voice. Make an impact. Find your fit - and your future.\nWe'reseekinga curious and drivenSite Reliability Engineer to join our Data & AI Platform Engineering team. In this role,you'llhelp keep our data and AI platforms running reliably, efficiently, and securely - platforms that power decisions across our global retail operations.\nYou'llwork alongside experienced SREs and engineers tomonitorproduction systems, respond to incidents, reduce operational toil, and build the automation that makes our infrastructure more resilient. This is an excellent opportunity to grow your SRE craft in a fast-paced, collaborative environment on Google Cloud Platform, with exposure to multi-cloud technologies and modern data engineering.\nAbout the Job\nReliability & Incident Response\nMonitor production systems usingobservability tooling - dashboards, alerts, and logs - to detect and triage issues before theyimpactend users\n\nParticipate inon-call rotations, respond to incidents following established runbooks, and escalate appropriately when needed\n\nContribute toblameless post-mortems, documenting rootcausesand follow-up action items to prevent recurrence\n\nHelpmaintainand improveSLO dashboards and alerting thresholds to ensure platform health is visible and measurable\n\nToil Reduction & Automation\nIdentifyrepetitive manual tasks and buildautomation toeliminatethem, reducing toil for yourself and the broader team\n\nWrite andmaintainscripts, tooling, andCI/CD pipeline components that improve deployment reliability and operational efficiency\n\nSupportself-serve infrastructure initiatives that allow engineering teams to safely provision and manage their own resources\n\nPlatform Operations & Cloud Infrastructure\nOperate andmaintainworkloads running onGCP - including GKE, Cloud Run,BigQuery, Pub/Sub, GCS, and Composer\n\nApply Infrastructure-as-Code practices (Terraform, Helm) to consistently and safely manage and version infrastructure changes\n\nSupportmulti-cloud awareness across GCP and Azure, following team standards for consistency and security across environments\n\nAdhere todata security and governance policies - IAM best practices, secrets management, encryption, and audit logging\n\nCollaboration & Growth\nWork closely with Data Engineering, AI Platform, and Software Engineering teams to ensure reliability is considered from design through deployment\n\nParticipate inreliability reviews, design discussions, and team ceremonies, contributing ideas and raising operational concerns early\n\nEngage withAI and agentic platform workloads, gaining exposure to the operational patterns of LLM-based systems and data pipelines\n\nContinuously develop your technical skills and SRE craft, supported by team knowledge-sharing, documentation, and hands-on experience\n\nAbout You\nRequired Qualifications\nBachelor's degree in Computer Science, Engineering, or related field (or equivalent practical experience)\n\n6+ years of experience in Site Reliability Engineering, DevOps, or Platform/Infrastructure Engineering in production environments\n\nHands-on experience with GCP services - particularly GKE, Cloud Run,BigQuery, Pub/Sub, and GCS\n\nWorkingproficiencywithInfrastructure-as-Code tools such as Terraform or Helm\n\nFamiliarity with observability tooling - metrics, logging, tracing, and alerting (e.g., Cloud Monitoring, Datadog, or Prometheus/Grafana)\n\nUnderstanding ofSLO/SLI concepts and how they relate to production reliability and on-call operations\n\nExposure to data security fundamentals: IAM, encryption, secrets management, and network policies\n\nProficiencyin at least one scripting or systems language (Python, Bash, or Go) for automation and operational tooling\n\nStrong communicationskills with the ability to clearly document incidents, runbooks, and technical processes\n\nTechnical Familiarity\nExperience withcontainer orchestration - Kubernetes or GKE - and the operational patterns around deploying and managing containerized workloads\n\nBasic understanding of CI/CD pipelines andGitOpsworkflows (ArgoCD, GitHub Actions, or similar)\n\nComfort working withdata platforms - familiarity with batch or streaming data pipelines is a plus\n\nAwareness ofmulti-cloud concepts, particularly across GCP and Azure\n\nDesirable Experience\nExperience working inretail, e-commerce, or consumer goods environments\n\nFamiliarity withGoogle's SRE principles - error budgets, toil tracking, and production readiness reviews\n\nExposure toAI or ML platform operations, including monitoring model serving infrastructure\n\nExperience withFinOps or cloud cost visibility tooling\n\nWhy Join Us?\nIfyou'rean engineer who is passionate about reliability, loves solving operational problems, and wants to grow your SRE craft at a global iconic brand,we'dlove to hear from you.\nLOCATION\nMexico, D.F., MexicoFULL TIME/PART TIME\nFull time Current LS&Co Employees, apply via your Workday account.","description_format":"text","description_chars":5165,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Mexico","iso":"MX","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Sportswear & Activewear"],"lifecycle":[{"event":"open","at":"2026-09-27T22:39:52Z"}],"liveness":{"score":75,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.833,"p_room":0.9,"age_days":9,"expected_fill_days":20,"reasons":["conf:5","velocity","win:mid","comp:brand"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/levi-strauss-site-reliability-engineer","json_url":"https://alion.io/job/levi-strauss-site-reliability-engineer.json","meta":{"generated_at":"2026-09-30T06:44:47Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4807,"day_limit":5000,"remaining_today":193,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}