{"id":1133917,"url":"https://alion.io/job/avra-member-of-technical-staff-observability-reliability","title":"Member of Technical Staff | Observability & Reliability","company":{"id":670137,"name":"Avra","domain":"avra.ai","url":"https://alion.io/company/avra","size_band":"201-500","is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","hiring_geo_confidence":"inferred","locations":["São Paulo, Brazil"],"countries":["BR"],"hiring_countries":["BR"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":77000,"max_usd":172000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":775},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Helm","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Ray","optional":false},{"name":"Terraform","optional":false},{"name":"Amazon EKS","optional":true},{"name":"AWS","optional":true},{"name":"GCP","optional":true},{"name":"Google GKE","optional":true}],"status":"live","first_seen_at":"2026-09-23T05:07:57Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-24T09:28:57Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About the role\nAt Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.\nIn this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments - so observability and reliability have to work the same way everywhere.\nWhat you'll do\nEvolve our observability stack for logs, metrics, traces, and alerting.\n\nMake sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.\n\nBring telemetry into customer clusters within a model where agents only make outbound connections.\n\nDetect drift between the desired state and what's actually running in each environment.\n\nMonitor the health of our deployment and runtime agents.\n\nProvide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.\n\nDefine SLOs, lead incident response and postmortems, and reduce MTTR - including when a fix requires coordinating with the customer.\n\nReduce telemetry cost: less redundant data, more useful signal.\n\nHow we measure success\n99.9% serving availability, with incidents trending down.\n\nMTTR, including on-premise incidents.\n\nNear-zero drift between desired and actual state.\n\nAll agents active and reporting, across every dataplane.\n\nWhat we're looking for\nDeep experience with OpenTelemetry and observability backends.\n\nHands-on practice with SLOs, error budgets, actionable alerting, and incident management.\n\nStrong experience with Kubernetes and infrastructure as code (Terraform / Helm ).\n\nExperience operating software in environments you don't fully control.\n\nProduction-quality code and reviews, and a willingness to operate what you build.\n\nNice to have\nShipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).\n\nGCP or GKE, AWS or EKS.\n\nML multi-node/multi-cluster workloads in production.\n\nFinancial services or regulated environments.","description_format":"text","description_chars":2335,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Brazil","iso":"BR","kind":"country"},{"name":"São Paulo","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","LLM & Generative AI","Foundation Models"],"lifecycle":[{"event":"open","at":"2026-09-23T05:58:52Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":18,"reasons":["conf:0","win:early"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/avra-member-of-technical-staff-observability-reliability","json_url":"https://alion.io/job/avra-member-of-technical-staff-observability-reliability.json","meta":{"generated_at":"2026-09-24T10:06:19Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}