{"id":1133916,"url":"https://alion.io/job/avra-member-of-technical-staff-inference-platform","title":"Member of Technical Staff | Inference Platform","company":{"id":670137,"name":"Avra","domain":"avra.ai","url":"https://alion.io/company/avra","size_band":"201-500","is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","hiring_geo_confidence":"inferred","locations":["São Paulo, Brazil"],"countries":["BR"],"hiring_countries":["BR"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":80000,"max_usd":179000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":775},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Fine-tuning","optional":false},{"name":"Kubernetes","optional":false},{"name":"Post-training","optional":false},{"name":"Python","optional":false},{"name":"Sophos","optional":false},{"name":"Amazon EKS","optional":true},{"name":"AWS","optional":true},{"name":"GCP","optional":true},{"name":"Google GKE","optional":true},{"name":"Ray","optional":true},{"name":"Ray Serve","optional":true}],"status":"live","first_seen_at":"2026-09-23T05:09:45Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-24T09:28:57Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About the role\nAt Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.\nIn this role, you'll join the Platform team to own where our models execute. Customers consume our models through large batches of millions of records and through real-time APIs, and they make business decisions on every response. You'll run governed model releases reliably and efficiently - in our cloud and on customer-hosted Kubernetes - and make inference fast, predictable, and cheap enough to serve both enterprise and mid-market customers.\nWhat you'll do\nEvolve Sophos, our online and batch inference runtime, built on Kubernetes.\n\nRun large batch inference on ephemeral jobs, with multi-dimensional admission control (CPU, memory, GPU) through Kueue.\n\nBuild and extend the Sophos controller and its Kubernetes custom resources.\n\nOptimize each model's inference engine and feature processing, using vectorized, columnar operations.\n\nServe graphs and data efficiently from Lance-based storage.\n\nOwn execution of training, post-training, and fine-tuning jobs, in our cloud and in customer dataplanes.\n\nDrive autoscaling, GPU serving, performance, and cost optimization, with telemetry for every model we run.\n\nSolve open problems such as deterministic job sizing, checkpointing and recovery for batch runs, per-customer encryption and isolation, resilience to difficult input files, and automatic profiling when a new model is accepted.\n\nHow we measure success\n99.9% serving availability.\n\np95/p99 latency for online inference and throughput for batch.\n\nCost per prediction and per training job.\n\nGPU utilization: paid capacity versus capacity actually used.\n\nTraining and batch jobs that finish on time and succeed without manual retries.\n\nWhat we're looking for\nExperience running model serving or large-scale batch compute on Kubernetes.\n\nExperience building Kubernetes controllers or operators.\n\nSkill at profiling and optimizing data-heavy Python pipelines.\n\nA clear sense of cost: you treat compute efficiency as a product feature.\n\nProduction-quality code and reviews, and a willingness to operate what you build.\n\nNice to have\nRay, Ray Serve, or KubeRay in production.\n\nKueue or other batch scheduling and admission-control systems.\n\nGPU serving and performance optimization.\n\nArrow, Parquet, Lance, or other columnar formats.\n\nShipping software to customer-hosted Kubernetes.\n\nGCP/AWS and GKE/EKS, and financial services or regulated environments.","description_format":"text","description_chars":2654,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Brazil","iso":"BR","kind":"country"},{"name":"São Paulo","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","LLM & Generative AI","Foundation Models"],"lifecycle":[{"event":"open","at":"2026-09-23T05:58:52Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":18,"reasons":["conf:0","win:early"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/avra-member-of-technical-staff-inference-platform","json_url":"https://alion.io/job/avra-member-of-technical-staff-inference-platform.json","meta":{"generated_at":"2026-09-24T10:12:37Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}