{"id":1133915,"url":"https://alion.io/job/avra-member-of-technical-staff-ml-systems","title":"Member of Technical Staff | ML Systems","company":{"id":670137,"name":"Avra","domain":"avra.ai","url":"https://alion.io/company/avra","size_band":"201-500","is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"staff","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","hiring_geo_confidence":"inferred","locations":["São Paulo, Brazil"],"countries":["BR"],"hiring_countries":["BR"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":78000,"max_usd":175000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":775},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"Ray","optional":false},{"name":"SkyPilot","optional":true}],"status":"live","first_seen_at":"2026-09-23T05:09:52Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-24T08:05:12Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About the role\nAt Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.\nIn this role, you'll join our ML Systems team, which owns Avra's ML core and the governance of every model we ship. Research produces candidate models and evidence; you build the reliable path from data and training to a governed, reproducible release that can run in our cloud or in any customer environment. ML Systems is an internal platform: its users are our researchers and platform engineers, and its success is measured by the leverage it creates for them.\nWhat you'll do\nBuild CUDA kernels and compute primitives for training and serving graph neural networks (GNNs).\n\nEvolve Monad, our sampler and distributed-training library, including neighbor sampling and training performance.\n\nSpecify our binary data formats (Lance, Arrow, CSR/CSC), and own materializations and feature backfills for training and evaluation.\n\nDefine data contracts and consumption requirements with the teams that build our customer and proprietary datasets.\n\nBuild and operate experiment tracking, checkpoints, and evaluation infrastructure, with reproducibility by default.\n\nOwn the model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models.\n\nDefine and run release gates, so every model running in production, batch, or on-premise maps to a governed release.\n\nMake it possible to audit exactly which data, code, configuration, and evidence produced each release.\n\nHow we measure success\nTime-to-experiment: how quickly a researcher goes from a hypothesis to materialized data, compute, and tracking.\n\nTime-to-governed-release: how quickly a validated candidate becomes an authorized release.\n\nTraining throughput per GPU on our foundation model training runs.\n\n100% of production models with complete release records and lineage - no ad hoc models in any environment.\n\nEvery release reproducible from its registered data, code, and configuration.\n\nWhat we're looking for\nStrong systems engineering skills and production-quality Python.\n\nExperience with distributed training (e.g., Ray, PyTorch distributed) and multi-node GPU workloads.\n\nExperience with columnar data formats and large-scale data materialization.\n\nFamiliarity with ML lifecycle tooling: experiment tracking, model registries, evaluation, and reproducibility.\n\nA product mindset: you treat an internal platform as a product with real users. You don't need to be a data scientist.\n\nNice to have\nCUDA kernel development or GPU performance optimization.\n\nGraph neural networks or graph sampling at scale.\n\nLance, Arrow, or other columnar/indexed storage formats.\n\nMulti-cloud GPU compute (e.g., SkyPilot).\n\nModel governance or audit requirements in financial services or other regulated environments.","description_format":"text","description_chars":2988,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Brazil","iso":"BR","kind":"country"},{"name":"São Paulo","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","LLM & Generative AI","Foundation Models"],"lifecycle":[{"event":"open","at":"2026-09-23T05:58:52Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":18,"reasons":["conf:0","win:early"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/avra-member-of-technical-staff-ml-systems","json_url":"https://alion.io/job/avra-member-of-technical-staff-ml-systems.json","meta":{"generated_at":"2026-09-24T09:21:38Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}