{"id":1241648,"url":"https://alion.io/job/revyl-staff-infrastructure-engineer-device-cloud","title":"Staff Infrastructure Engineer — Device Cloud","company":{"id":2913065,"name":"Revyl","domain":"revyl.com","url":"https://alion.io/company/revyl","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Work at a Startup","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"staff","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":{"min":150000,"max":250000,"currency":"USD","period":"year","gross":null,"usd_annual":250000},"salary_estimate":null,"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"GCP","optional":false},{"name":"Kubernetes","optional":false},{"name":"Platform Engineering","optional":false}],"status":"live","first_seen_at":"2026-09-25T16:49:20Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-26T23:51:17Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"As a Staff Infrastructure Engineer at Revyl, you'll own the technical strategy and systems required to scale the infrastructure that powers our mobile development platform.\nRevyl gives developers and AI agents access to cloud iOS and Android environments where they can build, run, inspect, and verify mobile applications. Behind that experience is a distributed compute platform spanning GCP, physical Mac infrastructure, simulators, emulators, build runners, streaming, and test execution.\nWe're entering a period where our infrastructure needs to support significantly more customers, workloads, and compute. You'll be responsible not only for operating that infrastructure, but for figuring out how it needs to evolve.\nYou'll dig into how our platform operates today, identify bottlenecks and scaling limits, understand where operational toil and reliability issues are coming from, model future capacity requirements, and develop the technical roadmap that gets us from where we are today to where we need to be.\nThis isn't a traditional DevOps or SRE role. We're looking for someone who can move fluidly between understanding the current system, designing the next version of it, and getting hands-on to build it.\nWhat you'll work on\nDefine our infrastructure scaling strategy. Understand the architecture and operational characteristics of Revyl's platform today, identify the systems that will break as load increases, and develop a roadmap for scaling them ahead of demand.\n\nCapacity planning and modeling. Translate expected customer and partnership growth into compute, storage, network, device, and hardware requirements. Build the models and instrumentation that let us understand when and where we need to add capacity.\n\nScaling our device cloud. Design and build the systems that provision, schedule, manage, and recycle large fleets of iOS simulators, Android emulators, and physical compute.\n\nDistributed workload orchestration. Evolve how builds, tests, agent sessions, and interactive workloads are scheduled across heterogeneous pools of compute as concurrency grows.\n\nCloud infrastructure. Scale and evolve the GCP infrastructure behind Revyl's APIs, control plane, workflow execution, storage, networking, and supporting services.\n\nIdentify and eliminate bottlenecks. Use production data to understand where we're constrained by CPU, memory, disk, network, scheduler throughput, database performance, host capacity, or architecture-and determine the right solution rather than simply adding more machines.\n\nBuild for step-function growth. Prepare Revyl for customers and partnerships that can introduce significantly more traffic than the platform handles today. Design load tests, failure tests, and capacity plans that give us confidence before traffic arrives.\n\nWhat we're looking for\nWe are looking for someone who has taken an infrastructure platform from one stage of scale to the next.\nYou likely have:\nHas 5+ years of software engineering, infrastructure, platform engineering, or SRE experience, with significant technical ownership.\nHas previously helped scale a production infrastructure or compute platform through a meaningful increase in load.\nCan analyze an existing architecture, identify scaling constraints, and turn those findings into a prioritized technical roadmap.\nUnderstands capacity planning and can translate business forecasts and expected workload growth into infrastructure requirements.\nIs capable of making architectural decisions under uncertainty and knows how to validate assumptions through instrumentation, benchmarking, load testing, and production data.\nHas built or operated distributed systems at meaningful scale.\nIs a strong software engineer who can build infrastructure services and control-plane systems, not just configure infrastructure tooling.\nUnderstands scheduling, queues, worker pools, concurrency, backpressure, resource allocation, failure recovery, and distributed systems fundamentals.\nHas deep experience with cloud infrastructure such as GCP, AWS, or Azure.\nWhat success looks like\nWithin your first year, we'd expect you to have materially changed both Revyl's infrastructure and our understanding of how it scales.\nThat means:\nWe have a clear model of the capacity and architecture required to support our next stages of growth.\nWe know where our major scaling limits are before customers encounter them.\nLarge new customers and partnerships have concrete capacity and load plans before launch.\nRevyl can handle an order of magnitude more concurrent compute without requiring an order of magnitude more operational work.\nCommon infrastructure failures are automatically detected and remediated.\nDevice scheduling and capacity management are predictable rather than reactive.\nWe can confidently answer how much additional traffic the platform can support.\nInfrastructure projects are prioritized based on real scaling constraints rather than whichever fire happened most recently.\nOur existing SRE and engineering team spend substantially less time firefighting.\nThe Device Cloud becomes a platform we can scale repeatedly rather than an infrastructure stack that needs to be reinvented every time Revyl grows.","description_format":"text","description_chars":5193,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","AI Agents","Ride Hailing"],"lifecycle":[{"event":"open","at":"2026-09-25T16:49:20Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":22,"reasons":["conf:0","win:early"],"computed_at":"2026-09-26T05:45:00Z"},"pay":{"stated_usd_annual":250000,"is_top_pay":true},"html_url":"https://alion.io/job/revyl-staff-infrastructure-engineer-device-cloud","json_url":"https://alion.io/job/revyl-staff-infrastructure-engineer-device-cloud.json","meta":{"generated_at":"2026-09-27T01:29:21Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1332,"day_limit":5000,"remaining_today":3668,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}