{"id":1486725,"url":"https://alion.io/job/prolaio-site-reliability-engineer","title":"Site Reliability Engineer","company":{"id":688731,"name":"Prolaio","domain":"prolaio.com","url":"https://alion.io/company/prolaio","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"junior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Chicago, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":66000,"max_usd":142000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":232},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"BigQuery","optional":false},{"name":"GCP","optional":false},{"name":"Google BigQuery","optional":false},{"name":"Python","optional":false},{"name":"SQL","optional":false},{"name":"Datadog","optional":true},{"name":"Grafana","optional":true},{"name":"Wi-Fi","optional":true}],"status":"live","first_seen_at":"2026-09-29T20:05:20Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-09-30T18:13:21Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Who Are We?\nProlaio believes that continuous learning and collaboration can make a significant difference in how heart care is administered. We are creating smarter ways to address heart disease and heart risks by uniting patients, care teams, and researchers on a secure, technology-enabled platform that drives clinical innovation and offers a path towards better patient outcomes.\nThis is precision cardiology, and we know it’s within reach.\nWhat Will You Do?\nThe Overview\nThe Site Reliability Engineer will ensure that Prolaio’s cardiovascular data platform reliably captures, transports, and delivers continuous biosensor data from participants to the clinical researchers and trial sponsors who depend on it. In this role, you will build and operate the monitoring, service level objectives, incident response, and automation needed to keep critical data flows healthy, identify failures before they impact studies, and ensure that device downtime never goes unnoticed simply because a shipment record says a device was delivered.\nThis role is ideal for a self-starter who thrives at the intersection of software reliability, connected devices, and clinical data. You will investigate incidents across the full participant-to-platform stack, combining service telemetry with device data and participant reports to determine whether an issue originated with the hardware, phone, firmware, mobile application, connectivity, or the platform itself. You will turn those investigations into durable improvements, automating manual checks, strengthening observability, reducing operational toil, and building the reliability practices that keep Prolaio’s clinical data complete and trustworthy at scale.\nThe Specifics\nBuild monitoring that treats missing data as a failure rather than a quiet day - freshness, volume and quality checks that verify delivery from outside the system instead of trusting a job's exit code.\nDefine and maintain SLIs, SLOs and error budgets for customer-facing data services, and report attainment honestly, including when the news is bad.\nJoin the on-call rotation, run incident response, and write blameless postmortems that produce owned follow-up actions.\nInvestigate device failures end to end: correlate discharge curves, Bluetooth disconnects and wear detection against the participant's complaint, and establish or exonerate each candidate cause on evidence.\nTreat the fleet as a population - failure rates per device-month, clustering by build lot, firmware version or site - rather than as a queue of individual tickets.\nClose the loop. When an analysis cannot reach a verdict, name the missing metric or threshold and route it back into the diagnostics pipeline so the next case is answerable from telemetry alone.\nBuild reliability tooling in Python and SQL on BigQuery, and keep monitors and probes defined as code rather than configured by hand.\nRetire manual checks by automating them. Operational toil should shrink quarter over quarter, not accumulate.\nWhy Prolaio?\n Impactful Work: You will join in the fight against heart failure (HF) and hypertrophic cardiomyopathy (HCM) with the goal of extending and saving the lives of our patients while also being at the forefront of changing the healthcare industry through technology.\nInnovative Environment: You will be part of an organization doing something that’s never been done before.\nProfessional Growth: You will join a growing team and have a substantial impact on our daily and future operations with the opportunity to continuously learn and grow.\nCollaborative Team: You will be part of a team of collaborative, curious, and committed individuals focused on the collective good, inclusiveness, scientific excellence, and advancing digital health for cardiology.\nWho You Are?\nBachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, Biomedical Engineering or a related technical field, or equivalent practical experience.\n2-5 years in reliability engineering, production operations, hardware or systems test, field failure analysis, or data engineering - in a role where you were accountable for something that had to keep working.\nWorking SQL and Python: able to pull a dataset, characterize it, and defend the query behind a number.\nExperience with a cloud data platform (Google Cloud and BigQuery preferred; AWS or Azure equivalents are fine).\nAbility to write up a technical investigation clearly - what was observed, what was ruled out and how, and what remains unknown.\nAdditional Qualifications (Nice to Haves)\nHands-on exposure to Bluetooth Low Energy (GATT services and characteristics, connection intervals, pairing and bonding, disconnect behaviour), Wi-Fi or cellular data standards - and how each one fails in the field rather than how it is specified to behave.\nBattery and power behaviour: discharge characterisation, charge-state interpretation, and telling a measurement artefact from a real capacity defect.\nFormal reliability method - FMEA and risk priority numbers, FRACAS practice, and standards such as IEC 60812 or ISO 13485.\nObservability tooling (Datadog, Grafana or similar) and synthetic monitoring.\nAndroid device telemetry, mobile application diagnostics, or embedded firmware experience.\nPrior work in medtech, digital health, clinical trials or another regulated, patient-facing environment.\nWhy You’ll Love Working Here\nMeaningful Compensation: Competitive salary, performance bonus, and equity so you can share in what we build.\nGreat Health Coverage: Medical, dental, and vision plans with multiple options and strong company contributions.\nFlexible Spending Perks: HSA, FSA, commuter benefits, and a $1,200 annual Lifestyle Spending Account to support wellness, commuting, family needs, and more.\nTime to Recharge: Generous paid time off, sick leave, and company holidays.\nFamily-First Benefits: Paid parental leave, caregiver leave, and support for growing families.\nSecurity & Peace of Mind: Company-paid life insurance and short- and long-term disability coverage.\n Plan for the Future: 401(k) plan to help you build long-term financial security.\n Care When You Need It: Easy access to telehealth and optional supplemental coverage for life’s unexpected moments.\nStarting Salary is at $ 118,000.00 (Exact Compensation may vary based on skills, experience, and location)","description_format":"text","description_chars":6336,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Continuous learning","Equity","Life insurance","Parental leave"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Patient Monitoring","Health Data & Interoperability"],"lifecycle":[{"event":"open","at":"2026-09-29T23:34:50Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":0,"expected_fill_days":15,"reasons":["conf:6","velocity","win:early","comp:junior"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/prolaio-site-reliability-engineer","json_url":"https://alion.io/job/prolaio-site-reliability-engineer.json","meta":{"generated_at":"2026-10-01T02:37:13Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1921,"day_limit":5000,"remaining_today":3079,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}