{"id":792731,"url":"https://alion.io/job/heliosintel-staffsenior-distributed-systems-engineer","title":"Staff/Senior Distributed Systems Engineer","company":{"id":679420,"name":"Helios","domain":"heliosintel.ai","url":"https://alion.io/company/heliosintel","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"staff","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":185000,"max":285000,"currency":"USD","period":"year","gross":null,"usd_annual":285000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"ElasticSearch","optional":false},{"name":"GCP","optional":false},{"name":"OCR","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"PostgreSQL","optional":false},{"name":"RabbitMQ","optional":false},{"name":"Redis","optional":false},{"name":"Redpanda","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Typesense","optional":false}],"status":"live","first_seen_at":"2026-08-24T21:31:54Z","employer_posted_date":"2026-08-24","last_verified_at":"2026-09-28T22:54:38Z","board_verified":true,"closed_at":null,"days_open":35,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":35},"description":"Staff/Senior Distributed Systems Engineer\nNew York City | Full-time | On-site in SoHo, five days per week\nReports to the CTO and works directly with the founding team.\nABOUT HELIOS\nHelios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer.\nGovernment shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer.\nOur core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments.\nWe bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry.\nMISSION\nAs a distributed systems engineer at Helios you will expand and operate the distributed execution, storage, scheduling, and reliability primitives that support data ingestion, document processing, search indexing, model inference, agents, and continuously running research workflows.\nThe primary responsibility centers around the orchestration of critical infrastructure to meet mission-critical performance and reliability guarantees. Our platform runs a multitude of network crawling, compute grids, memory-heavy document processing, OCR, inference, and latency-sensitive jobs, all of which you should be familiar with. Preferred areas of expertise include:\nQueue-, log-, workflow-, and actor-based distributed execution systems, including Kafka, Redpanda, Pulsar, SQS, Pub/Sub, RabbitMQ, Temporal, and equivalent technologies.\n\nEstablishment of at-least-once delivery with effectively once-only business outcomes through transactional outbox and inbox patterns, sagas, reconciliation, dead-letter handling, replay and historical backfill procedures.\n\nCPU-, memory-, disk-, network-, and GPU-aware scheduling, including workload classification, priority allocation, starvation prevention, tenant and dependency concurrency limits, placement constraints, resource quotas, and noisy-neighbor isolation.\n\nPostgreSQL, object storage, Redis or equivalent caches, search indices (Elasticsearch or Typesense), vector stores, change-data-capture, and graph-storage systems.\n\nApplication of distributed coordination and concurrency controls, including leader election, locks, leases, fencing tokens, optimistic concurrency, conflict resolution, and partition recovery.\n\nProvisioning and operation of AWS, GCP, or Azure infrastructure through terraform or similar IaC, vulnerability scanning, CI/CD and deployment methods.\n\nDefinition and administration of SLO/Is, error budgets, release controls, metrics, OpenTelemetry, Datadog, fault injection and load/failure testing.\n\nEvaluation of system performance and cost via e2e latency decomposition, resource and dependency profiling.\n\nEnforcement of platform security and tenant isolation policies.\n\nThe key objective of this work is to support the real-time access and availability of our entire data corpus as well as supporting the continued construction of the Helios Rapid Ontology System (H.R.O.S.), our long horizon memory data plane.\nKEY RESPONSIBILITIES\nOwn the operation and architecture of Proxi’s distributed execution platform, covering ingestion, document processing, search index performance and long-running research jobs.\n\nManage the deployment of our GovCloud and Air-Gapped resources for sensitive environments.\n\nManage the orchestration of long-running agent research tasks including scheduling, lease management and retention.\n\nQueue optimization and cross-cloud information pipeline scalability.\n\nBuild out dedicated resource-aware autoscaling architecture for fast search and document processing resources.\n\nManage CI/CD and compliance operations across the entire Helios platform.\n\nSupport global forward embedded customer infrastructure efforts.\n\nWORKING AT HELIOS\nThis is a full-time, in-person role based in our SoHo office in New York City. Team members are expected to work from the office five days per week.\nCertain customer engagements may require background checks, security reviews, access approvals, or eligibility for a U.S. government security clearance. Some projects may be subject to U.S. citizenship or other customer-specific access requirements.\nWe're not looking for passengers; we want driven innovators with a hunger to build from the ground up - obsessed with pushing the boundaries of natural language understanding, document processing, and personalized relevance.\nHelios is a fast-moving startup with ambitious goals. This is not a conventional 9-to-5 role. We expect flexibility during critical deployment periods, customer incidents, product launches, and other company-critical work. In return, this role offers unusual ownership, direct access to consequential institutions, and the opportunity to build systems that affect how major decisions are made.\nHOW WE WORK\nOwn outcomes, not just assigned tasks. Identify what needs to happen and drive it through completion.\n\nMove quickly without lowering the standard. Speed and rigor are complementary.\n\nStay close to the mission and the user. The best decisions begin with the real problem.\n\nWork across boundaries. Everyone contributes beyond the narrow limits of a job title.\n\nCommunicate directly. We value clear thinking, honest feedback, and low-ego collaboration.\n\nBuild for the real world. Our systems must perform in complex, regulated, and high-stakes environments.\n\nEQUAL OPPORTUNITY\nHelios is an equal opportunity employer. We evaluate candidates based on their abilities, experience, and potential to contribute to our mission. We do not discriminate on the basis of race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other status protected by applicable law.","description_format":"text","description_chars":6308,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-12T03:59:20Z"}],"liveness":{"score":56,"band":"ok","label":"Likely open","p_open":1,"p_active":0.743,"p_room":0.75,"age_days":34,"expected_fill_days":49,"reasons":["conf:3","win:late"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":285000,"is_top_pay":true},"html_url":"https://alion.io/job/heliosintel-staffsenior-distributed-systems-engineer","json_url":"https://alion.io/job/heliosintel-staffsenior-distributed-systems-engineer.json","meta":{"generated_at":"2026-09-28T23:27:45Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1091,"day_limit":5000,"remaining_today":3909,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}