{"id":1331653,"url":"https://alion.io/job/harell-data-software-engineer-ai-infrastructure","title":"Software Engineer, AI Infrastructure","company":{"id":3786358,"name":"Harell Data","domain":"harelldata.com","url":"https://alion.io/company/harelldata","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Rippling","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Palo Alto, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":158000,"max_usd":291000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":2304},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Fine-tuning","optional":false},{"name":"GCP","optional":false},{"name":"Kubernetes","optional":false}],"status":"live","first_seen_at":"2026-08-04T19:03:05Z","employer_posted_date":"2026-08-04","last_verified_at":"2026-09-30T22:27:00Z","board_verified":true,"closed_at":null,"days_open":57,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":57},"description":"About Harell Data\nAI has transformed digital industries, but progress in the physical sciences - drug discovery, materials science, climate modeling - has stalled. The bottleneck isn't compute or algorithms. It's data. The most valuable scientific datasets are locked in silos, unstructured, and inaccessible.\nWe're fixing that. Harell Data is a managed platform where organizations can securely share proprietary datasets, train models on high-performance GPUs, and deploy them for inference and application development. Dataset owners share their data for model training without giving up control of it. Researchers and engineers get access to compute to train, deploy, and use their models. It's the infrastructure layer that turns scattered scientific data into domain-specific foundation models.\n About the Role\nYou'll be an early engineer reporting directly to the CTO. You'll own the compute layer: the GPU clusters and the inference systems that run on them. You'll make the architectural decisions that define the platform. You'll also work directly with customers to understand what they actually need and turn that into infrastructure that works at scale.\nWhat You Will Do\nBuild the GPU compute layer - Orchestration for GPU workloads on Kubernetes: resource allocation, scheduling, multi-tenancy, and cost management.\nBuild the inference layer - Model loading, autoscaling, batching, and serving. You own the latency and throughput customers feel.\nOwn the ML pipeline end to end - Data ingestion, preprocessing, training and fine-tuning jobs, and recovery when multi-node jobs fail.\nWork directly with customers - Debug fine-tuning jobs that fail or run slow. Build the observability that tracks model performance and resource health in real time.\nOwn reliability - Incident response, on-call, and keeping the platform up as usage grows.\nShape technical direction - Lead build-vs-buy decisions on infrastructure and security. Set engineering standards. Help hire the team you want to work with.\nQualifications\n5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines\nHands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads.\nStrong CS fundamentals and system design chops\nComfortable with ambiguity - you've worked somewhere where the playbook didn't exist yet\nLocation note: this role is based in Palo Alto, CA. No relocation assistance available for this role.","description_format":"text","description_chars":2466,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Relocation assistance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":true,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-27T10:22:26Z"}],"liveness":{"score":18,"band":"cold","label":"Long shot","p_open":1,"p_active":0.503,"p_room":0.35,"age_days":57,"expected_fill_days":24,"reasons":["conf:7","win:tail"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/harell-data-software-engineer-ai-infrastructure","json_url":"https://alion.io/job/harell-data-software-engineer-ai-infrastructure.json","meta":{"generated_at":"2026-10-01T10:17:56Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1368,"day_limit":5000,"remaining_today":3632,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}