{"id":1485269,"url":"https://alion.io/job/coreweave-hardware-engineer-server-infrastructure","title":"Senior Hardware Engineer, Server Infrastructure","company":{"id":7348,"name":"CoreWeave","domain":"coreweave.com","url":"https://alion.io/company/coreweave","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"B","score":82,"open_postings":72,"ghost_share":0,"stale_share":0.514,"repost_share":0,"time_to_fill_p50_days":81,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Hardware","role_family":"Hardware","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"Ansible","optional":false},{"name":"CoreWeave","optional":false},{"name":"Python","optional":false},{"name":"Grafana","optional":true},{"name":"Kubernetes","optional":true},{"name":"Linux","optional":true},{"name":"Prometheus","optional":true}],"status":"live","first_seen_at":"2026-09-29T22:15:19Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-01T06:54:47Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.\nAbout the Role\nCoreWeave is seeking a highly skilled and motivated engineer to join our Hardware Engineering team. In this role, you will help design, develop, and optimize our server hardware infrastructure. You’ll collaborate closely with cross-functional teams, external vendors, and key stakeholders to deliver performant, reliable, and scalable hardware solutions that power CoreWeave’s rapidly growing infrastructure.\nYou will own server hardware from provisioning through decommission. This hands-on role combines engineering and operational support. Engineering work includes automation across the hardware lifecycle, hardware and firmware management services, monitoring and alerting, and qualification and bring-up of new platforms. Operational support includes acting as a senior point of contact for hardware escalations, driving deep root-cause analysis across hardware and firmware, and working quality and RMA issues through to resolution with server vendors and OEMs.\nBoth engineering and operational support are core responsibilities. You will support the systems you build and use what you learn from production failures to improve automation, telemetry, and platform design. You will work closely with data center operations teams, hardware technicians, and engineering teams to bring new regions online and keep existing infrastructure healthy.\nWhat You’ll Do\nDesign and develop server hardware infrastructure to support CoreWeave’s high-performance workloads.\nAutomate all aspects of the server hardware lifecycle, from provisioning and configuration through firmware management, monitoring, and decommissioning.\nDevelop and maintain hardware and firmware management services that ensure reliability at scale.\nDevelop and implement monitoring and alerting for server hardware health, improving alert quality to support reliable on-call response.\nServe as a senior point of contact for hardware escalations, performing deep troubleshooting and root-cause analysis across hardware and firmware to drive long-term fixes.\nParticipate in an on-call rotation for hardware escalations and improve the runbooks, alerts, and tooling that make the rotation sustainable.\nCollaborate with cross-functional teams to define hardware requirements, specifications, and system architecture.\nWork with server vendors and OEMs to evaluate, qualify, and deploy new platforms and resolve firmware, quality, and RMA issues.\nSupport new data center region bring-up and hardware qualification.\nAnalyze hardware system performance, identify bottlenecks, and implement improvements to efficiency and resilience.\nEstablish and continuously refine processes for internal hardware testing, deployment, and performance optimization.\nCreate and maintain accurate documentation of hardware designs, specifications, test procedures, and results.\nSupport data center operations teams and hardware technicians with troubleshooting guidance, runbooks, and training so common issues can be resolved without engineering escalation.\nTurn recurring production failures into automation, better telemetry, and improvements to platform design and vendor solutions.\nCommunicate status, trade-offs, and risks clearly to engineering, operations, and customer-facing stakeholders, including during active incidents.\nWho You Are\nDeep understanding of server hardware, components, and management technologies.\nProficiency in Ansible or Python, with hands-on experience programmatically interacting with server BMCs using Redfish or IPMI; Redfish preferred.\nExperience collaborating with hardware vendors and OEMs to evaluate, qualify, and deploy server solutions.\nDemonstrated experience supporting and troubleshooting production infrastructure, participating in on-call or escalation rotations, and driving incidents through to root-cause resolution.\nComfort with both building and automating systems and supporting infrastructure already in production.\nProven ability to stay current with technologies and trends in server and data center hardware.\nStrong interest in automation and infrastructure scalability, with a commitment to continuous improvement.\nExcellent technical documentation skills and attention to detail.\nStrong analytical and problem-solving abilities, with a bias toward systematic, data-driven decisions.\nExcellent written and verbal communication skills in English, with the ability to work effectively with technical teams and cross-functional stakeholders.\nPreferred Qualifications\nExperience bringing up new data center regions or standing up infrastructure in a new geography.\nExperience with GPU platforms and rack-scale systems such as NVIDIA GB200 or GB300.\nExperience with BMC technologies and management interfaces such as Redfish or IPMI at fleet scale.\nExperience with firmware lifecycle management or hardware qualification programs in large-scale environments.\nStrong Linux systems administration and debugging skills at fleet scale.\nFamiliarity with Kubernetes-based services, observability tools such as Prometheus and Grafana, and distributed production environments.\nExperience designing alerting, runbooks, and self-service tooling that enable operations teams to resolve common hardware failures without engineering escalation.\nExperience supporting external customers or partners in a technical troubleshooting capacity.\nWondering if You’re a Good Fit?\nWe believe in investing in our people and value candidates who bring diverse experiences to our teams-even if you aren’t a 100% skill or experience match. If some of this describes you, we’d love to talk.\nYou enjoy solving problems at the boundary of hardware, firmware, and software.\nYou want to both build systems and support them, and you use operational experience to improve what you build.\nYou enjoy tracing intermittent hardware failures to their root cause and validating that a fix addresses the underlying problem.\nYou look for opportunities to automate repetitive hardware procedures.\nYou work effectively with operations, engineering, and vendor teams to resolve complex technical issues.\nYou bring structure, accountability, and momentum to fast-moving environments.\nWhat We Offer\nThe range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.\nIn addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:\nMedical, dental, and vision insurance - 100% paid for by CoreWeave\nCompany-paid Life Insurance \nVoluntary supplemental life insurance \nShort and long-term disability insurance \nFlexible Spending Account\nHealth Savings Account\nTuition Reimbursement \nAbility to Participate in Employee Stock Purchase Program (ESPP)\nMental Wellness Benefits through Spring Health \nFamily-Forming support provided by Carrot\nPaid Parental Leave \nFlexible, full-service childcare support with Kinside\n401(k) with a generous employer match\nFlexible PTO\nCatered lunch each day in our office and data center locations\nA casual work environment\nA work culture focused on innovative disruption\nCalifornia Applicants\nCalifornia Consumer Privacy Act \nEqual Opportunity & Accommodations\nCoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.\nAs part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: .\nExport Control Compliance\nThis position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.","description_format":"text","description_chars":9418,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":["Life insurance","Parental leave","Vision insurance"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Data Centers & Colocation","Cloud Platforms (IaaS & PaaS)","AI Compute & Inference"],"lifecycle":[{"event":"open","at":"2026-09-29T22:17:49Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":1,"expected_fill_days":81,"reasons":["conf:12","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/coreweave-hardware-engineer-server-infrastructure","json_url":"https://alion.io/job/coreweave-hardware-engineer-server-infrastructure.json","meta":{"generated_at":"2026-10-01T11:27:56Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"assistant","counted_by":"address","units_charged":1,"used_today":604,"day_limit":2000,"remaining_today":1396,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}