{"id":1229656,"url":"https://alion.io/job/coreweave-principal-engineer-distributed-systems","title":"Principal Engineer, Distributed Systems","company":{"id":7348,"name":"CoreWeave","domain":"coreweave.com","url":"https://alion.io/company/coreweave","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"B","score":70,"open_postings":16,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":154,"computed_at":"2026-09-28T05:45:00Z"}},"role":"Backend","role_family":"Backend","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":227000,"max":303000,"currency":"USD","period":"year","gross":null,"usd_annual":303000},"salary_estimate":null,"experience_years_min":12,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"C++","optional":false},{"name":"CoreWeave","optional":false},{"name":"Go","optional":false},{"name":"Java","optional":false},{"name":"Python","optional":false},{"name":"Rust","optional":false},{"name":"Threat Modeling","optional":false},{"name":"Kubernetes","optional":true}],"status":"live","first_seen_at":"2026-09-14T19:42:10Z","employer_posted_date":"2026-09-17","last_verified_at":"2026-09-28T19:58:13Z","board_verified":true,"closed_at":null,"days_open":14,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":14},"description":"CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.\nAbout the role\nWe are looking for a Principal Engineer to provide technical leadership across Security Products. This is a senior individual-contributor role for an engineer who can define architecture, guide execution across multiple teams, and solve complex distributed-systems problems in security-critical infrastructure.\nThe systems you help design will support demanding production requirements: four nines of availability, high scalability, consistently low latency, strong security boundaries, and safe behavior under partial failure in a multi-region setup. You will work across the full lifecycle of these systems, from architecture and technical strategy through implementation guidance, operational readiness, incident learning, and long-term evolution.\nYour work will span across high-scale authorization systems, anomaly detection systems, bot-defense and security token service (STS), API authentication gateway, and the shared infrastructure required to operate these capabilities reliably across regions and deployment environments.\nYou will partner with engineering, security, infrastructure, networking, platform, product, and customer-facing teams. Success in this role requires both deep technical judgment and the ability to create alignment, raise engineering standards, and make complex architecture understandable and actionable for others.\nWhat you will do\nEstablish architectural approaches for distributed systems - multi-region operation, including service placement, failover, replication, traffic management, disaster recovery, and regional independence.\nDrive decisions around consistency models, caching, invalidation, propagation, revocation, idempotency, concurrency, and ordering where correctness and security are critical. \nDesign for data residency, tenant isolation, trust boundaries, blast-radius reduction, and controlled handling of sensitive security data.\nDefine fault-tolerance strategies for dependencies, networks, regions, storage systems, and control-plane components, including graceful degradation and safe recovery.\nDesign systems that meet four-nines availability goals while maintaining predictable low latency and high throughput under normal operation, traffic spikes, and partial failures.\nImprove the reliability and operability of critical services through SLOs, error budgets, metrics, logs, traces, audit events, alerting, incident response, and post-incident learning.\nGuide teams through architecture reviews, design reviews, implementation tradeoffs, capacity planning, load testing, performance analysis, and production readiness assessments.\nMentor senior and staff engineers, develop technical talent, and raise the quality of engineering practice across the organization.\nCommunicate architecture, tradeoffs, risks, and recommendations clearly to technical and executive audiences.\nLead the design of high-scale authorization systems that support policy authoring, policy evaluation, access-control enforcement, auditability, and integration across many services and tenants.\nProvide technical direction for a Security Token Service, including token issuance, validation, lifecycle management, trust relationships, key rotation, revocation, and secure service-to-service access.\nGuide the design and implementation of an API Authentication Gateway that provides consistent, secure, observable authentication and authorization for CoreWeave APIs.\nSet technical standards for API design, service contracts, threat modeling, cryptographic key management, secrets handling, observability, and operational readiness.\nPartner with service teams to make authorization and authentication capabilities easy to adopt through clear interfaces, SDKs, reference implementations, documentation, and reliable integration patterns.\nWhat you bring\n12+ years of experience designing and building production software, including substantial experience with distributed systems and platform infrastructure.\nA track record of serving as a principal, staff-plus, distinguished, or equivalent technical leader across multiple teams or a broad technical domain.\nDeep expertise in distributed-systems design, including scalability, availability, latency, consistency, concurrency, partition tolerance, caching, replication, and failure recovery.\nExperience designing and operating security-sensitive or reliability-critical services in production.\nStrong software engineering experience with one or more systems languages or backend ecosystems, such as Go, Java, Rust, C++, Python, or equivalent.\nDemonstrated ability to move from ambiguous requirements to clear architecture, sequenced execution plans, and durable engineering outcomes.\nExperience making architecture decisions that balance security, correctness, performance, operational simplicity, and time to value.\nStrong written and verbal communication skills, including the ability to influence without direct authority and build alignment across organizational boundaries.\nA track record of improving engineering standards, mentoring senior engineers, and helping teams deliver complex systems safely.\nPreferred qualifications\nExperience designing and implementing authorization systems at high scale, including policy engines, permission models, policy distribution, decision services, enforcement points, or authorization observability.\nExperience designing or operating a Security Token Service, identity platform, credential service, or comparable security-critical control plane.\nExperience designing or operating an API Authentication Gateway, service-mesh authorization layer, or centralized authentication and authorization platform.\nExperience operating systems with four-nines availability requirements, high request volume, strict latency objectives, and demanding reliability expectations.\nExperience with multi-region architecture, regional failover, active-active or active-passive operation, replication, disaster recovery, and regional isolation.\nExperience reasoning about strong, eventual, and bounded-staleness consistency models, especially for security policy, credentials, tokens, permissions, and revocation.\nExperience designing for data residency, regional data controls, tenant isolation, trust-domain separation, and compliance-sensitive deployment models.\nExperience building systems that remain secure and useful during dependency failures, network partitions, regional outages, stale data, degraded capacity, or partial compromise.\nExperience with authentication and identity standards such as OAuth 2.0, OpenID Connect, SAML, SCIM, JWT, mTLS, SPIFFE/SPIRE, or related protocols.\nExperience with cryptography, key management, HSMs, secrets management, certificate authorities, signing keys, token validation, or secure credential lifecycle management.\nExperience with Kubernetes, cloud infrastructure, service networking, distributed storage, messaging systems, and multi-cluster operations.\nExperience building security and platform services with strong observability, including metrics, logs, traces, audit records, security events, and actionable alerting.\nExperience working with enterprise customers, regulated workloads, or hybrid and multi-cloud environments.\nHow you will be successful\nIn your first year, you will:\nEstablish a clear technical north star for Security Products’ authorization, authentication, and security-infrastructure capabilities.\nBuild strong working relationships with engineering and product leaders, staff-plus engineers, security teams, and the service teams that depend on these platforms.\nAdvance the architecture and delivery plan for high-scale authorization systems, the Security Token Service, and the API Authentication Gateway.\nImprove the reliability, performance, observability, and operational maturity of critical security services against four-nines availability goals.\nHelp teams make explicit, durable decisions about multi-region design, data consistency, data residency, tenant isolation, fault tolerance, and security boundaries.\nReduce duplicated access-control and authentication patterns by enabling consistent, well-documented platform capabilities.\nRaise the quality of architecture reviews, threat modeling, performance testing, capacity planning, and production readiness across Security Products.\nMentor senior engineers and help grow a strong technical leadership community.\nWhy join Security Products\nSecurity Products is building the systems that make CoreWeave trustworthy at scale. You will work on foundational infrastructure that protects access to cloud resources, enables secure product experiences, and supports customers running critical AI workloads.\nThis is an opportunity to shape the architecture of a modern cloud security platform while solving problems at the intersection of distributed systems, security engineering, reliability, performance, and global-scale infrastructure.\nLocation requirement\nThis position is based in Sunnyvale, California or New York, New York and requires working on-site. Remote work is not available for this role.\nCoreWeave is an equal opportunity employer. We evaluate qualified applicants without regard to legally protected characteristics.\nThe base salary range for this role is $227,000 to $303,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). \nWhat We Offer\nThe range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.\nIn addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:\nMedical, dental, and vision insurance - 100% paid for by CoreWeave\nCompany-paid Life Insurance \nVoluntary supplemental life insurance \nShort and long-term disability insurance \nFlexible Spending Account\nHealth Savings Account\nTuition Reimbursement \nAbility to Participate in Employee Stock Purchase Program (ESPP)\nMental Wellness Benefits through Spring Health \nFamily-Forming support provided by Carrot\nPaid Parental Leave \nFlexible, full-service childcare support with Kinside\n401(k) with a generous employer match\nFlexible PTO\nCatered lunch each day in our office and data center locations\nA casual work environment\nA work culture focused on innovative disruption\nCalifornia Applicants\nCalifornia Consumer Privacy Act \nEqual Opportunity & Accommodations\nCoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.\nAs part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable acc...","description_format":"text","description_chars":12880,"description_truncated":true,"requirements":{"experience_years_min":12,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Life insurance","Parental leave","Vision insurance"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Data Centers & Colocation","Cloud Platforms (IaaS & PaaS)","AI Compute & Inference"],"lifecycle":[{"event":"open","at":"2026-09-25T14:29:23Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":13,"expected_fill_days":154,"reasons":["conf:4","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":303000,"is_top_pay":true},"html_url":"https://alion.io/job/coreweave-principal-engineer-distributed-systems","json_url":"https://alion.io/job/coreweave-principal-engineer-distributed-systems.json","meta":{"generated_at":"2026-09-29T03:03:10Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2716,"day_limit":5000,"remaining_today":2284,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}