{"id":1373673,"url":"https://alion.io/job/crowdstrike-director-engineering-data-infrastructure-reliability","title":"Director, Engineering - Data Infrastructure & Reliability","company":{"id":2153,"name":"CrowdStrike","domain":"crowdstrike.com","url":"https://alion.io/company/crowdstrike","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":89,"open_postings":20,"ghost_share":0.05,"stale_share":0.3,"repost_share":0.05,"time_to_fill_p50_days":29,"computed_at":"2026-09-28T05:45:00Z"}},"role":"Leadership","role_family":"Leadership","seniority":"head","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States","Austin, United States","Sunnyvale, United States","Redmond, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":183000,"max_usd":344000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":1670},"experience_years_min":12,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"Crowdstrike","optional":false},{"name":"Falcon","optional":false},{"name":"Go","optional":false},{"name":"LLM Guardrails","optional":false},{"name":"SIEM","optional":false},{"name":"AI Agents","optional":true},{"name":"Ansible","optional":true},{"name":"Apache Kafka","optional":true},{"name":"Chaos Engineering","optional":true},{"name":"Flink","optional":true},{"name":"GitOps","optional":true},{"name":"Grafana","optional":true},{"name":"Kubernetes","optional":true},{"name":"LLM","optional":true},{"name":"OpenSearch","optional":true},{"name":"OpenTelemetry","optional":true},{"name":"Prometheus","optional":true},{"name":"Pulumi","optional":true},{"name":"Python","optional":true},{"name":"SLI/SLO/SLA","optional":true},{"name":"Spark","optional":true},{"name":"Terraform","optional":true},{"name":"Trino","optional":true}],"status":"live","first_seen_at":"2026-09-28T05:38:11Z","employer_posted_date":"2026-09-28","last_verified_at":"2026-09-29T03:37:10Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed - we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.\nAbout the Role:\nCrowdStrike's Data Platform is the foundation beneath Falcon and Next-Gen SIEM: the ingestion, streaming, storage, and query systems that take in almost 3 trillion events per day, at millions of events per second, and make them searchable in seconds. It spans multiple clouds and regions, including sovereign and regulated environments, and it holds petabytes of security telemetry that customers rely on in the middle of active investigations.\nWe set a high bar for that platform. Reliability, cost efficiency, and provable compliance are engineering requirements here rather than afterthoughts, and we are raising that bar again as we move the organization toward domain ownership, data as a product, self-serve tooling, and federated governance.\nWe support planetary scale and our systems are stateful and order-dependent, which puts them outside what generic cloud build automation is designed for. Validating a change means tracing real data from the sensor through ingestion to query rather than exercising a service in isolation. The largest cost levers are spread across every team that queries or stores data. Each new sovereign region has to prove residency and controls on its own evidence. Work like this is worth far more solved once, well, and for everyone than solved separately on six teams.\nWe are forming a centralized Data Platform Enablement team to own that category of work, and we are hiring its founding engineering leader.\nYou will start by hiring and forming a small, senior team, then grow it into a multi-team organization as the charter expands. The team works as a horizontal function outside individual product team scope, in three modes. It builds platform-wide automation, tooling, and test infrastructure. It partners hands-on with individual data platform teams. And it enables the rest of the organization through guardrails, best practices, and self-service tooling. Every hour your team spends on cross-cutting reliability, testing, and cost work is an hour another data platform team gets back for work in its own domain.\nThe reach of this role is wider than its headcount. You will not own the product domains, but you will answer for reliability, cost, and compliance outcomes across all of them, which means most of your influence has to come from somewhere other than your reporting line. This one suits someone who would rather build a function than inherit one.\nThis role is hybrid, linked to one of the posted locations: Austin, TX; Sunnyvale, CA; Redmond, WA; or NY, NY.\nWhat You'll Do:\nOrganizational Leadership & Team Building\nStand up the team: define its operating model, hire net-new engineers, and bring in senior operators from existing data platform teams without leaving holes behind them.\n\nHire, develop, and lead engineering managers and technical leads as the team grows from one group into several domain-aligned ones.\n\nSet the technical bar. Your team's output is tooling, harnesses, and guardrails that engineers outside your org have to actually want to adopt.\n\nOwn org design as the charter matures: forming, rechartering, and sequencing teams against where the platform's risk actually sits.\n\nCloud Expansion Automation\nBuild the automation that makes stateful data platform systems first-class citizens in our cloud build tooling, including the ordering, state awareness, and dependency sequencing that generic SRE automation cannot handle.\n\nTake regional and sovereign cloud build-outs to full automation, so standing up a new environment is a repeatable, hands-off sequence.\n\nOwn time-to-launch per cloud and region as a headline metric, alongside the teams that own the destination architecture.\n\nEnd-to-End Validation & Continuous Health Checks\nBuild a testing system that generates and traces data from the CrowdStrike sensor through ingestion to query validation. The same harness does pre-launch validation for every new cloud and region, then continuous health checks in steady state.\n\nRaise the fidelity of generated test data so it exercises the same paths production traffic takes, and a passing run means what it claims.\n\nHold end-to-end (i..e sensor to query) coverage and escaped-defect rate as first-class metrics for every launch.\n\nObservability & Proactive Incident Response\nStandardize reliability and efficiency metrics across every data platform team, so the whole platform reports from one instrumented view.\n\nBuild detection that catches problems before customers report them, with automated routing to the right responders plus first-line triage and guidance for owning teams.\n\nSet and drive down targets for mean time to detect, respond, and recover, backed by instrumentation that proves the trend.\n\nPush the incident rate between releases and operations down by addressing systemic causes.\n\nCost Engineering\nRun fast detection and remediation for cost anomalies, so an efficiency regression surfaces in days rather than quarters.\n\nFind and execute optimization work across the ecosystem, either directly with partner teams or by handing them tooling. Some of it is straightforward compute migration. Some of it is pipeline and architecture rework.\n\nMove cost accountability to the layers that own the levers, so that application teams see what their query patterns, storage choices, and pipeline hops actually cost. This is as much an organizational change as a technical one, and you will lead it.\n\nResilience, Scale & Capacity Planning\nEstablish a recurring game day practice that stresses systems deliberately to establish their real limits.\n\nRun scale and stress testing ahead of projected demand.\n\nWork with data platform teams on forward-looking capacity planning, and hold forecast accuracy as a real metric.\n\nData Residency & Compliance Automation\nTurn residency verification and audit evidence into reusable platform capabilities that every new region inherits. GRC defines the controls and other teams own the sovereign architecture; your team automates the data platform's side of proving both.\n\nBuild automated validation that data lands, is processed, and is queried inside its declared residency boundary, wired into the same harness as the rest of the testing rather than a parallel one.\n\nExpress the data platform's share of sovereign and certification controls as executable checks that run continuously, and generate attestation evidence from telemetry so certification and re-certification draw on evidence that is always current.\n\nTreat a control that has quietly stopped holding as an incident, detected and routed under the same targets as any reliability event.\n\nExecutive Narrative & Cross-Functional Influence\nTurn your team's delivery into a clear executive story on reliability, cost, and compliance posture, and carry it into senior forums with candor about risk.\n\nEarn adoption from engineering leaders who do not report to you. A horizontal team runs on that trust.\n\nNegotiate scope deliberately. This charter will attract more requests than any team can absorb, and protecting your team's focus is part of the job.\n\nWhat You'll Need:\n12+ years of software engineering experience, including 5+ years leading engineering teams and at least 2 years leading through other managers.\n\nA track record of building, growing, and retaining high-performing platform, SRE, or infrastructure teams in a fast-paced, high-growth environment, including hiring senior engineers onto a team that did not exist yet.\n\nHands-on grounding in SRE practice at scale: SLOs, SLIs, error budgets, incident command, blameless postmortems, and capacity planning for high-throughput distributed systems.\n\nExperience owning reliability for large-scale stateful distributed systems, where ordering, data state, and recovery semantics rule out generic automation. This is the hardest technical part of the job.\n\nWorking fluency with distributed data infrastructure: streaming platforms, OLAP and search engines, object storage, and large-scale query systems. Enough to hold a credible design conversation and to recognize a bad proposal.\n\nOwnership of a substantial infrastructure cost portfolio, including driving optimization across organizational boundaries and shifting cost accountability to the teams that control the spend.\n\nProven ability to influence without authority. You have changed how engineering teams outside your reporting line operate, and you can explain how you earned that adoption instead of mandating it.\n\nExperience operating under regulatory, residency, or certification constraints, and comfort turning control requirements into engineering work.\n\nStrong executive communication. You can compress a complicated reliability or cost story into something a leadership team can decide on, and you do not let your team's work go unseen.\n\nA specific point of view on applying AI to reliability and operations: where it pays off, where it does not, and what you would build first.\n\nProven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.\n\nBachelor's degree in Computer Science or related field, or equivalent work experience.\n\nBonus Points:\nExperience standing up sovereign, air-gapped, or regulated cloud environments, and automating the evidence that proves they comply.\n\nHands-on depth with Kafka, Flink, Spark, Cassandra, OpenSearch, Pinot, Trino, or comparable systems at petabyte scale.\n\nPython and/or Golang, Infrastructure as Code (Terraform, Ansible, Pulumi), Kubernetes at fleet scale, and GitOps workflows.\n\nAdvanced observability practice with Prometheus, Grafana, OpenTelemetry, distributed tracing, and large-scale log aggregation, weighted toward SLO dashboards and reliability scorecards rather than vanity metrics.\n\nHaving built and run a chaos engineering or game day practice that other teams joined voluntarily.\n\nFinOp, or having led a cost optimization and/or re-attribution program to completion.\n\nHaving shipped LLM-native or agentic tooling for incident prevention, triage, or remediation.\n\nThis role will require the candidate to periodically undergo and pass additional background and fingerprint check(s) consistent with government customer requirements.Benefits of Working at CrowdStrike:\nMarket leader in compensation and equity awards\n\nComprehensive physical and mental wellness programs\n\nCompetitive vacation and holidays for recharge\n\nPaid parental and adoption leaves\n\nProfessional development opportunities for all employees regardless of level or role\n\nEmployee Networks, geographic neighborhood groups, and volunteer opportunities to build connections\n\nVibrant office culture with world class amenities\n\nGreat Place to Work Certified™ across the globe...","description_format":"text","description_chars":14155,"description_truncated":true,"requirements":{"experience_years_min":12,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["401k plan","Equity","Health insurance","Professional development","Wellness"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Threat Intelligence","Endpoint Security","Security Operations"],"lifecycle":[{"event":"open","at":"2026-09-28T05:38:11Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":29,"reasons":["conf:0","win:early","comp:brand"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/crowdstrike-director-engineering-data-infrastructure-reliability","json_url":"https://alion.io/job/crowdstrike-director-engineering-data-infrastructure-reliability.json","meta":{"generated_at":"2026-09-29T04:46:05Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4565,"day_limit":5000,"remaining_today":435,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}