{"id":1674112,"url":"https://alion.io/job/ddn-staff-engineer-lustre","title":"Staff Engineer, Lustre","company":{"id":672909,"name":"DDN","domain":"ddn.com","url":"https://alion.io/company/ddn","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"A","score":100,"open_postings":3,"ghost_share":0,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":39,"computed_at":"2026-10-04T05:45:00Z"}},"role":"Industrial Engineering","role_family":"Industrial Engineering","seniority":"staff","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Santa Clara, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":200000,"max":250000,"currency":"USD","period":"year","gross":null,"usd_annual":250000},"salary_estimate":null,"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"InfiniBand","optional":false},{"name":"Linux","optional":false},{"name":"HPC","optional":true}],"status":"live","first_seen_at":"2026-10-01T19:47:16Z","employer_posted_date":"2026-10-01","last_verified_at":"2026-10-04T22:46:02Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"We are seeking a Staff Engineer with 10+ years of experience in distributed storage and Linux-based systems engineering. This is a hands-on senior technical role focused on design, debugging, performance, and operational excellence across LustreFS and adjacent stack components. The ideal candidate brings strong expertise in one or more Lustre subsystems, can independently drive complex investigations, and collaborates effectively across engineering, QE, support and release teams. Engineers who are comfortable using AI to accelerate triage, debugging, code comprehension and new feature design will be especially valuable.\nKey Responsibilities\nDesign, develop and debug LustreFS features, fixes and enhancements across relevant subsystems such as llite, MDS/MDT, OSS/OST, LDLM and LNet.\n\nInvestigate customer and scale-related defects, drive root-cause analysis and implement high-quality fixes with strong attention to correctness and maintainability.\n\nContribute to performance tuning, failure analysis and reliability improvements for large-scale Lustre deployments.\n\nParticipate actively in code reviews, design reviews and subsystem discussions, bringing rigor to testing and operational readiness.\n\nWork closely with QE and support to reproduce issues, improve diagnostic data quality and increase coverage for high-risk failure scenarios.\n\nHelp document subsystem behavior, debugging approaches, known failure patterns and operational best practices.\n\nUse AI-assisted tools where appropriate to speed up issue triage, summarize logs, improve code understanding and capture reusable lessons learned.\n\nRequired Qualifications\n10+ years of experience in systems software, distributed systems, storage, Linux kernel or filesystem engineering.\n\nStrong experience in LustreFS development, support or performance engineering with depth in at least one major subsystem.\n\nStrong C programming and Linux systems debugging skills.\n\nWorking knowledge of Linux kernel internals, filesystem semantics, networking and performance analysis.\n\nExperience with LNet and/or high-performance transports such as RDMA, InfiniBand, RoCE or TCP-based storage networking.\n\nAbility to debug and resolve issues spanning multiple layers including client, server, network and backend storage.\n\nStrong collaboration skills and the ability to work across functions in a fast-moving engineering environment.\n\nPreferred Skills\nExperience in HPC, AI infrastructure or large-scale parallel storage environments.\n\nExposure to metadata-heavy and throughput-heavy workload characterization and tuning.\n\nFamiliarity with ZFS, ldiskfs, NVMe-backed storage and related observability / performance tooling.\n\nExperience creating test plans, reproducer frameworks, runbooks or diagnostic automation.\n\nComfort using AI tools to accelerate debugging, code reviews, triage, documentation and early-stage design ideation.\n\nExperience mentoring junior engineers or leading focused technical efforts within a subsystem.\n\nWhat You Will Work On\nHands-on development and debugging of LustreFS defects, performance issues and subsystem enhancements.\n\nCustomer-facing and scale-related issue investigation across llite, metadata, object storage, LNet and transport layers.\n\nCollaborative design and implementation of reliability, observability and serviceability improvements.\n\nReviewing and validating fixes through targeted tests, failure injection, log analysis and performance characterization.\n\nUsing AI-assisted workflows to accelerate triage, debug loops, code understanding and documentation quality.\n\nContributing to team redundancy by strengthening documentation, code review quality and subsystem knowledge sharing.\n\nWhy This Role Matters\nThis role is central to building durable engineering redundancy in LustreFS: expanding deep subsystem ownership, reducing concentration risk, and accelerating next-generation delivery through strong engineering fundamentals and AI-enabled execution.","description_format":"text","description_chars":3955,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Enterprise Storage Systems","Servers & Data Center Hardware"],"lifecycle":[{"event":"open","at":"2026-10-02T07:42:22Z"}],"visa":[],"liveness":{"score":99,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.993,"p_room":1,"age_days":2,"expected_fill_days":39,"reasons":["conf:0","urgency","velocity","win:early"],"computed_at":"2026-10-04T05:45:00Z"},"pay":{"stated_usd_annual":250000,"is_top_pay":true},"html_url":"https://alion.io/job/ddn-staff-engineer-lustre","json_url":"https://alion.io/job/ddn-staff-engineer-lustre.json","meta":{"generated_at":"2026-10-05T00:58:31Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1259,"day_limit":5000,"remaining_today":3741,"minute_limit":60,"resets_at":"2026-10-06T00:00:00Z"}}}