{"id":716943,"url":"https://alion.io/job/node-digital-senior-data-engineer","title":"Senior Data Engineer","company":{"id":5581,"name":"Node.Digital","domain":"node.digital","url":"https://alion.io/company/node-digital","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workable","truth_index":{"grade":"A","score":92,"open_postings":3,"ghost_share":0,"stale_share":0.333,"repost_share":0,"time_to_fill_p50_days":51,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Washington, United States"],"countries":["US"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":118000,"max_usd":203000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":212},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Machine Learning","optional":false},{"name":"MS SQL","optional":false},{"name":"Pandas","optional":false},{"name":"Polars","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"SQL","optional":false},{"name":"Bicep","optional":true},{"name":"CI/CD","optional":true},{"name":"Rest API","optional":true},{"name":"Spark","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-07-30T00:00:00Z","employer_posted_date":"2026-07-30","last_verified_at":"2026-09-30T23:08:12Z","board_verified":true,"closed_at":null,"days_open":63,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":63},"description":"Senior Data Engineer\nLocation: Herndon, VA (Remote Work)\nMust have an Public Trust Clearance\nKEY RESPONSIBILITIES\nProvide authoritative expertise on data engineering methods and best practices, including code first development approaches and modern pipeline design patterns.\nDesign, implement, and maintain the data architecture that supports products and end users, with all assets managed under source control.\nDesign, implement, and maintain ELT and ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning, using both SDK V1 and SDK V2.\nMigrate source data identified by SBA OIG into Azure Data Lake Storage.\nNormalize entity attributes such as addresses, phone numbers, and other common fields.\nReview, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.\nEstablish quality controls across all pipelines and introduce error handling, logging mechanisms, and validation checks.\nIncorporate source control across all pipelines and analytics codebases so code can evolve iteratively without destabilizing the architecture.\nOptimize ingestion, processing, and storage across a wide variety of datasets and data types, including modern columnar formats such as Parquet.\nDevelop self service capabilities that let SBA OIG analysts query and export data for investigations and audits.\nAuthor robust standard operating procedures governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.\nProduce detailed documentation of the data architecture, including data dictionaries, entity relationship diagrams, and pipeline process maps.\nMaintain and expand the environment with additional datasets and services on request, following a defined intake and testing process before production deployment.\nStay current with emerging AI tooling relevant to data engineering and contribute to exploratory work evaluating automation and language model assisted capabilities.\nRequirements\nEducation\nBachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field. Alternatively, five years of applied work experience in any of the same fields.\n5 years - Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.\n5 years - Designing, implementing, and maintaining ELT and ETL processes in cloud based data analytics environments.\n3 years - Working in Azure Synapse and Azure Machine Learning with the modern data stack. Certifications preferred, DP-203 or equivalent.\n3 years -Manipulating data in Python. Pandas is required. PySpark and Polars preferred. Experience developing reusable, modular code preferred.\nPREFERRED QUALIFICATIONS\nDP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification.\nImplementing pipelines and infrastructure using code first approaches: Python SDK, CLI, REST APIs, or infrastructure as code tooling such as Terraform or Bicep.\nImplementing source control and continuous integration and delivery workflows for data assets.\nDemonstrated familiarity with AI coding assistants and large language model integration patterns.\nPySpark or Polars at production scale.\nEntity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.\nBuilding self service analytic access for non engineering users.\nBenefits\nWe are proud to offer competitive compensation and benefits packages to include\nMedical \nDental\nVision\nBasic Life \nHealth Saving Account\n401K matching\nThree weeks of PTO/Sick\n11 Paid Holidays\nPre-Approved Online Training","description_format":"text","description_chars":3730,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":true,"languages":[]},"benefits":["401k plan"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Cybersecurity"],"lifecycle":[{"event":"open","at":"2026-09-11T06:14:15Z"}],"liveness":{"score":34,"band":"fade","label":"Fading","p_open":1,"p_active":0.778,"p_room":0.44,"age_days":63,"expected_fill_days":51,"reasons":["conf:6","velocity","win:tail","crowd:"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/node-digital-senior-data-engineer","json_url":"https://alion.io/job/node-digital-senior-data-engineer.json","meta":{"generated_at":"2026-10-01T11:10:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2632,"day_limit":5000,"remaining_today":2368,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}