{"id":1473203,"url":"https://alion.io/job/seneca-holdings-senior-data-engineer","title":"Senior Data Engineer","company":{"id":684887,"name":"Seneca Holdings","domain":"senecaholdings.com","url":"https://alion.io/company/senecaholdings","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"B","score":75,"open_postings":19,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":26,"computed_at":"2026-10-06T05:45:30Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"inferred_payroll_markers","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":["US"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":82000,"max_usd":180000,"period":"year","method":"global_role_seniority_cell","sample_n":1839},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"Databricks","optional":false},{"name":"GitLab","optional":false},{"name":"Pandas","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Amazon CloudWatch","optional":true},{"name":"Amazon Kinesis","optional":true},{"name":"Amazon S3","optional":true},{"name":"Apache Kafka","optional":true},{"name":"AWS Lambda","optional":true},{"name":"AWS Step Functions","optional":true},{"name":"CloudFormation","optional":true},{"name":"Delta Lake","optional":true},{"name":"FedRAMP","optional":true},{"name":"IAM","optional":true},{"name":"Scrum","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-29T17:13:06Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-07T01:09:25Z","board_verified":true,"closed_at":null,"days_open":7,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":7},"description":"Great Waters Federal i s part of the Seneca Nation Group (SNG) portfolio of companies . SNG is Seneca Holdings' federal government contracting business that meets mission-critical needs of federal civilian, defense, and intelligence community customers. Our portfolio comprises multiple subsidiaries that participate in the Small Business Administration 8(a) program. To learn more about SNG, visit the website and follow us on LinkedIn.\nOur team of talented individuals is what makes us successful. To support our team, we provide a balanced mix of benefits and programs. Your total rewards package includes competitive pay, benefits, and perks, flexible work-life balance, professional development opportunities, and performance and recognition programs. We offer a comprehensive benefits package that includes medical, dental, vision, life, and disability, voluntary benefit programs (critical illness, hospital, and accident), health savings and flexible spending accounts, and retirement 401K plan. One of our fundamental principles is to offer competitive health and welfare benefits to our team members, providing coverage and care for you and your family. Full-time employees working at least 30 hours a week on a regular basis are eligible to participate in our benefits and paid leave programs. We pride ourselves on our collaborative work environment and culture, which embraces our mission of providing financial and non-financial benefits back to the members of the Seneca Nation.\nWe are seeking an experienced Data Engineer to design, build, and maintain scalable data pipelines within a Databricks E2 environment hosted on AWS and governed by FISMA High and multi-tenant security and compliance standards.\nThe ideal candidate has strong hands-on experience with Databricks, Apache Spark, and the broader AWS data ecosystem. This individual should be comfortable working across the full data lifecycle-from ingestion and transformation to the delivery of analytics-ready datasets.\nKey Responsibilities\nAnalyze and collect data from various sources, including relational databases, APIs, external data providers, and real-time streaming sources.\nDesign and implement efficient, scalable data pipelines to cleanse, transform, and aggregate data for downstream reporting and analytics.\nBuild and maintain Databricks solutions using medallion architecture, including Bronze, Silver, and Gold layers, to create structured, reliable, and progressively refined datasets.\nDevelop source-to-target mapping documentation for all data pipelines and conduct thorough unit testing to validate data accuracy and pipeline logic.\nWrite advanced SQL to create aggregations across complex datasets and analyze data for anomalies, quality issues, and inconsistencies.\nImplement real-time and near-real-time data ingestion using AWS Database Migration Service (DMS) and other AWS-native services.\nDevelop data-processing logic using Python and/or R, leveraging Spark and related packages such as PySpark and Pandas for large-scale data transformation.\nManage source code, versioning, and deployments using GitLab, and support automated build and release processes through CI/CD pipelines.\nCollaborate with cross-functional teams in an Agile project environment, participating in sprint planning, stand-ups, and iterative delivery.\nEnsure that data engineering practices align with FISMA High and multi-tenant security, governance, and compliance requirements.\nLeverage AI automation tools to support data engineering pipeline development, testing, and validation and to accelerate delivery.\nRequired Qualifications\nProven, hands-on experience as a Data Engineer or Databricks Developer building production-grade data pipelines.\nStrong working knowledge of Databricks on AWS, including E2 architecture, cluster configuration, job orchestration, and workspace management.\nHands-on experience with Apache Spark, PySpark, and Pandas for large-scale, distributed data processing.\nSolid programming skills in Python; working knowledge of R is a plus.\nExperience with Databricks Auto Loader for scalable, incremental file ingestion.\nPractical experience with AWS Database Migration Service (DMS) for change data capture and real-time data replication.\nStrong experience with Delta Lake and Delta tables, including schema evolution, time travel, and optimization techniques such as OPTIMIZE, Z-ORDER, and VACUUM.\nExperience with Amazon RDS and other relational database sources used in data extraction and integration workflows.\nAdvanced SQL skills, including complex joins, window functions, aggregations, and performance tuning.\nSolid understanding of medallion architecture and modern data lakehouse design principles.\nExperience with GitLab and CI/CD pipelines for automated testing, build, and deployment of data engineering code.\nExperience working in Agile/Scrum project environments.\nPreferred Skills\nBroader knowledge of AWS services beyond DMS and RDS, including S3, Glue, Lambda, Step Functions, CloudWatch, and IAM.\nExperience operating within FISMA High, FedRAMP, or other regulated, multi-tenant compliance environments.\nFamiliarity with data quality frameworks, anomaly detection techniques, and automated data validation.\nExposure to infrastructure-as-code tools such as Terraform or CloudFormation for provisioning data platform resources.\nKnowledge of streaming technologies such as Kafka, Kinesis, or Spark Structured Streaming.\n\nEqual Opportunity Statement:\nSeneca Holdings provides equal employment opportunities to all employees and applicants without regard to race, color, religion, sex/gender, sexual orientation, national origin, age, disability, marital status, genetic information and/or predisposing genetic characteristics, victim of domestic violence status, veteran status, or other protected class status. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leave of absence, compensation and training. The Company also prohibits retaliation against any employee who exercises his or her rights under applicable anti-discrimination laws. Notwithstanding the foregoing, the Company does give hiring preference to Seneca or Native individuals. Veterans with expertise in these areas are highly encouraged to apply.","description_format":"text","description_chars":6342,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"phd","optional":false},"security_clearance":false,"languages":[]},"benefits":["401k plan","Flexible schedule","Professional development","Retirement plans"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-29T17:39:31Z"}],"visa":[],"liveness":{"score":61,"band":"ok","label":"Likely open","p_open":1,"p_active":0.615,"p_room":1,"age_days":6,"expected_fill_days":26,"reasons":["conf:1","stale_co","velocity","win:early"],"computed_at":"2026-10-06T05:45:30Z"},"pay":null,"html_url":"https://alion.io/job/seneca-holdings-senior-data-engineer","json_url":"https://alion.io/job/seneca-holdings-senior-data-engineer.json","meta":{"generated_at":"2026-10-07T01:53:43Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3060,"day_limit":5000,"remaining_today":1940,"minute_limit":60,"resets_at":"2026-10-08T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":684887},"rest":"https://alion.io/mcp/rest/get_company?id=684887"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fseneca-holdings-senior-data-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fseneca-holdings-senior-data-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fseneca-holdings-senior-data-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/seneca-holdings-senior-data-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fseneca-holdings-senior-data-engineer"}]}