{"id":1292982,"url":"https://alion.io/job/virtualitics-data-engineer","title":"Data Engineer","company":{"id":2043111,"name":"Virtualitics","domain":"virtualitics.com","url":"https://alion.io/company/virtualitics","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Columbia, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":91000,"max_usd":174000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":285},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon S3","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"Docker","optional":false},{"name":"Explainable AI","optional":false},{"name":"IAM","optional":false},{"name":"Kubernetes","optional":false},{"name":"Machine Learning","optional":false},{"name":"MySQL","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Amazon Redshift","optional":true},{"name":"Apache Kafka","optional":true},{"name":"Dagster","optional":true},{"name":"Databricks","optional":true},{"name":"dbt","optional":true},{"name":"NIST 800-171","optional":true},{"name":"NIST 800-53","optional":true}],"status":"live","first_seen_at":"2026-09-02T18:16:50Z","employer_posted_date":"2026-09-02","last_verified_at":"2026-09-28T18:37:39Z","board_verified":true,"closed_at":null,"days_open":26,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":26},"description":"Virtualitics is the category leader in AI-native readiness applications for defense, government, and critical infrastructure. Founded on a decade of Caltech research in partnership with NASA/JPL, we are led by scientists, strategists, and servicemembers united by a single mission: to solve the world’s most complex, mission-critical challenges with AI.\nOur Readiness AI solutions deliver operational certainty - giving leaders and operators a clear picture of what’s ready, what’s at risk, and what to do next. By identifying risks early, diagnosing root causes, and recommending prioritized actions with transparent, explainable AI, we help organizations move from data complexity to decision advantage.\nBehind that impact is relentless innovation. Inventors at heart, we hold 15+ U.S. patents and are leading the shift toward agent-driven readiness. But what truly sets us apart is our culture - relentless about results, grounded in transparency, and driven by compassion for the mission and the people it serves.\nIf you’re motivated by impact, inspired by technical depth, and ready to build AI that performs where it matters most - you’ll find your mission here.\nOur team is excited to find our next Data Engineer to join the company.\nAs a Data Engineer at Virtualitics, you will build and own the data foundation that our AI-native readiness applications run on - the pipelines, models, and infrastructure that turn messy, high-volume government and defense data into reliable, analysis-ready inputs for our platform. You will work shoulder-to-shoulder with Machine Learning engineers, Data Scientists, and Platform teams, adapting ingestion and transformation workflows to the specific, often unique data environments of each customer. This is a chance to work close to the mission and see your pipelines directly shape how analysts and decision-makers solve hard problems.\nThis is a Washington, D.C.-based role with on-site work at customer sites and secure facilities as required, plus remote flexibility.\n*Candidates must possess an active US Government Security Clearance (SECRET or higher).\nResponsibilities\nDesign, build, and maintain scalable, secure data pipelines for ingestion, transformation, and delivery across cloud and hybrid environments.\n\nModel and manage structured and unstructured data to support AI/ML workloads, analytics, and application features.\n\nPartner with Full Stack Engineers, Machine Learning Engineers, and Data Scientists to productionize data workflows, feature pipelines, and model inputs.\n\nBuild and improve automation for data quality, validation, lineage, and monitoring.\n\nOptimize storage, query performance, and cost across relational and object stores.\n\nImplement and enforce data security, access controls, and governance in line with government and DoD compliance requirements.\n\nSupport ingestion and integration of data within DoD and commercial customer environments, including deployments into secure and disconnected settings.\n\nInvestigate and resolve data, pipeline, and performance issues.\n\nCollaborate with Platform Engineers, AI Engineers, DevSecOps, and QA to ensure reliable, scalable, and secure data architectures.\n\nSkills & Qualifications\nActive US Government Security Clearance (SECRET or higher) - required.\n\nBS in Computer Science, Engineering, or related field.\n\n4+ years of experience in data engineering, or building production data pipelines.\n\nStrong programming skills in Python and SQL.\n\nExperience with distributed data processing frameworks (e.g., Spark).\n\nStrong experience with relational databases (PostgreSQL, MySQL) and data modeling.\n\nAdvanced AWS experience (S3, RDS, EMR, Lambda, IAM, CloudWatch); comfortable working within secure cloud environments.\n\nExperience building for data quality, validation, and observability.\n\nFamiliarity with containerized workloads (Docker, Kubernetes) and Infrastructure as Code (Terraform).\n\nNice to Have\nExperience deploying into DoD environments (Platform One, Palantir FedStart, etc.).\n\nKnowledge of NIST 800-53, NIST 800-171, and CMMC 2.0 controls.\n\nExperience supporting Machine Learning or Data Science teams with feature pipelines and ML data infrastructure.\n\nExperience with data pipeline and orchestration tools (e.g., Airflow, dbt, Dagster, or similar).\n\nExperience with streaming data (Kafka, etc.) and real-time processing.\n\nExperience with data warehousing and lakehouse architectures (i.e. AWS Athena, Redshift, Databricks, Stardog, etc.).\n\nStartup experience on a growing team, and a desire to mentor more junior engineers.","description_format":"text","description_chars":4555,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":true,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Data & Analytics","Military"],"lifecycle":[{"event":"open","at":"2026-09-26T08:32:57Z"}],"liveness":{"score":76,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.85,"p_room":0.9,"age_days":25,"expected_fill_days":42,"reasons":["conf:3","velocity","win:mid"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/virtualitics-data-engineer","json_url":"https://alion.io/job/virtualitics-data-engineer.json","meta":{"generated_at":"2026-09-28T22:55:15Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":552,"day_limit":5000,"remaining_today":4448,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}