{"id":1556915,"url":"https://alion.io/job/gather-ai-software-development-engineer-ii-data-engineer","title":"Software Development Engineer II – Data Engineer","company":{"id":31685,"name":"Gather AI","domain":"gather.ai","url":"https://alion.io/company/gather","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":95,"open_postings":5,"ghost_share":0,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":77,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":["IN"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":15500,"max_usd":38000,"period":"year","method":"role_seniority_country_cell","sample_n":17},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Dagster","optional":false},{"name":"Databricks","optional":false},{"name":"dbt","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"Docker","optional":false},{"name":"Git","optional":false},{"name":"Kubernetes","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Snowflake","optional":false},{"name":"SQL","optional":false},{"name":"Apache Kafka","optional":true},{"name":"Fivetran","optional":true},{"name":"pySpark","optional":true},{"name":"Spark","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-30T20:18:32Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-02T00:04:27Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About Us\nAre you ready to build the future of the supply chain? At Gather AI, we're not just creating software; we're pioneering a new era of warehouse intelligence. We've developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining \"on-time, in full\" delivery.\nIf you're looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We're leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time.\nAbout the Team\nYou'll join the Full Stack group within Cloud Services, which is being built now. Today, production and analytical workloads share a single database, and each product defines its own metrics. This team exists to fix that by designing and building the warehouse, transformation layers, and semantic model from the ground up. You'll work closely with Full Stack, ML, Product, Security, and Customer Success.\nAbout the Role\nMost data engineers early in their careers maintain one slice of a large, mature platform, or build pipelines against a model someone else designed years ago. This role is the opposite.\nAs an SDE II, Data Engineer, you'll be one of the first engineers building Gather AI's data foundation from scratch. The work includes:\nMoving analytics off our production PostgreSQL database\nWriting the layered dbt models every product will share\nDefining metrics once so they stay consistent everywhere\nLinking structured records back to the drone images and video they came from\nYou'll start by shipping well-scoped pipelines and models on our Drone product alongside senior engineers. From there, you'll grow toward owning a data domain end to end: its design, its quality, and its on-call.\nThis is a fully remote role on our India-based team. We work across multiple time zones, so clear written communication, initiative, and the ability to make progress independently matter.\nWhat You'll Do\n Build extraction pipelines. Move data from production PostgreSQL into the analytical warehouse (raw, refined, serving) using incremental loads that protect the production database.\nBuild and extend the shared data model. Write and maintain dbt models, such as dimensions, reusable metric building blocks, and serving tables, following the architecture set by our Principal Data Engineer.\nImplement metrics. Add metrics to the semantic layer so a measure like \"scan accuracy by site\" is defined once and matches across every dashboard.\nKeep data trustworthy. Add tests, freshness checks, and alerts to the pipelines you build, and help run safe backfills when data needs correcting.\nBuild in tenant isolation. Apply our access-control and tenant-separation patterns to every model and pipeline you ship.\nMaintain lineage. Keep metrics traceable to their source records and drone images, so any number can be checked against the original capture.\nWork at the ingestion boundary. Partner with the integration team to validate incoming WMS data against agreed contracts.\nDocument your work. Keep models documented and registered in the data catalog.\nShip safely. Deliver through CI/CD, use AI-assisted development, and join on-call for the pipelines you own.\nGrow with the platform. Start on the Drone product, then help onboard MHE Vision, 3D case counting products onto the same data foundation over the year.\nWhat You'll Need\nExperience: 2-5 years building and running production data pipelines, and a degree in Computer Science or equivalent practical experience.\n SQL: Strong SQL, including joins, window functions, and CTEs, plus a working understanding of dimensional modelling (facts, dimensions, grain). PostgreSQL experience preferred.\nTransformation: Hands-on experience building tested models in dbt, Snowflake Dynamic Tables, Databricks Lakeflow Declarative Pipelines, or equivalent.\nPython: Production pipeline code with tests and code review, beyond notebooks or one-off scripts.\nOrchestration: Experience running pipelines in Airflow, Dagster, Databricks Lakeflow Jobs, Snowflake Tasks, or equivalent, including handling failures, re-runs, and backfills.\nData movement: Has loaded data from an operational database into a warehouse or lakehouse such as Snowflake or Databricks.\nCloud: Production experience on Azure or another major cloud, including object storage, Git, CI/CD, and Docker/Kubernetes.\nData quality mindset: Writes tests alongside code and cares whether data is correct, not just whether the job ran.\nOwnership and growth: Hands-on and curious. Digs into problems beyond their own code, brings ideas to the team, and takes review feedback well.\nCommunication: Clear written and spoken English. Writes good PRs and documentation and raises blockers early in a distributed team.\nNice to Have\nChange data capture (Debezium, Fivetran) or streaming (Kafka, Event Hubs)\nPySpark or Snowpark\nTerraform or other infrastructure as code\nSemantic layers (dbt Semantic Layer, Cube) or data catalogs (Purview, DataHub)\nMulti-tenant data platforms or row-level security\nWorking with image, video, or sensor data alongside structured records\nLogistics, warehousing, or robotics domain experience","description_format":"text","description_chars":5473,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Drones & UAVs","Warehousing","Logistics & Warehouse Robotics","Drone AI"],"lifecycle":[{"event":"open","at":"2026-10-01T02:24:04Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":77,"reasons":["conf:3","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/gather-ai-software-development-engineer-ii-data-engineer","json_url":"https://alion.io/job/gather-ai-software-development-engineer-ii-data-engineer.json","meta":{"generated_at":"2026-10-02T00:32:32Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":484,"day_limit":5000,"remaining_today":4516,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}