{"id":1513033,"url":"https://alion.io/job/stratzi-ai-data-engineer","title":"Data Engineer","company":{"id":4474,"name":"Stratzi.ai","domain":"stratzi.ai","url":"https://alion.io/company/stratzi-ai","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":15000,"max_usd":37000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":13},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon ECS","optional":false},{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"ClickHouse","optional":false},{"name":"Dagster","optional":false},{"name":"dbt","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Flink","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"InfluxDB","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"MQTT","optional":false},{"name":"OPC UA","optional":false},{"name":"Prefect","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Time Series Forecasting","optional":false},{"name":"TimescaleDB","optional":false},{"name":"PostgreSQL","optional":true}],"status":"live","first_seen_at":"2026-09-30T08:35:59Z","employer_posted_date":null,"last_verified_at":"2026-09-30T08:35:59Z","board_verified":false,"closed_at":null,"days_open":1,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":1},"description":"We're building the data backbone for large-scale SCADA deployments, moving telemetry from field devices, RTUs, PLCs, and historians into reliable, queryable platforms that operations, maintenance, and engineering teams depend on daily. This is not a generic ETL role. You'll work where operational technology meets modern data infrastructure: high-frequency time-series at scale, protocol-level integration, strict network segmentation, and data that people use to make real-time decisions about physical assets. If you've dealt with tag mapping, out-of-order sensor data, or the reality that a historian's timestamps and your ingestion clock rarely agree, you'll be at home here.\nResponsibilities:\nDesign, build, and operate ingestion pipelines from SCADA sources, historians, RTUs, PLCs, IEDs, and edge gateways into our central data platform.\nWork with industrial protocols and interfaces (OPC UA/DA, Modbus, DNP3, IEC 60870-5-104, IEC 61850, MQTT/Sparkplug B) and the connectors that sit on top of them.\nBuild streaming and batch pipelines for high-volume time-series data, handling late arrivals, deadband compression, gaps, backfills, and clock drift.\nModel asset hierarchies and tag namespaces so that raw point IDs become something an engineer can actually query.\nOwn data quality: validation, reconciliation against source historians, alerting on stale or silent tags, and clear lineage from field device to dashboard.\nPartner with OT and network teams to move data across segmented environments (Purdue-model zones, DMZs, unidirectional gateways) without compromising security posture.\nSupport downstream consumers with condition monitoring, predictive maintenance, energy analytics, KPI, and regulatory reporting.\nInstrument and monitor your own pipelines; participate in an on-call rotation for critical data flows [adjust or remove].\nRequirements:\n3+ years building production data pipelines.\nStrong Python and SQL.\nHands-on experience with a distributed streaming or processing framework (Kafka, Flink, Spark Structured Streaming, or equivalent).\nTime-series data at scale: You've worked with a purpose-built store (TimescaleDB, InfluxDB, ClickHouse, or a historian) and understand why a general-purpose RDBMS struggles here.\nWorkflow orchestration (Airflow, Dagster, Prefect, or similar).\nComfort with Linux, containers, Git, and CI/CD.\nPractical grasp of data modelling, schema evolution, and idempotent pipeline design.\nAbility to work with engineers who speak in tag names and equipment IDs rather than tables and columns.\nStrongly preferred:\nDirect experience with SCADA or industrial historian platforms AVEVA/OSIsoft PI, Wonderware, GE Proficy/iFIX, Siemens WinCC, Schneider EcoStruxure/ClearSCADA, and Ignition.\nRailway or metro domain experience in traction power SCADA, tunnel ventilation, ECS/BMS, signalling and interlocking data, ATS, AFC, rolling stock condition monitoring, depot systems, or passenger information systems.\nFamiliarity with OT security practices and standards (IEC 62443, network segmentation, read-only data extraction patterns).\nCloud data platforms (Azure IoT Hub / AWS IoT / GCP) and edge-to-cloud architectures.\nExposure to rail assurance standards (EN 50126/50128/50129) or working inside a regulated engineering environment, dbt, Kubernetes, or Terraform.","description_format":"text","description_chars":3307,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Predictive Analytics","AI Consulting & Integration"],"lifecycle":[{"event":"open","at":"2026-09-30T08:35:59Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":18,"reasons":["seen:0","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/stratzi-ai-data-engineer","json_url":"https://alion.io/job/stratzi-ai-data-engineer.json","meta":{"generated_at":"2026-10-01T20:58:57Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3408,"day_limit":5000,"remaining_today":1592,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}