First seen by Alion on Sep 30, 2026.
We're building the data backbone for large-scale SCADA deployments, moving telemetry from field devices, RTUs, PLCs, and historians into reliable, queryable platforms that operations, maintenance, and engineering teams depend on daily. This is not a generic ETL role. You'll work where operational technology meets modern data infrastructure: high-frequency time-series at scale, protocol-level integration, strict network segmentation, and data that people use to make real-time decisions about physical assets. If you've dealt with tag mapping, out-of-order sensor data, or the reality that a historian's timestamps and your ingestion clock rarely agree, you'll be at home here.
Responsibilities:
- Design, build, and operate ingestion pipelines from SCADA sources, historians, RTUs, PLCs, IEDs, and edge gateways into our central data platform.
- Work with industrial protocols and interfaces (OPC UA/DA, Modbus, DNP3, IEC 60870-5-104, IEC 61850, MQTT/Sparkplug B) and the connectors that sit on top of them.
- Build streaming and batch pipelines for high-volume time-series data, handling late arrivals, deadband compression, gaps, backfills, and clock drift.
- Model asset hierarchies and tag namespaces so that raw point IDs become something an engineer can actually query.
- Own data quality: validation, reconciliation against source historians, alerting on stale or silent tags, and clear lineage from field device to dashboard.
- Partner with OT and network teams to move data across segmented environments (Purdue-model zones, DMZs, unidirectional gateways) without compromising security posture.
- Support downstream consumers with condition monitoring, predictive maintenance, energy analytics, KPI, and regulatory reporting.
- Instrument and monitor your own pipelines; participate in an on-call rotation for critical data flows [adjust or remove].
Requirements:
- 3+ years building production data pipelines.
- Strong Python and SQL.
- Hands-on experience with a distributed streaming or processing framework (Kafka, Flink, Spark Structured Streaming, or equivalent).
- Time-series data at scale: You've worked with a purpose-built store (TimescaleDB, InfluxDB, ClickHouse, or a historian) and understand why a general-purpose RDBMS struggles here.
- Workflow orchestration (Airflow, Dagster, Prefect, or similar).
- Comfort with Linux, containers, Git, and CI/CD.
- Practical grasp of data modelling, schema evolution, and idempotent pipeline design.
- Ability to work with engineers who speak in tag names and equipment IDs rather than tables and columns.
Strongly preferred:
- Direct experience with SCADA or industrial historian platforms AVEVA/OSIsoft PI, Wonderware, GE Proficy/iFIX, Siemens WinCC, Schneider EcoStruxure/ClearSCADA, and Ignition.
- Railway or metro domain experience in traction power SCADA, tunnel ventilation, ECS/BMS, signalling and interlocking data, ATS, AFC, rolling stock condition monitoring, depot systems, or passenger information systems.
- Familiarity with OT security practices and standards (IEC 62443, network segmentation, read-only data extraction patterns).
- Cloud data platforms (Azure IoT Hub / AWS IoT / GCP) and edge-to-cloud architectures.
- Exposure to rail assurance standards (EN 50126/50128/50129) or working inside a regulated engineering environment, dbt, Kubernetes, or Terraform.

