Salary
≈ $28k – $56k per year (Estimated)
Location
In office (Pune)
Seniority
Architect · 5+ years exp
First seen by Alion on Sep 30, 2026.
Overview
Company
Impact
Profile match
Stratzi.ai is an AI consulting firm building custom AI systems for businesses, offering Gen-AI, predictive analytics, automation and enterprise search solutions to help organizations optimize and scale operations.
The core responsibilities for the job include the following:
Data lake architecture:
- Own the lakehouse design: storage layout, open table format (Iceberg, Delta Lake, or Hudi), catalog, and the medallion/zone model from raw ingest through to curated, query-ready datasets.
- Define partitioning, clustering, and compaction strategies tuned for time-series access patterns, long time-range scans over narrow tag selections, and recent-window queries.
- Solve the small-files problem that streaming SCADA ingestion inevitably creates, and set file sizing, sort order, and maintenance job standards.
- Design schema evolution and versioning so tag additions, renames, and vendor upgrades don't break downstream consumers.
- Build rollup, downsampling, and retention tiers: raw at full fidelity, aggregates for dashboards, archival for regulatory, and RAMS analysis.
- Choose and own the query layer (Trino/Presto, Spark SQL, DuckDB, or warehouse engine) and tune it for the workloads that actually matter.
Database engineering:
- Own the relational and time-series databases behind the platform schema design, indexing, query tuning, connection, and resource management.
- Capacity planning, HA/replication, backup, and tested restore procedures, upgrade paths.
- Define how the lake and the operational stores coexist: what stays hot in a time-series DB, what lands in the lake, and how the two reconcile against source historians.
Governance and operations:
- Data cataloguing, lineage, access control, and audit, including how OT-sourced data is classified and who can reach it.
- Cost and storage growth management [cloud spend / on-prem capacity, as applicable].
- Set standards, review designs, and mentor the data engineering team.
- Work with OT and network teams on data movement across segmented zones (Purdue model, DMZ, unidirectional gateways) without weakening the security boundary.
Requirements:
- 5+ years in data engineering, database engineering, or data architecture, with 3+ years specifically on data lake or lakehouse platforms in production.
- Deep expertise in at least one open table format (Apache Iceberg, Delta Lake, or Apache Hudi), including maintenance operations, snapshot management, and failure modes.
- Strong object storage fundamentals (S3, ADLS, GCS, or on-prem MinIO/Ceph layout, consistency, lifecycle policies, cost, or capacity implications).
- Distributed query engines: Trino/Presto, Spark, or equivalent, with real tuning experience.
- Expert SQL and strong RDBMS administration (PostgreSQL, Oracle, SQL Server, or similar): you can read a query plan and fix what's wrong.
- Time-series data at scale: TimescaleDB, InfluxDB, ClickHouse, or a historian, and a clear view of where each fits.
- Metadata and catalog systems (Hive Metastore, AWS Glue, Unity Catalog, Nessie, Polaris, or similar).
- Python and a JVM language, or equivalent depth for platform-level work.
- Infrastructure literacy: Linux, containers, Kubernetes, IaC, CI/CD.
- Track record of setting technical standards other engineers build on.
Strongly preferred:
- Industrial or OT data experience historians (AVEVA/OSIsoft PI, Wonderware, GE Proficy, Siemens WinCC, Ignition), tag namespaces, and asset hierarchies.
- Railway or metro domain experience in traction power, signalling and interlocking, tunnel ventilation, ECS/BMS, rolling stock condition monitoring, AFC, and depot systems.
- Air-gapped or on-premise deployments and the constraints that come with them, OT security standards (IEC 62443), and segmented network architectures.
- ISA-95 or equivalent asset information modelling, dbt, data quality frameworks (Great Expectations, Soda), and lineage tooling (OpenLineage, DataHub, Amundsen).
- Rail assurance standards (EN 50126/50128/50129) or other regulated engineering environments.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,061,370 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Data Science
Similar stack
Same company
Pune
OT Integration & Data Engineer
1 day ago
≈ $45k – $108k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Madrid
SQL
Databases
InfluxDB
AI/ML
Time Series Forecasting
IoT
MQTT
OPC UA
Apply
In office
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
AI/ML
Spark
Airflow
dbt
DevOps
GCP
Azure
CI/CD
AWS
Analytics
ETL/ELT
Apply
ETIC, AWS Data Engineer, Senior Associate
1 year ago
≈ $61k – $160k per year (Estimated) • In office • Full-Time • Cairo
AI/ML
Hadoop
DevOps
AWS
Analytics
Azure Data Factory
Management
Agile
Apply
≈ $20k – $39k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Mumbai
Databases
Google BigQuery
BigQuery
AI/ML
Hadoop
Airflow
Machine Learning
DevOps
GCP
AWS
Analytics
ETL/ELT
Azure Data Factory
Management
Agile
Apply
ETIC, Fabric Data Engineer, Manager
1 year ago
≈ $60k – $158k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Cairo
Python
SQL
Python
pySpark
Databases
Microsoft Fabric
AI/ML
Hadoop
Spark
Machine Learning
DevOps
CI/CD
Git
AWS
Analytics
Power BI
Azure Data Factory
Dimensional Modeling
Management
Agile
Apply
.NET разработчик, Middle, удаленно, Аутстафф
2 days ago
up to $29k per year • Remote (likely EAEU) • Full-Time • 1+ year exp
JavaScript
TypeScript
SQL
C#
Node JS
C#
.NET
Entity Framework Core
Databases
PostgreSQL
MS SQL
Apache Kafka
Frontend
React.js
DevOps
Rest API
Git
Kubernetes
Management
Agile
Apply
.NET специалисты (AI), удаленно, Аутстафф
2 days ago
up to $29k per year • Remote (likely EAEU) • Full-Time
JavaScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Dapper
Databases
PostgreSQL
RabbitMQ
MS SQL
Apache Kafka
AI/ML
Copilot
Cursor
Claude
Claude Code
Function Calling
LLM
OpenAI Codex
DevOps
Rest API
CI/CD
Git
Docker
TeamCity
GitHub
QA
Swagger
Postman
SoapUI
Apply
Data Scientist - Clearance Required
2 days ago
$130k – $160k per year • In office • Secret • 2+ years exp • Bachelor's Degree • Tysons
Python
SQL
AI/ML
Scikit-learn
NumPy
Time Series Forecasting
Machine Learning
DevOps
Git
Analytics
Tableau
Power BI
Management
Confluence
Agile
Scrum
Apply
Аналитик данных
2 days ago
$19k – $23k per year (net) • Hybrid • Bachelor's Degree • Moscow
Python
SQL
AI/ML
NumPy
Apply
Senior Product Analyst
2 days ago
$59k per year • Hybrid • Full-Time • London
Python
SQL
SAS
Databases
Snowflake
Analytics
Tableau
Power BI
Microsoft Excel
Apply
Data Scientist
2 days ago
≈ $15k – $35k per year (Estimated) • In office • 4+ years exp • Pune
Python
SQL
AI/ML
Weights & Biases
MLFlow
Fine-tuning
Embeddings
Scikit-learn
Prompt Engineering
TensorFlow
NumPy
PyTorch
LLM
RAG
Time Series Forecasting
Structured Outputs
LLM Guardrails
Apply
Data Engineer
2 days ago
≈ $15k – $37k per year (Estimated) • In office • 3+ years exp • Pune
Python
SQL
Databases
PostgreSQL
ClickHouse
TimescaleDB
InfluxDB
Apache Kafka
AI/ML
Spark
Dagster
dbt
Prefect
Flink
Time Series Forecasting
DevOps
Terraform
GCP
Azure
CI/CD
Git
AWS
Kubernetes
Amazon ECS
Linux
Analytics
ETL/ELT
IoT
MQTT
OPC UA
Apply
Senior Data Scientist(6+ years with Machine Learning, NLP, Applied AI, or AI Product Development)
1 day ago
≈ $17k – $34k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Pune
Python
SQL
Python
FastAPI
AI/ML
LangGraph
AutoGen
LangChain
Embeddings
Prompt Engineering
AI Agents
NLP
CrewAI
RAG
Machine Learning
DevOps
Azure
CI/CD
Git
AWS
Docker
Management
Agile
Apply
Lead Data Scientist
1 day ago
≈ $29k – $51k per year (Estimated) • In office • Full-Time • Pune
AI/ML
Machine Learning
Apply
Lead Data Engineer (Data Platforms)
1 day ago
≈ $34k – $74k per year (Estimated) • In office • Full-Time • 8+ years exp • Pune
Databases
Snowflake
Databricks
Apache Iceberg
Delta Lake
DevOps
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Azure AKS
FinOps
Amazon S3
IAM
Apply
Data Engineering Operations Lead
1 day ago
≈ $24k – $43k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Pune
Python
Java
SQL
Scala
Python
pySpark
Databases
Databricks
Delta Lake
MS SQL
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
Airflow
MLFlow
Pandas
DevOps
GCP
CI/CD
Git
Incident Management
SLI/SLO/SLA
Analytics
ETL/ELT
Management
Agile
Apply
IN_Manager_-SAP Service Management_SAP_Advisory_Pune
12 hours ago
≈ $8.5k – $25k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Pune
ABAP
Apply
This is one of many
1,061,370 more open roles from verified company boards, updated every day.

