368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$104k – $204k per year (Estimated)
Location
In office (Reading)
Seniority
Senior · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Hitachi is a major Japanese multinational conglomerate specializing in industrial equipment, power systems, railway infrastructure, and digital technologies. Through its Social Innovation Business and Lumada platform, the company combines operational technology with enterprise IT to modernize critical infrastructure globally. Operating across diverse sectors - including energy, mobility, healthcare, and green technologies - Hitachi drives digital transformation and sustainable development worldwide.

Function

Cloud & Data Engineering

Our Company

We’re Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world’s potential. We’re people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what’s now to what’s next. We make it happen through the power of acceleration.

Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don’t expect you to ‘fit’ every requirement - your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.

Job description

Job Summary

We are seeking a highly skilled Hands-On Data Engineering Lead with deep expertise in Apache Flink, AWS Managed Service for Apache Flink, Kafka, and AWS. This role requires a technical leader who actively architects, develops, debugs, optimizes, and supports production-grade real-time streaming platforms. The ideal candidate combines hands-on engineering depth with experience leading teams that deliver scalable, resilient, high-throughput, and low-latency data solutions in 24x7 environments.

Key Responsibilities

Hands-On Technical Leadership

  • Lead and mentor data engineers and architects while remaining directly involved in solution design, implementation, and troubleshooting.
  • Serve as the technical authority for Apache Flink, AWS Managed Service for Apache Flink, Kafka-based streaming, and associated AWS services.
  • Lead architecture, design, code, configuration, deployment, and production-readiness reviews.
  • Establish engineering standards, coding practices, test discipline, and production-support procedures.
  • Coordinate technical decisions across application, data, cloud-platform, performance, and operations teams.

Real-Time Streaming Architecture & Engineering

  • Architect, design, and develop production-scale streaming solutions using Apache Flink, AWS Managed Service for Apache Flink, Java or PyFlink, Flink SQL, and Kafka.
  • Apply stateful and event-time processing patterns, including keyed and broadcast state, state TTL, timers, windows, joins, watermarks, and delivery-semantics controls.
  • Design resilient checkpointing, savepoint, restart, recovery, and state-migration strategies.
  • Design Kafka topics, partitions, shard keys, routing, consumer groups, offsets, transactions, schemas, and source/sink integrations.
  • Design and operate sharded streaming architectures, including workload decomposition, shard-key selection, state distribution, rebalancing, cross-shard processing, parallelism, failure isolation, and recovery.
  • Build fault-tolerant streaming applications that meet defined throughput, latency, availability, and recovery objectives.

Performance Engineering & Optimization

  • Define measurable entry, exit, and acceptance criteria for load, peak, burst, replay, recovery, and long-running soak tests.
  • Verify that test inputs, replay behavior, duration, measurement windows, metric units, and outputs are representative and comparable.
  • Analyze throughput, latency, backpressure, state growth, checkpoints, recovery, resource utilization, data skew, and operating headroom.
  • Identify operator and stage-level bottlenecks and implement validated code, configuration, partitioning, or scaling improvements.
  • Document test conditions, findings, qualifications, risks, and recommendations using reproducible evidence.

Production Engineering & Troubleshooting

  • Act as a senior escalation point for complex production issues involving Apache Flink, AWS Managed Service for Apache Flink, and Kafka.
  • Diagnose checkpoint failures, savepoint recovery issues, backpressure, state growth, idle partitions, data skew, memory pressure, garbage collection, restarts, and throughput or latency degradation.
  • Analyze JobManager and TaskManager events, runtime configuration, logs, metrics, deployment history, and service behavior.
  • Lead evidence-based root-cause analysis and define corrective and preventive actions for material incidents.
  • Use controlled experiments to confirm or reject technical hypotheses and validate remediation effectiveness.
  • Engage AWS Support and service specialists when deeper platform analysis or service-limit clarification is required.

Observability & Operational Readiness

  • Define and implement monitoring for throughput, lag, backpressure, state size, checkpoints, restarts, failures, service events, and recovery.
  • Standardize metric definitions, units, aggregation windows, data sources, thresholds, alerting, and escalation paths.
  • Develop and review deployment, incident-triage, rollback, snapshot/savepoint recovery, scaling, and change-control procedures.
  • Prepare technical documentation, operational runbooks, and knowledge-transfer materials for engineering and support teams.

Software Engineering Excellence

  • Write, debug, test, optimize, and deploy production-grade Java, Python/PyFlink, and SQL code.
  • Perform code reviews and promote automated testing, code quality, infrastructure-as-code, and DevOps practices.
  • Build and improve CI/CD pipelines and deployment automation for streaming applications.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.
  • 10+ years of software engineering, data engineering, or distributed-systems experience, including 7+ years of recent hands-on Apache Flink engineering.
  • Deep production experience with AWS Managed Service for Apache Flink, including deployments, runtime behavior, scaling, service limits, snapshots, monitoring, troubleshooting, and recovery.
  • Strong command of the Flink DataStream API and Flink SQL, including stateful processing, event time, watermarks, windows, joins, timers, checkpoints, savepoints, and exactly-once concepts.
  • Demonstrated experience diagnosing backpressure, checkpoint behavior, state growth, skew, idle partitions, memory and GC issues, restarts, Kafka lag, and performance degradation.
  • Strong programming skills in Java, Python/PyFlink, and SQL, with experience developing and supporting production-grade streaming applications.
  • Hands-on experience with Kafka or Confluent Kafka and sharded distributed systems, including topic and shard-key design, partitioning, routing, consumer groups, rebalancing, state migration, cross-shard processing, failure isolation, and recovery.
  • Working knowledge of AWS services and controls supporting streaming platforms, including CloudWatch, S3, IAM, networking, service quotas, CI/CD, and infrastructure as code.
  • Experience building and operating highly available, high-throughput, low-latency platforms in 24x7 production environments.
  • Experience with monitoring, alerting, observability, incident response, root-cause analysis, and recovery planning.
  • Ability to communicate technical findings clearly and distinguish observed facts, estimates, hypotheses, proposals, and approved decisions.
  • Demonstrated ability to lead technical teams while remaining hands-on in engineering and troubleshooting activities.

Preferred Qualifications

  • Experience supporting global, mission-critical streaming platforms and mentoring engineering teams.
  • AWS certification or demonstrated equivalent AWS platform expertise.
  • Experience processing very large event volumes using multi-shard architectures and managing multi-terabyte state, workload skew, state migration, or disaster-recovery design.

About us

We’re a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.

Fostering innovation through diverse perspectives

Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.

We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.

How we look after you

We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We’re also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We’re always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you’ll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.

We’re proud to say we’re an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Reading
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$16k – $60k per year (Estimated) • In office • Full-Time • PhD • Mumbai
Python
SQL
Python
pySpark
Databases
Presto
Snowflake
AI/ML
Dagster
Prefect
Spark
DevOps
Amazon S3
AWS
CI/CD
Analytics
ETL/ELT
Power BI
Tableau
Apply
$93k – $126k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
Node JS
SQL
TypeScript
JavaScript
Databases
MS SQL
Oracle
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
GCP
Git
GitLab
GitLab CI
Jenkins
Management
Jira
QA
JMeter
Playwright
Postman
Rest-Assured
TestNG
Apply
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
In office • Full-Time • Bachelor's Degree • Gurgaon
C++
COBOL
Java
SQL
Apply
$35k – $78k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Apply
HMAX Senior Developer 12 hours ago
$42k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Genoa • Turin
Java
Java
Spring Boot
Databases
Apache Kafka
InfluxDB
PostgreSQL
DevOps
Azure
CI/CD
Docker
Grafana
Kubernetes
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Beijing
Analytics
Power BI
Apply
In office • Full-Time • Master's Degree • Västerås
Apply
$69k – $153k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Ludvika
Apply
Data Architect-68464 18 days ago
$119k – $247k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Reading
Python
SQL
Databases
Amazon Aurora
Amazon Redshift
Apache Kafka
DynamoDB
PostgreSQL
Snowflake
AI/ML
dbt
Spark
DevOps
AWS
CI/CD
Docker
Git
Kubernetes
Rest API
Amazon S3
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.