Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 25, 2026. Finnomena scores B on the Alion truth index.
Job Description
Objectives of this role / About this job:
The Data Analytics Engineer Associate supports the development and day-to-day operation of data pipelines that provide reliable data for analytics and business operations. Working with senior engineers, you will help ingest, transform, and validate data using SQL, Python, and PySpark, monitor batch and streaming pipelines, and investigate data issues within clearly defined tasks.
This role is suitable for fresh graduates or candidates with up to two years of relevant experience who are interested in data engineering. You will gain hands-on experience with the team’s Google Cloud data stack, including BigQuery, Dataform, Google Cloud Storage, Managed Service for Apache Airflow, Managed Service for Apache Spark, Pub/Sub, and Datastream, with guidance and code reviews from experienced team members.
Responsibilities:
- Develop and maintain assigned batch ETL/ELTjobs using SQL, Python, and PySpark, following existing patterns and guidance from senior engineers.
- Support data ingestion from databases, APIs, files, and Google Workspace sources, including basic data cleansing, transformation, and source-to-target validation.
- Assist with the development and maintenance of streaming data pipelines using services such as Google Cloud Pub/Sub and Datastream, following established patterns and documented procedures.
- Implement and maintain data quality checks to identify missing, duplicate, invalid, or delayed data, and help investigate discrepancies.
- Monitor batch and streaming pipelines, data freshness, and pipeline health; review logs, troubleshoot routine issues, and escalate complex incidents with clear supporting information.
- Assist with approved reruns and data reprocessing following documented procedures and senior guidance.
- Prepare datasets for downstream analytics and reporting, and help maintain dashboards used to monitor pipeline status and data quality.
- Document data mappings, transformation logic, validation rules, and troubleshooting steps; use Git and participate in code reviews and testing before changes are released.

