Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 19, 2026. Infosys scores B on the Alion truth index.
About the job:
Join a team where data engineering meets intelligent automation to power next-generation analytics and AI experiences. In this role, you’ll help build reliable, scalable data pipelines and enable GenAI-ready datasets that support experimentation, model development, and business decision-making. You’ll collaborate closely with analysts, data scientists, and platform teams to turn raw, fragmented data into trusted, well-modeled, and accessible assets. If you enjoy solving complex data challenges, improving performance, and bringing structure to fast-moving AI initiatives, this is a great opportunity to grow your impact. You’ll work in a culture that values ownership, continuous learning, and practical innovation-where clean data, strong engineering, and thoughtful collaboration come together to deliver real outcomes.
Responsibilities
Key Responsibilities:
- Design, build, and maintain robust ETL/ELT pipelines to ingest, transform, and curate data from multiple sources.
- Develop and optimize data models and curated datasets to support analytics, reporting, and AI/ML workloads.
- Implement data quality checks, validation rules, and monitoring to ensure accuracy, completeness, and reliability.
- Enable GenAI initiatives by preparing high-quality datasets for downstream consumption (e.g., feature-ready and retrieval-ready data).
- Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver reusable data assets.
- Troubleshoot pipeline failures and performance bottlenecks; improve scalability, latency, and cost efficiency.
- Maintain documentation for pipelines, transformations, lineage, and operational runbooks to support maintainability.
Minimum Qualifications:
- BTECH, MTECH, MCA, MSC or equivalent education.
- 3-5 years of experience in data engineering with hands-on ownership of production-grade pipelines.
- Strong experience in ETL processes including extraction, transformation, orchestration, and scheduling.
- Working exposure to GenAI-oriented data preparation needs and supporting AI/ML data workflows.
- Ability to collaborate with stakeholders to translate requirements into scalable data solutions.
Technical requirements
Good to have skills:
SQL, Python, Apache Spark, Airflow, Data Modeling, data engineering, ai
Additional responsibilities
Preferred Qualifications:
- Experience designing scalable data architectures and implementing reusable data frameworks for multiple use cases.
- Familiarity with building datasets for GenAI use cases such as retrieval workflows and knowledge augmentation patterns.
- Proven ability to improve pipeline reliability through automation, alerting, and proactive monitoring.
- Strong problem-solving skills with a track record of optimizing transformations and reducing end-to-end processing time.
- Experience working in agile teams and contributing to code reviews, documentation, and engineering best practices.
Preferred skills
Technology->Data Engineering->Databricks,Technology->AI-Responsible AI->Responsible AI->explainable ai
Education
MCA,MSc,MTech,Bachelor of Engineering,BTech

