This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Junior Data Engineer based in India.
As a Junior Data Engineer, you will contribute to the development and maintenance of data systems that support analytics, reporting, business insights, and AI-driven decision-making. Working closely with experienced engineers, you will gain hands-on exposure to data pipelines, SQL, Python, and modern cloud data platforms. You will help transform and validate data while learning how scalable and reliable data workflows are designed and maintained. The role offers structured mentorship and an opportunity to build practical data engineering skills from the beginning of your career. You will collaborate with analytics and business teams to understand data requirements and help deliver useful, high-quality datasets. Along the way, you will participate in code reviews, documentation, testing, and continuous improvements to engineering practices. This is a strong opportunity for a curious early-career engineer looking to grow within a modern AI and data environment.
Accountabilities
- Assist senior engineers in building, maintaining, and improving data pipelines.
- Write, test, and optimize SQL queries to extract, transform, validate, and analyze data.
- Support batch data ingestion and basic Extract, Transform, Load (ETL) and Extract, Load, Transform (ELT) workflows using Python.
- Perform basic data quality checks and identify, document, and flag inconsistencies or anomalies.
- Support the maintenance and operation of cloud-based data warehouse or lakehouse environments such as Snowflake and Databricks.
- Document data flows, technical specifications, test cases, and other relevant engineering processes.
- Participate in code reviews and apply feedback to improve code quality, reliability, and maintainability.
- Collaborate with analytics and business teams to understand data requirements and support data-related initiatives.
- Proactively learn new tools, programming languages, technologies, and data engineering best practices.
- Communicate project progress, technical challenges, and blockers clearly with team members and senior engineers.
- For fresh graduates: Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related field.
- For candidates with experience: 1-2 years of hands-on experience in data engineering, analytics engineering, or a similar role.
- Basic to working knowledge of SQL, including SELECT statements, JOINs, aggregations, and, for experienced candidates, basic query performance tuning.
- Basic to hands-on knowledge of Python programming.
- Understanding of relational databases, data structures, data warehousing concepts, tables, schemas, and fundamental ETL principles.
- For experienced candidates, exposure to PySpark is an advantage.
- Exposure to cloud data warehouse or lakehouse platforms such as Snowflake, Databricks, Amazon Redshift, or Google BigQuery is preferred for candidates with experience.
- Basic exposure to data pipeline or ETL tools such as Airflow, dbt, or Fivetran is a plus.
- Familiarity with version control systems such as Git, GitHub, or Bitbucket is preferred for experienced candidates.
- Basic understanding of at least one cloud platform, such as Amazon Web Services (AWS), Microsoft Azure, or Google Cloud Platform (GCP).
- Strong analytical thinking, problem-solving ability, curiosity, and attention to detail.
- Good verbal and written communication skills.
- Ability to collaborate effectively and work within an Agile team environment.
- Academic or internship projects involving data pipelines or analytics are a plus.
- Familiarity with streaming technologies such as Kafka, NoSQL databases, Business Intelligence (BI) and visualization tools such as Power BI, Tableau, or Looker, or Docker and containerization concepts is advantageous.
- Full-time opportunity within an AI and data-focused engineering environment.
- Remote working arrangement.
- Strong mentorship and guidance from experienced data engineers.
- Hands-on exposure to modern data engineering tools, cloud data platforms, and engineering practices.
- Opportunity to work with technologies such as SQL, Python, Snowflake, Databricks, and cloud platforms.
- Practical experience supporting data pipelines, analytics, reporting, and AI-driven decision-making.
- Opportunity to develop technical skills through code reviews, documentation, testing, and collaboration with cross-functional teams.
- Exposure to a broad range of data engineering technologies and opportunities to learn new tools and practices.

