{"id":1533109,"url":"https://alion.io/job/fam-data-engineering-intern","title":"Data Engineering Intern","company":{"id":3777403,"name":"Fam","domain":"famapp.com","url":"https://alion.io/company/fam-7","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"intern","employment_type":"internship","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Roorkee, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":15500,"max_usd":45000,"period":"year","method":null,"sample_n":3503},"experience_years_min":null,"visa_sponsorship":true,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Airflow","optional":false},{"name":"Claude","optional":false},{"name":"Claude Code","optional":false},{"name":"Copilot","optional":false},{"name":"Cursor","optional":false},{"name":"Java","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"Python","optional":false},{"name":"Scala","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Trino","optional":false},{"name":"Amazon S3","optional":true},{"name":"AWS","optional":true},{"name":"Chroma","optional":true},{"name":"Embeddings","optional":true},{"name":"Git","optional":true},{"name":"Hallucination","optional":true},{"name":"IAM","optional":true},{"name":"Linux","optional":true},{"name":"llama.cpp","optional":true},{"name":"LLM","optional":true},{"name":"Metabase","optional":true},{"name":"Ollama","optional":true},{"name":"OpenSearch","optional":true},{"name":"pgvector","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Power BI","optional":true},{"name":"Qdrant","optional":true},{"name":"RAG","optional":true},{"name":"Superset","optional":true}],"status":"live","first_seen_at":"2026-09-30T17:09:26Z","employer_posted_date":null,"last_verified_at":"2026-09-30T17:09:26Z","board_verified":false,"closed_at":null,"days_open":1,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":1},"description":"About Fam (previously FamPay)\nFam is India's first payments app for everyone above 11. FamApp helps make online and offline payments through UPI and FamCard. We are on a mission to raise a new, financially aware generation, and drive 250 million+ young users in India to kickstart their financial journey super early in their life.\nWe're reimagining how the next generation experiences fintech-going beyond payments to build a lifestyle brand that blends money, identity, and everyday experiences into one seamless, intuitive journey.\nFounded in 2019 by IIT Roorkee alumni, Fam is backed by some of the most respected investors around the world like Elevation Capital, Y-Combinator, Peak XV (Sequoia Capital) India, Venture Highway, Global Founder's Capital and the likes of Kunal Shah, Amrish Rao as angel investors.\nAbout the role\nWe are looking for Data Engineering Interns to work closely with the Fam Data Team in building and operating our data platform. You will contribute to our lakehouse, data pipelines and data quality checks, which power analytics, product and compliance reporting for a UPI/fintech platform serving millions of users. You will begin with well-scoped tasks under the guidance of a mentor and progressively take end-to-end ownership of individual pipeline tasks.\nOn the Job\nPlatform Understanding: Learn how our OLTP source systems feed the OLAP lakehouse. Read table schemas, including column types and partition columns, and trace the end-to-end flow of an individual Airflow DAG task.\nSQL & Transforms: Write and modify SQL queries and simple Spark/SQL transformations under guidance, and validate the correctness of their output.\nPipeline Quality: Add basic row-count and null checks, along with meaningful logging, to the pipeline tasks you own. Understand the importance of idempotent pipelines and apply this principle in your changes.\n Pipeline On-call (Shadow): Participate in the pipeline on-call rotation alongside an experienced engineer. Identify failed runs and data freshness breaches, escalate with the relevant DAG and run details, and execute existing backfill playbooks under guidance.\nCost-aware Querying: Apply partition filters instead of full table scans, and understand that storage grows with data volume and every query incurs compute cost.\nAnalyst Support: Collaborate with product analysts as a peer by fixing queries, directing them to the right datasets in the data catalog, and publishing simple, reusable views.\nAI-assisted Data Work: Use the tools provided, such as the Trino MCP, Text-to-SQL and data discovery tools, in day-to-day work, while ensuring that PII is never included in prompts.\n AI-assisted Development: Use AI coding assistants (Cursor, Claude Code, Copilot or similar) to accelerate development tasks such as writing SQL, transformations, tests and scripts. Review, test and fully understand all generated code before raising a pull request; accountability for the code remains with you.\nExperimentation: Identify small improvement ideas, build prototypes and present them during sprint reviews. The focus is on learning rather than measurable impact.\nMust-haves:\nFinal-year student or recent graduate in Computer Science, Information Technology or a related field, or equivalent practical experience.\nClear understanding of OLTP vs OLAP systems and their respective use cases.\nAbility to write correct, readable SQL, including joins, aggregations, filters and basic window functions.\nWorking knowledge of Python; familiarity with Scala or Java is a plus.\nAbility to read a table schema and explain the data it represents.\nBasic understanding of how LLM / GPT models work, including tokens, context windows, prompting and the causes of hallucination.\nPractical experience using AI coding assistants for development tasks, with the judgement to verify their output.\nAt least one academic project or internship involving data or backend engineering that you can explain in detail.\nProficiency with Git, Linux and the command line.\nA curious mindset: you question unexpected data patterns and escalate issues early.\nGood to have\nExposure to Apache Spark, Airflow or any data processing or orchestration framework.\nUnderstanding of idempotency, partitioning and columnar file formats such as Parquet.\nAbility to read a simple query plan and identify full table scans.\nExperience running an open-weight LLM locally (Ollama, llama.cpp or similar) or building a small embedding and similarity search prototype.\nBasic understanding of Retrieval-Augmented Generation (RAG), including chunking, embeddings, retrieval and grounding responses in source data.\nExposure to vector databases such as pgvector, Qdrant, Chroma or OpenSearch.\nFamiliarity with AWS fundamentals such as S3 and IAM.\nExposure to BI tools such as Superset, Metabase or Power BI.\nInterest in fintech, payments or data privacy.","description_format":"text","description_chars":4864,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-30T17:09:26Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":21,"reasons":["seen:0","win:early","comp:junior"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/fam-data-engineering-intern","json_url":"https://alion.io/job/fam-data-engineering-intern.json","meta":{"generated_at":"2026-10-01T19:12:25Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1663,"day_limit":5000,"remaining_today":3337,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}