{"id":1272296,"url":"https://alion.io/job/draup-big-data-engineer","title":"Big Data Engineer","company":{"id":2764326,"name":"Draup","domain":"draup.com","url":"https://alion.io/company/draup","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Darwinbox","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"junior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":14000,"max_usd":34000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":11},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Airflow","optional":false},{"name":"Amazon S3","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Flink","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false}],"status":"live","first_seen_at":"2026-09-03T05:27:38Z","employer_posted_date":null,"last_verified_at":"2026-09-03T05:27:38Z","board_verified":false,"closed_at":null,"days_open":37,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":37},"description":"Role : Big Data Engineer\n\nJob Description :\n\nThe Big Data Engineer at Draup is responsible for building scalable techniques and processes for data storage, transformations, and analysis. The role includes designing, and implementation of the optimal, generic, and reusable data-platforms and pipelines. You will work with a very proficient, inquisitive, enthusiastic, and experienced team of developers, engineers, and product team. You will have ability to shape up the data products.\n\nWhat You Will Do :\n\n- Build scalable architectures for data storage, transformations, and analysis.\n\n- Work on data engineering projects to ensure pipelines are reliable, efficient, testable, & maintainable.\n\n- Design and develop solutions which are scalable, generic, and reusable.\n\n- Build and execute data warehousing, data mining and data modelling activities using agile development techniques.\n\n- Leading and taking ownership of big data projects successfully from scratch to production.\n\n- Bring the problem-solving attitude with focus on understanding the basics and implementing the use cases of organizational impact.\n\n- Collaborate with various teams including data science, backend, data harvesting and product teams.\n\nWhat You'll Need :\n\n- Proficient understanding of big data and distributed systems principles.\n\n- Must have good programming experience in Python.\n\n- Proficiency in Apache Spark (PySpark) is a must.\n\n- Good work experience in developing ETL and ELT solutions in a scalable way from multiple data sources.\n\n- Understanding in technologies like Relational and NoSQL datastores.\n\n- Working and conceptual Knowledge of MapReduce, HDFS, Amazon S3.\n\n- Ability to code and think in functional programming paradigm.\n\n- Enthusiastic about optimizing the code performance and resources of the system.\n\n- Ability to communicate complex technical concepts to both technical and non-technical audiences.\n\n- Takes ownership of all technical aspects of software development for assigned projects.\n\nWhat Will Give You an Advantage :\n\n- Expertise in big data infrastructure, distributed systems, data modelling, query processing and relational.\n\n- Involved in the design of big data solutions with Spark/HDFS/MapReduce/Flink.\n\n- Worked with different types of file-storage formats like Parquet, ORC, Avro, Sequence files etc.\n\n- Experience and understanding of cluster managers like YARN, Spark Standalone, Mesos or Kubernetes, etc.\n\n- Strong knowledge of data structures and algorithms.\n\n- Understands how to apply technologies to solve big data problems and to develop innovative big data solutions.\n\n- Someone with problem solving mind-set with good design and architectural patterns will be preferred.\n\n- Experience in working with AWS tools like EMR, Lambda, Glue or equivalent tools on other cloud systems.\n\n- Knowledge of workflow orchestration tools like Airflow, Jenkins.\n\nWho You Are :\n\n- B.E / B.Tech / M.E / M.Tech / M.S in Computer Science or software engineering.\n\n- Experience of 2-4 Years working with Big Data technologies.\n\n- Open to embrace the challenge of dealing with terabytes and petabytes of data on daily basis.\n\nSkills\nBig Data, Data Storage, Data Analytics, Data Warehousing, Data Mining, Data Modeling, Distributed Systems, Python, PySpark, ETL","description_format":"text","description_chars":3274,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-26T00:07:01Z"}],"visa":[],"liveness":{"score":19,"band":"cold","label":"Long shot","p_open":0.6,"p_active":0.566,"p_room":0.55,"age_days":37,"expected_fill_days":30,"reasons":["seen:37","win:tail","comp:junior"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/draup-big-data-engineer","json_url":"https://alion.io/job/draup-big-data-engineer.json","meta":{"generated_at":"2026-10-11T02:59:48Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":194,"day_limit":5000,"remaining_today":4806,"minute_limit":60,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2764326},"rest":"https://alion.io/mcp/rest/get_company?id=2764326"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fdraup-big-data-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fdraup-big-data-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fdraup-big-data-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/draup-big-data-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fdraup-big-data-engineer"}]}