{"id":1280215,"url":"https://alion.io/job/supersourcing-senior-data-engineer","title":"Senior Data Engineer","company":{"id":59528,"name":"Supersourcing","domain":"supersourcing.com","url":"https://alion.io/company/supersourcing","size_band":"11-50","is_staffing_agency":true,"employer_type":"agency","is_intermediary":true,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":28000,"max_usd":58000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":9},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Airflow","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Dagster","optional":false},{"name":"Dask","optional":false},{"name":"Docker","optional":false},{"name":"Embeddings","optional":false},{"name":"Feature Store","optional":false},{"name":"FSDP","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"Machine Learning","optional":false},{"name":"ONNX Runtime","optional":false},{"name":"Prefect","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"RAG","optional":false},{"name":"Ray","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"TorchServe","optional":false},{"name":"Triton","optional":false},{"name":"Amazon Kinesis","optional":true},{"name":"Amazon SageMaker","optional":true},{"name":"Apache Hudi","optional":true},{"name":"Apache Iceberg","optional":true},{"name":"Apache Kafka","optional":true},{"name":"Delta Lake","optional":true},{"name":"Flink","optional":true},{"name":"Hugging Face","optional":true},{"name":"Kubeflow","optional":true},{"name":"Kubernetes","optional":true},{"name":"LLM","optional":true},{"name":"LoRA","optional":true},{"name":"MLFlow","optional":true},{"name":"PEFT","optional":true},{"name":"pgvector","optional":true},{"name":"Pinecone","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Qdrant","optional":true},{"name":"QLoRA","optional":true},{"name":"Terraform","optional":true},{"name":"Transformers","optional":true},{"name":"Vertex AI","optional":true},{"name":"Weaviate","optional":true},{"name":"Weights & Biases","optional":true}],"status":"live","first_seen_at":"2026-08-20T05:31:30Z","employer_posted_date":null,"last_verified_at":"2026-08-20T05:31:30Z","board_verified":false,"closed_at":null,"days_open":38,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":38},"description":"Position : Senior Data Engineer PyTorch / ML Data Platforms\n\nCompany : CodeWalnut\n\nEmployment Type : Full-Time with Supersourcing\n\nExperience : 5+ Years\n\nLocation : Bangalore\n\nWork Mode : Work From Office 5 Days a Week\n\nFunction : Data & AI Engineering\n\nRole Overview : \n\nWe are looking for a hands-on Senior Data Engineer who is equally strong in production data engineering and PyTorch. You will own the data layer supporting machine learning and agentic AI systems, covering data ingestion, transformation, feature engineering, training data curation, and production ML infrastructure.\n\nKey Responsibilities : \n\n- Design, build, and operate batch and streaming data pipelines at production scale.\n\n- Build and optimize PyTorch training and inference workflows, including custom Datasets, DataLoaders, DDP/FSDP, mixed precision, checkpointing, and reproducibility.\n\n- Own feature engineering and feature store design, ensuring training/serving parity and preventing data leakage.\n\n- Curate, version, and validate training datasets with data quality and drift monitoring.\n\n- Deploy and serve ML models using TorchServe, ONNX Runtime, Triton, or equivalent platforms.\n\n- Build data infrastructure for RAG and agentic AI systems, including embedding pipelines, vector stores, chunking, and evaluation datasets.\n\n- Implement pipeline observability, lineage, data contracts, monitoring, and alerting.\n\n- Collaborate with backend, frontend, platform, and client engineering teams to deliver end-to-end AI solutions.\n\n- Review code and mentor mid-level engineers.\n\nMandatory Skills : \n\n- 5+ years of Data Engineering / ML Engineering experience\n\n- Strong Python development experience\n\n- Hands-on PyTorch experience with production model deployment\n\n- Strong SQL and Data Modelling fundamentals\n\n- Experience with Apache Spark, Ray, Dask, or equivalent distributed processing technologies\n\n- Experience with Airflow, Dagster, Prefect, or similar orchestration tools\n\n- Experience with AWS / GCP / Azure data platforms\n\n- Docker, Git, and CI/CD\n\n- Strong communication and client-facing skills\n\nGood to Have : \n\n- LLM / Agentic AI experience\n\n- RAG, embeddings, and Vector Databases\n\n- Pinecone, Weaviate, Qdrant, pgvector\n\n- LoRA / QLoRA and Hugging Face Transformers\n\n- MLflow, Weights & Biases, Kubeflow, SageMaker, or Vertex AI\n\n- Kubernetes / Terraform\n\n- Kafka, Kinesis, Flink, or Spark Structured Streaming\n\n- Delta Lake, Apache Iceberg, or Hudi\n\n- GPU performance tuning and inference optimization\n\nSkills\nData Engineering, PyTorch, Data Infrastructure, Data Pipeline, Agentic AI, RAG, Data Modeling, SQL, Apache Flink, Spark","description_format":"text","description_chars":2627,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-26T02:00:24Z"}],"liveness":{"score":10,"band":"cold","label":"Long shot","p_open":0.6,"p_active":0.288,"p_room":0.55,"age_days":38,"expected_fill_days":31,"reasons":["seen:38","agency","velocity","win:tail"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/supersourcing-senior-data-engineer","json_url":"https://alion.io/job/supersourcing-senior-data-engineer.json","meta":{"generated_at":"2026-09-28T02:44:35Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1503,"day_limit":5000,"remaining_today":3497,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}