{"id":1227571,"url":"https://alion.io/job/hiringblaze-ai-data-engineer","title":"AI Data Engineer","company":{"id":3800132,"name":"HiringBlaze","domain":"hiringblaze.com","url":"https://alion.io/company/hiringblaze","size_band":null,"is_staffing_agency":true,"employer_type":"agency","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Hyderabad, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":23000,"max_usd":47000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":51},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Airflow","optional":false},{"name":"Amazon S3","optional":false},{"name":"Apache Kafka","optional":false},{"name":"C++","optional":false},{"name":"Cohere SDK","optional":false},{"name":"Embeddings","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Flink","optional":false},{"name":"Haystack","optional":false},{"name":"Knowledge Graph","optional":false},{"name":"LangChain","optional":false},{"name":"LlamaIndex","optional":false},{"name":"LLM","optional":false},{"name":"Milvus","optional":false},{"name":"Neo4j","optional":false},{"name":"OpenAI","optional":false},{"name":"Pinecone","optional":false},{"name":"Prefect","optional":false},{"name":"Python","optional":false},{"name":"Qdrant","optional":false},{"name":"Quantization","optional":false},{"name":"RAG","optional":false},{"name":"Rust","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Synthetic Data","optional":false},{"name":"Unstructured.io","optional":false},{"name":"Weaviate","optional":false},{"name":"Windows","optional":false}],"status":"live","first_seen_at":"2026-09-24T04:11:11Z","employer_posted_date":null,"last_verified_at":"2026-09-24T04:11:11Z","board_verified":false,"closed_at":null,"days_open":6,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":6},"description":"Description : \n\nRole : AI Data Engineer\n\n- Location : Rai Durg, Hyderabad\n\n- Work mode- Hybrid model working (3 days work from office)\n\n- Experience : 5 - 8 Years (Minimum 5 years- AI Data Engineer)\n\n- Mandatory Skills : DVC (Data Version Control) and Airflow, Apache Spark, Flink, and Kafka, Advanced level Python and AI logic and Rust (or C++), Vector Database Mastery like configuration of HNSW indexes, scalar quantization, and metadata filtering strategies\n\n- Qualification : Bachelor of Engineering - Bachelor of Technology (B.E./B.Tech.)\n\n- Notice period : Immediate / early joiners (Max. 15-30 days)\n\n- Interview Process : 2 - 3 Technical rounds\n\nImportant Note : \n\n- We are currently prioritizing immediate / early joiners (maximum 15-30 days- notice period above 30 days will be automatically rejected.).\n\n- All mandatory technical skills must be clearly highlighted within the project descriptions in your resume, not just listed in the Skills or Roles & Responsibilities sections .\n\nPosition Overview : \n\nWe are seeking a hardcore, hands-on AI Data Engineer to build the high-performance data infrastructure required to power autonomous AI agents. You won't just be moving data from A to B; you will be architecting Dynamic Context Windows, managing Real-time Semantic Indexes, and building Self-Cleaning Data Pipelines that feed our \"Super Employee\" agents.\n\nKey Responsibilities : \n\n- Vector & Graph ETL : Design and maintain pipelines that transform unstructured data (PDFs, emails, logs, chats) into optimized embeddings for Vector Databases (Pinecone, Weaviate, Milvus).\n\n- Semantic Data Modeling : Engineer data structures that optimize for Retrieval-Augmented Generation (RAG), ensuring agents find the \"needle in the haystack\" in milliseconds.\n\n- Knowledge Graph Construction : Build and scale Knowledge Graphs (Neo4j) to represent complex relationships in our trading and support data that standard vector search misses.\n\n- Automated Data Labeling & Synthetic Data : Implement pipelines using LLMs to auto-label datasets or generate synthetic edge cases for agent training and evaluation.\n\n- Stream Processing for Agents : Build real-time data \"listeners\" (Kafka/Flink) that feed live context to agents, allowing them to react to market or support events as they happen.\n\n- Data Reliability & \"Drift\" Detection : Build monitoring for \"Embedding Drift\", identifying when the statistical distribution of your data changes and the agent's \"knowledge\" becomes stale.\n\nQualifications : \n\n- Vector Database Mastery : Expert-level configuration of HNSW indexes, scalar quantization, and metadata filtering strategies within Pinecone, Milvus, or Qdrant.\n\n- Advanced Python & Rust : Proficiency in Python for AI logic and Rust (or C++) for high-performance data processing and custom embedding functions.\n\n- Big Data Ecosystem : Hands-on experience with Apache Spark, Flink, and Kafka in a high-throughput environment (Trading/FinTech preferred).\n\n- LLM Data Tooling : Deep experience with Unstructured.io, LlamaIndex, or LangChain for document parsing and chunking strategy optimization.\n\n- MLOps & DataOps : Mastery of DVC (Data Version Control) and Airflow/Prefect for managing complex, non-linear AI data workflows.\n\n- Embedding Models : Understanding of how to fine-tune embedding models (e.g., BGE, Cohere, or OpenAI) to better represent domain-specific (Trading) terminology.\n\nAdditional qualifications : \n\n- Chunking Strategy Architect : You don't just \"split text.\" You implement Semantic Chunking and Parent-Child retrieval strategies to maximize LLM context relevance.\n\n- Cold/Warm/Hot Storage Strategy : Managing cost and latency by tiering data between Vector DBs (Hot), SQL/NoSQL (Warm), and S3/Data Lakes (Cold).\n\n- Privacy & Redaction Pipelines : Building automated PII (Personally Identifiable Information) redaction into the ingestion layer to ensure agents never \"see\" or \"leak\" sensitive user data.\n\nWhy Join ?\n\n- Opportunity to lead transformative initiatives, modernizing legacy systems and shaping the future of trading technology.\n\n- Work with cutting-edge technologies in a dynamic, fast-paced environment.\n\n- Competitive compensation, professional growth opportunities, and the chance to work with industry-leading experts.\n\nSkills\nData Engineering, Artificial Intelligence, Data Infrastructure, Data Pipeline, ETL, LLM, MLOps, LangChain, Python","description_format":"text","description_chars":4383,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Growth opportunities"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T13:06:44Z"}],"liveness":{"score":33,"band":"fade","label":"Fading","p_open":1,"p_active":0.421,"p_room":0.788,"age_days":6,"expected_fill_days":8,"reasons":["seen:6","agency","velocity","win:late"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/hiringblaze-ai-data-engineer","json_url":"https://alion.io/job/hiringblaze-ai-data-engineer.json","meta":{"generated_at":"2026-10-01T03:00:39Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2436,"day_limit":5000,"remaining_today":2564,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}