{"id":1268282,"url":"https://alion.io/job/spyne-data-engineer","title":"Data Engineer","company":{"id":2082579,"name":"Spyne","domain":"spyne.ai","url":"https://alion.io/company/spyne","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"AI Agents","optional":false},{"name":"Amazon S3","optional":false},{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"AWS Step Functions","optional":false},{"name":"ClickHouse","optional":false},{"name":"ETL/ELT","optional":false},{"name":"IAM","optional":false},{"name":"Metabase","optional":false},{"name":"MySQL","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"SQL","optional":false},{"name":"Computer Vision","optional":true}],"status":"live","first_seen_at":"2026-06-16T08:23:50Z","employer_posted_date":"2026-06-16","last_verified_at":"2026-09-25T22:59:18Z","board_verified":true,"closed_at":null,"days_open":101,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":101},"description":"Who are we?\nWe are Spyne, redefining how cars are marketed and sold with cutting-edge Generative AI. What started as a bold idea-using AI-powered visuals to help dealers sell online faster-has evolved into a full-fledged AI-first automotive retail ecosystem.\nBacked by $16 M in Series A funding from Vertex Ventures, Accel, and other top investors, we're scaling fast:\nExpanded across the US & EU markets\nLaunched industry-first AI-powered Image & 360° solutions\nAchieved a 5× revenue surge in 15 months, aiming for 3-4× growth this year\nKnow Our Journey\n2020: Launched as a visual merchandising platform\n2023: Pivoted to AI-driven automotive retail solutions\n2024: Achieved 5× revenue growth in 15 months, aiming for 3-4× more\nToday: Driving the GenAI revolution with AI-powered sourcing, pricing, CRM, and Agentic AI for dealerships\nRead more about us:\nStudio AI Product\nVini AI Product\nSeries A Announcement on YourStory\nSeries A Coverage in Economic Times\nAutocar Pro News\nWhat Are We Looking For?\nWe're seeking a highly skilled Data Engineer to establish and own Spyne's dedicated data engineering function-our first. As Spyne AI accelerates market penetration across US rooftops and processes massive, high-velocity streams of unstructured computer vision payloads (Studio AI) and conversational state events (Vini AI), our data warehousing volume and complexity have scaled exponentially.\nThis is not a maintenance role. You will actively restructure our entire data platform-taking absolute ownership from our DevOps team and raising our data infrastructure to true tech-industry standards. You'll build the foundational data layer that powers our BI platforms, ML observability, and long-term analytics roadmap.\nLocation: Gurugram (Work from Office, 5 days a week)\nRole: Full-Time, Data Engineer\nWhat Will You Do?\nData Warehousing Architecture & Modeling: Design, scale, and own our core enterprise Data Warehouse built on ClickHouse Cloud-implementing robust data modeling methodologies, efficient time-based partitioning, and the centralized foundational data layer that powers all downstream BI platforms.\nCDC & Schema Evolution: Spearhead our transition from self-hosted Debezium/Kafka to managed ClickPipes; design performant ELT pipelines using ClickHouse Materialized Views, JSONExtract, and arrayJoin functions to parse deep, complex MongoDB Atlas JSON arrays into clean, flattened analytical tables.\n Advanced ClickHouse Engine Tuning: Manage SharedReplacingMergeTree tables partitioned by time; handle complex edge cases including cross-partition physical deletions (MongoDB tombstone events) and eliminate Cartesian explosions during array joins and LEFT JOIN operations.\n Event Sourcing for ML Pipelines: Maintain and optimize our append-only observability architecture (SQS → Lambda → ClickHouse Async Inserts) to track GPU and CPU ML workloads orchestrated via AWS Step Functions and AWS Batch, leveraging AggregatingMergeTree and anyLast state combinators to unify partial state updates.\nPerformance Optimization & OOM Prevention: Troubleshoot and optimize heavy analytical queries to prevent concurrent Out-Of-Memory crashes when Metabase dashboards fire heavy models simultaneously; aggressively push down filters and leverage GLOBAL IN hash-lookups to eliminate broadcast overhead.\nAWS Data Networking: Navigate secure cross-account data transit within public-internet-denied cloud perimeters using AWS PrivateLink, VPC Lattice Service Networks, and MSK Multi-VPC connectivity secured via IAM authentication.\nData Platform Ownership: Define and drive our long-term data warehousing roadmap; document architecture decisions, establish data quality standards, and reduce bandwidth currently falling on the DevOps team.\nCollaboration: Partner closely with ML Engineering, Product, and DevOps teams to ensure data pipelines are reliable, observable, and aligned with evolving product requirements.\nWhat You Must Have?\nExperience: 3-5 years in a dedicated data engineering role, with proven ownership of production-grade data warehouse or analytics infrastructure.\n ClickHouse: Deep, hands-on expertise with ClickHouse-including engine selection (ReplacingMergeTree, AggregatingMergeTree), Materialized Views, partitioning strategies, and query optimization. ClickHouse Cloud experience is strongly preferred.\n ELT & CDC Pipelines: Demonstrated experience designing and operating Change Data Capture pipelines using Debezium, Kafka, or managed equivalents (ClickPipes, AWS DMS); strong command of schema evolution and data transformation patterns.\n Complex Data Modeling: Proficiency in parsing and flattening deeply nested JSON structures (JSONExtract, arrayJoin); experience modeling data from NoSQL sources (MongoDB Atlas) into analytical schemas.\nEvent-Driven & Streaming Architectures: Hands-on experience with event sourcing patterns, async insert architectures, and streaming systems such as Apache Kafka / AWS MSK and AWS SQS/Lambda.\nAWS Infrastructure: Strong working knowledge of AWS data services-S3, Lambda, Step Functions, Batch, MSK-and AWS networking constructs including PrivateLink, VPC, and IAM-based authentication.\nPerformance Debugging: Proven ability to diagnose and resolve OOM errors, slow analytical queries, and pipeline bottlenecks at scale.\nScripting & Automation: Proficient in Python and SQL for pipeline development, data validation, and operational tooling.\n Multi-Database Expertise (Strong Plus): Architecture experience and performance tuning across MySQL, PostgreSQL, MongoDB, and Kafka-comfort navigating polyglot data environments is highly valued.\nEducation: Bachelor's or master's degree in Computer Science, Data Engineering, or a related field.\nWhy is Spyne an Employee-Centric Company?\nComprehensive Health & Life Coverage - GMC, GPA, and GTLI benefits for you and your family\nPerformance-Driven Growth - Fast career progression, ownership from Day 1, and stock options for top performers\nElevate Learning & Development - Access LinkedIn Learning, mentorship programs, and hands-on AI data projects to upskill daily\nCollaborative Office Culture - Thrive in our energetic, innovation-first workplace\nWhy Spyne?\nStrong Culture: A supportive, transparent environment with high autonomy\nCompetitive Compensation: Market-leading salary, equity, and benefits\nDynamic Growth: Join us at a pivotal growth stage-be the architect of our entire data platform, not just a contributor\n Cutting-Edge Tech: Work with ClickHouse Cloud, distributed ML pipelines, and real-time automotive AI data at a scale very few engineers encounter","description_format":"text","description_chars":6583,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Equity","Stock options"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Conversational AI","E-learning"],"lifecycle":[{"event":"open","at":"2026-09-25T22:59:18Z"}],"liveness":{"score":15,"band":"cold","label":"Long shot","p_open":1,"p_active":0.526,"p_room":0.28,"age_days":101,"expected_fill_days":45,"reasons":["conf:4","win:tail","crowd:"],"computed_at":"2026-09-26T03:59:12Z"},"pay":null,"html_url":"https://alion.io/job/spyne-data-engineer","json_url":"https://alion.io/job/spyne-data-engineer.json","meta":{"generated_at":"2026-09-26T03:59:12Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4307,"day_limit":5000,"remaining_today":693,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}