{"id":1266869,"url":"https://alion.io/job/tendios-senior-data-engineer","title":"Senior Data Engineer","company":{"id":690448,"name":"Tendios","domain":"tendios.com","url":"https://alion.io/company/tendios","size_band":"11-50","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Teamtailor","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":["ES"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":91000,"max_usd":203000,"period":"year","method":"global_role_seniority_cell","sample_n":1536},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Apache Kafka","optional":false},{"name":"ClickHouse","optional":false},{"name":"Confluence","optional":false},{"name":"ElasticSearch","optional":false},{"name":"Embeddings","optional":false},{"name":"Function Calling","optional":false},{"name":"Hybrid Search","optional":false},{"name":"LLM","optional":false},{"name":"Metabase","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"Nest.JS","optional":false},{"name":"Node JS","optional":false},{"name":"OpenSearch","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prefect","optional":false},{"name":"Python","optional":false},{"name":"Qdrant","optional":false},{"name":"RabbitMQ","optional":false},{"name":"RAG","optional":false},{"name":"Reranking","optional":false},{"name":"Semantic Search","optional":false},{"name":"Semantic Search","optional":false},{"name":"SQL","optional":false},{"name":"Tool Use","optional":false},{"name":"Turborepo","optional":false},{"name":"TypeScript","optional":false},{"name":"AWS","optional":true},{"name":"dbt","optional":true},{"name":"DigitalOcean","optional":true},{"name":"Docker","optional":true},{"name":"Hetzner","optional":true},{"name":"JavaScript","optional":true},{"name":"Langfuse","optional":true},{"name":"Qwen","optional":true}],"status":"live","first_seen_at":"2026-09-23T11:41:37Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-28T00:17:19Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":5},"description":"Senior Data Engineer Tendios\nRemote (Spain) or hybrid from Barcelona · Full-time\nAbout Tendios\nTendios is the tender intelligence platform for the Spanish public procurement market. Companies use Bid to find, qualify and win public tenders. Public institutions use Create to draft and manage their own. Behind both products is one of the most complete datasets on Spanish public procurement: tenders, lots, CPVs, contracting bodies, resolutions and awards. It’s collected continuously from hundreds of public sources and enriched with AI.\nThat data is our product. We’re looking for the person who will own it end to end.\nThe role\nThis is not a classic data architect role. You won’t just design schemas and hand them off. You’ll own Tendios data from the moment it’s scraped to the moment a customer sees it in Bid, asks Vera about it, or queries it through our MCP server. You’ll write the code, propose the features, and make the calls on how data is modelled, stored and exposed, including how it feeds our AI features.\nYou’ll report directly to the Head of Technology and work closely with product, the AI & Data squad and the backend squads.\nWhat you’ll do\nTurn data into product\nPropose and build data-driven features. Examples: market and competitor intelligence from award data, pricing benchmarks, contracting-body profiles, better tender matching and alerting, and data-quality signals shown to customers.\n\nOwn the data side of our AI features. Design and improve the retrieval layer behind Vera and our Dynamic RAG service: document ingestion and conversion (Docvert), chunking, embeddings, Qdrant indexing, hybrid search with Elasticsearch, and reranking.\n\nUse LLMs where they add real value in the pipeline: structured extraction from tender documents, classification (e.g. CPVs), entity resolution and summarisation. Build evaluation sets to measure quality instead of guessing.\n\nWork with the AI team so Vera and the Tendios MCP server get clean, well-structured, well-documented data and tools.\n\nDefine and measure data and retrieval quality: coverage per source, freshness, parsing accuracy, deduplication and RAG answer quality. Treat it as a product metric, not an afterthought.\n\nOwn the data platform\nTake ownership of the collection pipeline (Crawl Manager, Tenders Discovery, Collector, Tenders and Resolution Parsers, RabbitMQ consumers) and make it more reliable, observable and cheaper to run..\n\nKeep our read models consistent with the system of record: Elasticsearch for search and filtering, and Qdrant for semantic search and RAG.\n\nBuild out the analytics layer on ClickHouse, orchestrated with Prefect and exposed through Metabase. This includes replacing our manual, Excel-based SaaS metrics reporting (MRR, churn, NRR) with a proper warehouse.\n\nWrite and ship code\nWrite production code in Python for the pipeline and data services and in TypeScript/Node.js in our Turborepo backend (NestJS) where data features touch the API.\n\nReview code, set standards for data modelling and migrations, and document decisions in Confluence.\n\nContribute to data governance and security as part of our compliance work (ENS, ISMS), including data ownership, retention, access and auditability.\n\n What we’re looking for\nMust have\n6+ years in software engineering, with at least 3 years focused on data-intensive systems.\n\nStrong Python and solid SQL. Deep, hands-on PostgreSQL experience: modelling, performance and migrations at scale.\n\nComfortable working across the stack when needed, including reading and writing TypeScript/NestJS.\n\nExperience building and running production data pipelines, including event-driven or queue-based architectures (RabbitMQ, Kafka or similar).\n\nExperience with search or analytical stores such as Elasticsearch/OpenSearch or ClickHouse.\n\nA solid, hands-on understanding of how LLM applications work: RAG, embeddings and vector search, chunking strategies, prompt design, tool calling, and how to evaluate retrieval and answer quality. You’ve shipped at least one LLM or RAG feature to production.\n\nA product mindset. You look at a dataset and see features, and you can write a clear proposal and defend it with product and business stakeholders.\n\nFluent Spanish and English\n\n Nice to have\nWeb scraping and document parsing at scale, including PDFs and messy semi-structured sources.\n\nQdrant or other vector databases at scale, hybrid search and reranking.\n\nLLM observability and evaluation tooling (e.g. Langfuse), or running self-hosted open models (e.g. Qwen) on GPU infrastructure.\n\nOrchestration and ELT tooling such as Prefect, Airflow, dlt or dbt.\n\nExperience with large or legacy data migrations (MongoDB to PostgreSQL is a big bonus).\n\nKnowledge of public procurement, open data or regulated environments.\n\nDocker, and experience with cloud and hybrid infrastructure (Hetzner, DigitalOcean, AWS).\n\nWhy join\nYour work is the product. Better data means better tender matching, better answers from Vera and better decisions for our customers, and you’ll see that directly.\n\nReal ownership. You’ll help define the data strategy of a growing SaaS company, not execute someone else’s.\n\nInteresting problems. You’ll work on scraping at scale, entity resolution across public bodies and suppliers, polyglot persistence, and AI on top of a unique domain dataset.\n\nAn engineering team organised in autonomous squads, with a modern stack and a pragmatic culture. Work fully remote from anywhere in Spain, or hybrid from our Barcelona office.","description_format":"text","description_chars":5489,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"Spanish","level":"Advanced (C1)","optional":false},{"language":"English","level":"Advanced (C1)","optional":false}]},"benefits":[],"hiring_locations":[{"name":"Spain","iso":"ES","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Public Finance","Supply Chain"],"lifecycle":[{"event":"open","at":"2026-09-25T22:39:15Z"}],"liveness":{"score":81,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":0.945,"age_days":4,"expected_fill_days":7,"reasons":["conf:5","win:mid"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/tendios-senior-data-engineer","json_url":"https://alion.io/job/tendios-senior-data-engineer.json","meta":{"generated_at":"2026-09-28T22:58:56Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":618,"day_limit":5000,"remaining_today":4382,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}