{"id":1901873,"url":"https://alion.io/job/civic-marketplace-data-engineer","title":"Data Engineer","company":{"id":674738,"name":"Civic Marketplace","domain":"civicmarketplace.com","url":"https://alion.io/company/civicmarketplace","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":null,"employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":["London, United Kingdom"],"countries":["GB"],"hiring_countries":["US","GB"],"hiring_countries_total":2,"salary":null,"salary_estimate":{"min_usd":72000,"max_usd":166000,"period":"year","method":"role_country_seniority_unknown","sample_n":116},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agentic Workflows","optional":false},{"name":"AI Agents","optional":false},{"name":"Claude","optional":false},{"name":"Copilot","optional":false},{"name":"ElasticSearch","optional":false},{"name":"LLM","optional":false},{"name":"LLM Evaluation","optional":false},{"name":"Milvus","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"OpenSearch","optional":false},{"name":"Pinecone","optional":false},{"name":"Python","optional":false},{"name":"RAG","optional":false},{"name":"SQL","optional":false},{"name":"dbt","optional":true},{"name":"Fivetran","optional":true},{"name":"Metabase","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Snowflake","optional":true}],"status":"live","first_seen_at":"2026-10-05T08:53:25Z","employer_posted_date":"2026-10-05","last_verified_at":"2026-10-09T00:55:43Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Why this role exists\nEvery city, county, and school district buys things. Roads, software, cleaning services, IT infrastructure. The total is somewhere north of two trillion dollars a year. And almost all of it moves through procurement processes designed for a different era: slow; paper-heavy; opaque and exhausting for everyone involved.\nCivic Marketplace was built to fix that. We're a modern, data-driven platform where government agencies discover, evaluate, and engage suppliers. Where businesses, especially smaller and growing ones, can actually find and win public sector work without needing a dedicated contracts team to navigate the maze.\nWe're past the point of proving this works. Agencies are live on the platform, real money moves through it, and we're now combining that marketplace infrastructure with AI in ways that could genuinely transform how procurement works. Not just incrementally but structurally.\nNone of it works without data that can be trusted. Procurement is a data problem before it is a software problem: who the suppliers actually are, what they can genuinely deliver, which contracts an agency is already entitled to buy from, what this thing cost the county next door. That information exists, scattered across thousands of agency portals, state registries and PDF attachments, in no agreed format, and nobody has assembled it properly. Whoever does gets to define how public money is spent for the next decade. That's this role.\nThe problem you'd own\nThe hard part isn't moving data from one place to another. Any engineer here will tell you the pipelines are the easy half.\nThe hard part is that public procurement has no shared vocabulary. The same supplier turns up as four different legal entities across three registries, and two of the spellings are wrong. Commodity codes are applied inconsistently, or not at all. A cooperative contract one agency can buy from today is invisible to the agency next door because it was published as a PDF on a portal with no API. Every source you touch is incomplete, inconsistent, or both, and most of it is public record, so you can't quietly correct it. You have to model the mess honestly.\nAnd the bar just moved. We recently launched an MCP integration that lets agencies request quotes through Claude, GPT and Copilot. When an agent answers a procurement question, a wrong answer doesn't read like a bug, it reads like advice, and in public spending, bad advice ends up in a council meeting. That puts weight on freshness, lineage and provenance that most product data layers never carry. As we build out our agentic procurement capabilities, the reliability of that data, and how we prepare it for retrieval, becomes our most critical engineering challenge.\nSo this is an entity resolution and data trust problem dressed as a pipeline problem, and it's yours to solve. You'd work out what is actually wrong with the data, which is rarely what everyone assumes, then build the fix so it holds for every source we add next. Not a script per source held together by whoever wrote it.\nAbout the role\nYou would be our first dedicated data engineering hire, reporting to Mikey Mo, our Head of Engineering. Data work today sits with the product engineering team and gets done alongside shipping features. It works, but it belongs to whoever last touched it, and that isn't a foundation for what comes next.\nTo give you a sense of the gap: awarded quotes that slip through the platform get manually caught by our commercial team and hand-entered into HubSpot instead. Supplier onboarding data follows the same pattern. So as and when the platform and HubSpot become out of sync, it makes it difficult to say with confidence which one is right.\nThis is a build-and-do role. You'd collaborate closely with our Head of Engineering to shape the architecture, and you'd also write the pipelines, carry the pager for them, and go and read the raw source when a number looks wrong. If you're looking for a role where you set direction in isolation and hand implementation to other people, this isn't it. If you want substantial responsibility and scope as a function lead alongside our Head of Engineering while still enjoying the craft yourself, it very much is.\nYou would not be starting from nothing. There's a live platform with real transactions running through it, a product engineering team who know where the bodies are buried, and access most people in this field would have to scrape for: Civic Marketplace is a member of the NIGP Business Council, with real partnerships across councils of governments. Most people doing this job spend their first year getting the data access. You'd start with the door already open.\nWhat you'd own\nFive things, and the freedom to decide how.\nTrust in the data. The north star. A golden, trusted dataset that powers Civic Marketplace's analytics, so when someone asks how many quotes were awarded last quarter, or how much GMV we've captured, there's one number, and it's right. You'd own the diagnosis and the fix.\n\nIngestion. Pipelines that pull supplier, agency, solicitation and contract data out of sources never designed to be read by anything but a person. Reliable, observable, and cheap enough to add the next source without a meeting about whether it's worth it.\n\nThe canonical model. One supplier, one record, across every spelling, trading name, subsidiary and registry identifier. One shared definition of a contract vehicle, a commodity, an agency, that product, sales and customer success all use and none of them argue with. This is the unglamorous part that makes everything downstream possible.\n\nThe system underneath. Testing, lineage and freshness, so when the platform tells an agency something we know where it came from and when. And self-serve data for product, customer success and analytics, so nobody queues behind an engineer for a number. Everything you build should work for the next source without you in the loop. This is where AI earns its place, turning what is currently hand-checked into something repeatable.\n\nAgentic data infrastructure. As we scale our use of AI-native procurement, you will own the data layer that powers our agentic workflows. This goes beyond standard retrieval. You will design ingestion and storage strategies, including optimizing vector pipelines like Pinecone for RAG, to ensure our agents have the high-fidelity, context-aware data necessary to perform reliably in production.\n\nYou'd work closely with product engineering, customer success and our Growth Lead. Application feature work stays with the product engineering team, so you're not competing for that ground. You own the layer they build on top of.\nHow we use AI\nWe use Claude daily, across content pipelines, community response and internal knowledge, and increasingly inside the product itself. That's real, not a line in a job ad.\nFor this role it cuts two ways. There's how you work: source profiling, schema mapping, test generation, the documentation that otherwise never gets written. Most of that is now automatable, and we expect you to automate it. And there's what you build: the data layer that decides whether agentic procurement is trustworthy or embarrassing. The second is the harder and more interesting problem.\nWhat we care about is not whether you can use it. Everyone says they can. It's whether you reach for it to make something repeatable. The difference between someone who writes pipelines faster and someone who builds the tooling that means the next hundred sources don't each need hand-holding is the difference we're hiring for.\nIf you've built data quality tooling, evaluation harnesses, retrieval pipelines or automated documentation from scratch, tell us about them. If AI has changed how you think about the work and not just how fast you produce it, we especially want to hear that.\nOur stack: Snowflake, Postgres, dbt, Fivetran, HubSpot, PostHog, Metabase.\nWhat you bring\nYou'll need:\nData engineering experience in a startup or scale-up, where you've built the platform rather than inherited one\n\nStrong SQL and Python, and the judgement to know when the warehouse is the wrong place to solve a problem\n\nA data model you designed that other people had to live with, including the parts you would do differently now\n\nReal experience of messy external sources: half-documented APIs, flat-file drops, scraped pages, PDFs, and data you neither control nor can correct\n\nEntity resolution or record linkage experience, or clear evidence you'd be good at it. This is the centre of the job, not an edge case\n\nA commercial head. You can tell the difference between a data problem that's blocking revenue and one that's merely interesting, and you sequence accordingly\n\nThe instinct to find out why a number is wrong before proposing a fix. Sometimes it's the pipeline, sometimes the model, sometimes the source was always like that, and those need different answers\n\nThe instinct to build systems that scale rather than pipelines that run once\n\nGenuine AI fluency, in the sense above\n\nThe determination to stay positive through the ups and downs of a fast-moving startup, and to find a path through problems rather than wait for conditions to be right\n\nComfort operating remotely across timezones with real autonomy and not much oversight\n\nCuriosity about why public sector data is different. You don't need to have done it. You do need to want to understand it\n\nYou'll stand out if you have:\nGovtech, public sector or civic tech background, or experience working with public records and open data\n\nFamiliarity with supplier and procurement data: SAM.gov and UEI, DUNS, NAICS or UNSPSC, cooperative purchasing, COG or NIGP\n\nExperience with search and retrieval, whether OpenSearch, Elasticsearch or vector pipelines (like Pinecone or Milvus) feeding an LLM product.\n\nBilingual English and Spanish, especially valuable for supplier data quality and onboarding\n\nExperience being the first data hire on a team that had been doing it themselves\n\nWhy you'll love this role\nThe problem is genuinely important. Public procurement touches every road, school and hospital. Getting it right matters in ways most B2B SaaS doesn't. And you'd be widening access for small and local businesses, which is the part of this that keeps us up at night in a good way.\nThe dataset doesn't exist yet. Most data engineering jobs are being the fifth person to tidy the same warehouse. This one is assembling something nobody has assembled properly: a clean, current picture of who supplies the public sector and what public money actually buys.\nThe scope is unusual. First dedicated data hire, reporting to Mikey, taking substantial responsibility for the function and co-designing the architecture decisions that come with it.\nYou'd own a real piece of it. Equity is part of the package here, and we mean it as more than a line in an offer letter. We're looking for someone who wants to build something they have a stake in, not just a job with a good salary attached.\nThe AI opportunity is real. We're not bolting AI onto existing workflows. We're rethinking what's possible when procurement becomes genuinely agentic, and the data layer is the part that decides whether any of it can be trusted.\nPeople who care. Small team, direct access to founders, genuine investment in doing things well. We move fast, but we don't move sloppy.\nHow we work\nBuild bridges to help customers win. We are obsessively focused on helping both agencies and suppliers succeed. That's not a value statement. It's the job.\nHigh velocity, high ownership. We move quickly, make decisions, and take responsibility for outcomes. There's no one to hand things off to and no one to hide behind.\nIn the arena. We stay close to our users. We learn from them directly. We build based on what we see, not what we assume.\nLearning quotient. Rapid iteration isn't just a process. It's a mindset. We'd rather be wrong fast and right eventually than slow and cautious throughout.\nAnd bec...","description_format":"text","description_chars":13289,"description_truncated":true,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"Spanish","level":"Proficiency (C2)","optional":false},{"language":"English","level":"Proficiency (C2)","optional":false}]},"benefits":["Equity","Vision insurance"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"},{"name":"United Kingdom","iso":"GB","kind":"country"},{"name":"London","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Commerce","Marketplaces","Public Finance","Supply Chain"],"lifecycle":[{"event":"open","at":"2026-10-05T11:27:48Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":2,"expected_fill_days":41,"reasons":["conf:1","win:early"],"computed_at":"2026-10-08T05:49:30Z"},"pay":null,"html_url":"https://alion.io/job/civic-marketplace-data-engineer","json_url":"https://alion.io/job/civic-marketplace-data-engineer.json","meta":{"generated_at":"2026-10-09T02:39:30Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":948,"day_limit":5000,"remaining_today":4052,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":674738},"rest":"https://alion.io/mcp/rest/get_company?id=674738"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fcivic-marketplace-data-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fcivic-marketplace-data-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fcivic-marketplace-data-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/civic-marketplace-data-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fcivic-marketplace-data-engineer"}]}