{"id":1198155,"url":"https://alion.io/job/collective-ai-data-engineer-data-platform","title":"AI Data Engineer, Data Platform","company":{"id":171769,"name":"Collective","domain":"collective.com","url":"https://alion.io/company/collective-3","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":{"grade":"A","score":85,"open_postings":5,"ghost_share":0,"stale_share":0.6,"repost_share":0,"time_to_fill_p50_days":45,"computed_at":"2026-10-10T05:45:15Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":180000,"max":230000,"currency":"USD","period":"year","gross":null,"usd_annual":230000},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"BigQuery","optional":false},{"name":"CI/CD","optional":false},{"name":"Dagster","optional":false},{"name":"dbt","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"Fivetran","optional":false},{"name":"Google BigQuery","optional":false},{"name":"LLM","optional":false},{"name":"Metabase","optional":false},{"name":"Python","optional":false},{"name":"SQL","optional":false},{"name":"Apache Kafka","optional":true},{"name":"Claude Code","optional":true},{"name":"Datadog","optional":true},{"name":"GCP","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-24T16:42:45Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-10-11T19:14:25Z","board_verified":true,"closed_at":null,"days_open":17,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":17},"description":"About Collective:\nCollective is on a mission to redefine the way businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business incorporation to accounting, bookkeeping, tax services, and access to a thriving community, all in one integrated platform. We believe in empowering self-employed people to enjoy the same tax savings that big companies get, so they can focus on their passion, not paperwork.\nFeatured in Forbes, Business Insider, Yahoo, Bloomberg, Financial Times, TechCrunch, and more. We are backed by General Catalyst, Sound Ventures, QED Investors, Google’s Gradient Ventures, Expa, and other investors who have financed iconic companies like YouTube, Substack, Twitch, Box, Hims, Instacart, and Lyft.\nAbout the role:\nWe are looking for a Data Engineer to own and scale the data platform that powers analytics, reporting, and AI across Collective. You will design, build, and maintain the pipelines that move data from our product, financial, and third-party systems into our BigQuery warehouse; model that data into clean, well-documented, reliable tables; and set the engineering standards that keep the platform trustworthy as the company grows.\nYou will join the Data Engineering team within Engineering and work closely with product engineers, analysts, and business stakeholders across Operations, Finance, and Go-to-Market. This is a hands-on role for someone who cares about data quality, takes ownership of production systems end-to-end, and wants their work to be the foundation the rest of the company builds on.\nWhat you'll do:\nDesign and build data pipelines. Develop, deploy, and maintain scalable batch and event-driven pipelines that ingest data from application databases, SaaS tools, and external APIs into BigQuery using managed connectors (Fivetran), custom Python loaders, and orchestration tooling.\n\nModel the data. Design and implement dimensional and analytical data models in dbt, following a layered architecture (raw, staging, marts) with clear grain, naming conventions, and documentation that analysts and downstream tools can rely on.\n\nOwn data quality and reliability. Implement testing, monitoring, alerting, and data contracts across the pipeline; define and meet freshness and accuracy SLAs; triage and resolve pipeline failures and data incidents to root cause.\n\nOptimize performance and cost. Tune warehouse queries, partitioning, and clustering; manage BigQuery spend; and keep pipelines efficient as data volume grows.\n\nEstablish engineering standards. Drive best practices for version control, code review, CI/CD, and infrastructure-as-code across the data stack; document systems and runbooks so the platform is maintainable by the team.\n\nGovern and secure data. Implement access controls, PII handling, and data retention practices appropriate for a financial services company; partner with Security and Legal on compliance requirements.\n\nEnable the business. Partner with product engineers on source schema design and change management, and with analysts and stakeholders to translate business questions into reliable datasets, metric definitions, and self-serve reporting in Metabase.\n\nSupport AI and analytics use cases. Maintain the semantic layer, metric definitions, and documentation that allow LLM-based tools and internal agents to query the warehouse accurately and consistently.\n\nWhat you'll bring:\nExperience: 5+ years of professional experience in data engineering, analytics engineering, or a closely related role, ideally at a B2B SaaS or fintech company.\n\nSQL and Python: Expert-level SQL and strong Python skills for building pipelines, transformations, and tooling; comfortable writing tested, production-grade code.\n\nModern data stack: Hands-on production experience with a cloud data warehouse (BigQuery strongly preferred), dbt or equivalent transformation framework, managed ingestion tools (Fivetran or similar), and an orchestrator (Airflow, Dagster, Cloud Composer, or similar).\n\nData modeling: Deep understanding of dimensional modeling, layered warehouse architecture, and schema design, with strong opinions on grain, naming, and consistency.\n\nData quality and observability: Experience implementing testing frameworks, lineage, monitoring, and alerting for data pipelines, and operating them in production including on-call.\n\nEngineering fundamentals: Fluency with git-based workflows, code review, CI/CD, and infrastructure-as-code; you treat data infrastructure as software.\n\nOwnership: A track record of taking ambiguous, high-impact problems and delivering reliable systems end-to-end, with a focus on outcomes rather than just implementation.\n\nCommunication: Ability to explain technical trade-offs to non-technical stakeholders and drive alignment on data definitions across teams.\n\nNice to have:\nExperience with streaming or event data (Pub/Sub, Kafka, or similar) and product analytics tooling (Amplitude or similar).\n\nExperience with Terraform and Google Cloud Platform infrastructure.\n\nExperience with observability platforms such as Datadog.\n\nExposure to financial, accounting, tax, or payroll data and the correctness requirements that come with it.\n\nExperience building semantic layers or metric stores consumed by LLM-based tools, or supporting LLM evaluation programs.\n\nAI-assisted development experience (Claude Code or similar).\n\nWhat we offer:\nHybrid Work Model: Based in San Francisco with a balance of in-office and remote flexibility.\n\nFresh Lunch: Provided on in-office days.\n\nCommuter Support: $150 monthly reimbursement for transit expenses.\n\nHealth & Wellness: $200 quarterly reimbursement to support your well-being.\n\nTime Off: Flexible PTO plus 14 company holidays.\n\nComprehensive Coverage: 100% medical, dental, and vision for employees; 75% coverage for dependents.\n\nParental Leave: 16 weeks fully paid.\n\nRetirement & Ownership: 401k plan plus an equity package.\n\nTeam Connection: Quarterly virtual events and an annual in-person summit.","description_format":"text","description_chars":6043,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["401k plan","Equity","Hybrid work","Parental leave","Retirement plans"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Design & Creative"],"lifecycle":[{"event":"open","at":"2026-09-24T19:30:39Z"}],"visa":[],"liveness":{"score":83,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.825,"p_room":1,"age_days":15,"expected_fill_days":45,"reasons":["conf:0","velocity","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":{"stated_usd_annual":230000,"is_top_pay":true},"html_url":"https://alion.io/job/collective-ai-data-engineer-data-platform","json_url":"https://alion.io/job/collective-ai-data-engineer-data-platform.json","meta":{"generated_at":"2026-10-11T22:30:39Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":12482,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":171769},"rest":"https://alion.io/mcp/rest/get_company?id=171769"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fcollective-ai-data-engineer-data-platform"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fcollective-ai-data-engineer-data-platform"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fcollective-ai-data-engineer-data-platform"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/collective-ai-data-engineer-data-platform\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fcollective-ai-data-engineer-data-platform"}]}