824,589open jobs
53,146companies
135,057added this week
Browse all
Salary
≈ $71k – $118k per year (Estimated)
Location
In office (Paris)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Sep 22, 2026.

Overview
Company
Impact
Profile match
The industry moves slow. You don’t have to.

TL;DR Davis is hiring a Founding Data Engineer to build the data behind our foundation model, trained from scratch to generate buildings as geometric graphs. You will own the data end to end, from raw floorplans and synthetic generation to a canonical representation, curation, large-scale pretraining and the ablations that tell us which data actually moves the model. Here the dataset is part of the algorithm.

About Davis

Davis is an AI-native real estate company accelerating early-stage development and architectural design. Today developers coordinate four to five fragmented stakeholders over weeks or months. Soon they will need only one: Davis.

We turn every input that shapes a development decision into decision-ready outputs: investor-grade feasibility studies, investment analysis, and architect-certified designs, delivered in days. Every stage pairs our proprietary AI systems with expert review, so velocity never comes at the cost of reliability.

We closed a $5.5M pre-seed co-led by Heartcore Capital and Balderton Capital, with Yellow, Evantic and Entrepreneur First, alongside angels from the founding teams of Spacemaker, Black Forest Labs, Hugging Face, Supabase, Cleo and Spore Bio. We already work with leading developers and expect to support hundreds of projects over the coming year, deepening our research, our hiring, and our coverage of the development process end to end.

The Role

You will own the data our foundation model learns from, a model we train from scratch to generate buildings as geometric graphs. Part of the corpus comes from real floorplans as images and PDFs that have to become clean, standardized graphs. A large part will be synthetic, procedurally generated building graphs, geometry and rendered floorplans, with controlled variation in style, scan noise, annotations and furniture, each kept with its ground-truth graph automatically.

You will think about the whole loop, from raw and synthetic data to a canonical structured representation, curation and validation, the training dataset, large-scale pretraining, evaluation, data ablations, and back to improving the generator and the data mixture. You are senior enough to architect the data stack and set the data strategy, and hands-on enough to write the pipelines, run the experiments, train the models and do the ablations yourself.

What you will be working on

  • Corpus from raw sources. Turn real floorplans (images, PDFs, scans) into clean, standardized building graphs, with the geometry and semantics that make them trainable.

  • Synthetic data generation. Explore strategies to expand the dataset with synthetic data.

  • Canonical representation and curation. Define the standardized representation, then filter, deduplicate, quality-score and validate at scale, with versioning, provenance and lineage.

  • Pretraining data and mixtures. Assemble the training datasets, design the mixture and the curriculum, and blend synthetic and real data for large-scale pretraining.

  • Data ablations. Train models to learn which data actually helps, read the results, and feed them back into the generator and the mixture.

  • Evaluation. Build the eval harness (datasets, metrics, regression tests, monitoring) that tracks data and model quality over time.

What We Are Looking For

  • You have built the dataset, not just trained on it. You have personally built or generated the data used for a large pretraining run, from raw or synthetic sources, rather than only training on a dataset someone handed you.

  • Senior and deeply hands-on. At least 5 years of strong experience, senior enough to architect the data stack and set strategy, but still coding the pipelines, running the experiments and doing the ablations yourself.

  • Data as a first-class problem. A track record where the data itself is the object: curation, filtering, deduplication, quality scoring, mixtures, synthetic generation.

  • Strong engineering. Deep Python, clean and typed code, async and concurrency, distributed data pipelines, TDD culture.

Nice to have

  • Computer vision and document understanding, images to structured output, OCR, layout extraction, segmentation, geometry extraction, vectorization, raster to vector, image to scene graph, 3D or CAD.

  • Graph and structured scientific data, molecular, protein or scene graphs, meshes, CAD, BIM, 3D geometry, or relational and structured world data.

  • Synthetic worlds and simulation, a structured state to a simulator to a renderer to synthetic images with perfect labels, then a perception model, with an eye on the sim-to-real gap.

  • Foundation models from scratch, real involvement in a large pretraining run, not only fine-tuning.

  • Public evidence of data ownership, lead on a dataset, a Hugging Face release, a dataset card, a technical blog on your pipeline, or a talk on data curation or synthetic data.

  • GIS and geometry, parcels, zoning layers, projections, computational geometry or constrained optimization.

  • Multi-country data, heterogeneous sources, localization and varying rules.

Why Join Davis

  • Own the data end to end, from the generator to the pretraining mixture, as the person who defines what the model learns from.

  • Shape a foundation model from scratch, on a structured representation no one else is training on.

  • Work on genuinely hard problems, image to graph, synthetic worlds and large-scale pretraining, with your work reaching clients within days.

  • Competitive salary and meaningful equity, at a founding level, in an early-stage company.

  • Join a world-class team, a mix of AI researchers, engineers and architects backed by world-class VCs.

More information about Davis, the team and the market we're going after:

Team

Mehdi (Co-founder & CEO) grew up in a family of architects and has lived this problem firsthand. He is a repeat founder who bootstrapped his first startup at 20, and graduated from Sciences Po and HEC Paris. Amine (Co-founder & CTO) is an AI researcher from École Polytechnique who worked extensively on discrete diffusion and turned down a PhD with Google DeepMind to build Davis. They started working together in July 2025 at Entrepreneur First's first European residency, a two-month lock-in in a German castle.

Today we are a team of twelve: technical profiles from Polytechnique, ENS and INRIA alongside architects and deep real estate expertise. We are small with an extremely high bar. If you want to work deeply on hard problems and see your work reach clients within days, you are the one we need.

Why We Will Win

Real estate is a $13 trillion industry that technology has largely bypassed. The professional services that feed it (design, engineering, feasibility, permitting) represent hundreds of billions in spend that no one has seriously automated.

Proptech spent the last decade selling SaaS on the edges of these workflows. It did not work, for two reasons: no professional wants another tool to learn, and no tool can automate work that runs on expert judgment. Davis makes a different bet. We do not sell tools, we sell the work, AI-generated, expert-validated, delivered in days instead of weeks. Every project compounds our data advantage across typologies, geographies and regulatory contexts.

Why No One Has Solved Architectural Design

Real estate development bleeds time and money in architectural design loops. Architects cycle through dozens of floorplan revisions to meet regulatory and client constraints, each round taking days, each missed constraint restarting the loop. Traditional CAD and BIM tools offer zero generative capability, and parametric tools only check constraints after generation, which leaves designs that frequently break under new zoning rules or irregular sites.

Generative AI has the potential to solve this, but does not yet. Fine-tuning image diffusion models on floorplans produces layouts that look plausible but fall apart under scrutiny: hallucinated rooms, mislabeled spaces, code violations no architect would accept. Pixel-space models have no concept of what a wall is or why a corridor needs to connect two things. Compliance-guidance techniques typically require segmentation at each noisy timestep, compounding errors and making major edits impractical. These layouts may look plausible yet break building codes, or need heavy post-processing before they are usable.

This is exactly the gap Davis closes. Because our model generates on structured representations rather than pixels, it produces compliant, editable plans from the start, and the data you build feeds that model directly.

We care about who you are, not just what's on your CV.

If you are drawn to what we are building but do not meet every requirement, we still want to hear from you. Studies show that women in particular tend to apply only when they meet all of the criteria. If that is you, please do not let it hold you back. We would love to receive your application.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Paris
$63k – $79k per year • In office • Full-Time • 4+ years exp • Paris
Python
SQL
Databases
MySQL
PostgreSQL
Redis
Cassandra
DynamoDB
Amazon Aurora
Google Cloud Spanner
Azure SQL Database
DevOps
GCP
CI/CD
AWS
Cybersecurity
GDPR
Apply
Data Engineer 5 months ago
$62k – $77k per year • In office • Full-Time • 3+ years exp • Paris
Python
SQL
Databases
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
Analytics
Tableau
Power BI
ETL/ELT
A/B Testing
Looker
Apply
≈ $137k – $226k per year (Estimated) • Remote (United States) • Public Trust • Full-Time • Bachelor's Degree • Salt Lake City
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Apache Kafka
Microsoft Fabric
AI/ML
Spark
Machine Learning
DevOps
Terraform
Azure
CI/CD
Git
Cybersecurity
HIPAA
NIST 800-53
FedRAMP
Analytics
Power BI
Azure Data Factory
Collibra
Apply
≈ $32k – $88k per year (Estimated) • In office • Full-Time • Hanoi
Python
SQL
Databases
PostgreSQL
Oracle
Apache Kafka
DevOps
CI/CD
Git
AWS
Linux
Cybersecurity
CyberArk
Analytics
ETL/ELT
DataStage
Oracle Data Integrator
Apply
≈ $32k – $86k per year (Estimated) • In office • Full-Time • Hanoi
Python
PowerShell
Bash
Databases
MySQL
PostgreSQL
Oracle
MS SQL
Apache Kafka
DevOps
Terraform
Ansible
CI/CD
Git
AWS
Linux
Apply
In office • Full-Time • 2+ years exp • Noida
Python
Management
Agile
QA
TestRail
Selenium
Appium
Apply
Software Testing Lead 7 hours ago
≈ $23k – $65k per year (Estimated) • In office • Full-Time • Noida
Python
JavaScript
Databases
Milvus
Pinecone
AI/ML
LangChain
MLFlow
Prompt Engineering
Kubeflow
Great Expectations
TensorFlow
LLM
RAG
DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Shift-Left
GitLab
Cybersecurity
Shift-Left Security
Management
Agile
Scrum
QA
Selenium
JMeter
Cypress
Playwright
Postman
Rest-Assured
Locust
Apply
$90k – $220k per year • In office • 2+ years exp • United States
Python
Apply
$490k – $560k per year • In office • 1+ year exp • Seattle
Python
AI/ML
Pandas
Machine Learning
Apply
$200k – $280k per year • In office • Washington
Python
TypeScript
AI/ML
Machine Learning
Apply
≈ $79k – $147k per year (Estimated) • In office • Full-Time • 2+ years exp • PhD • Paris
Python
Databases
Supabase
AI/ML
Fine-tuning
Diffusion Models
NLP
LLM
RAG
Hugging Face
Context Engineering
LLM Evaluation
Multi-Agent Systems
Mobile
Clean Architecture
Apply
≈ $69k – $159k per year (Estimated) • In office • Full-Time • Master's Degree • Paris
Python
Databases
Supabase
AI/ML
PyTorch Lightning
Fine-tuning
RLHF
Reinforcement Learning
Multimodal AI
Diffusion Models
PyTorch
Hugging Face
GRPO
Reward Modeling
Machine Learning
Apply
Residential Architect 10 days ago
≈ $69k – $150k per year (Estimated) • In office • Full-Time • 3+ years exp • PhD • Paris
Databases
Supabase
AI/ML
Hugging Face
Design
AutoCAD
Apply
Generalist Architect 10 days ago
≈ $69k – $150k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Paris
Databases
Supabase
AI/ML
Hugging Face
Design
AutoCAD
Apply
Data Center Architect 11 days ago
≈ $83k – $184k per year (Estimated) • In office • Full-Time • 3+ years exp • PhD • Paris
Databases
Supabase
AI/ML
Hugging Face
Design
AutoCAD
Apply
≈ $31k – $52k per year (Estimated) • In office • Paris
Apply
≈ $48k – $115k per year (Estimated) • Remote (France) • Full-Time • 2+ years exp • Paris
Apply
≈ $60k – $126k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Paris
Apply
≈ $53k – $90k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Paris
Apply
In office • Paris
Databases
Snowflake
Neo4j
Databricks
ArangoDB
AI/ML
AI Agents
Edge AI
DevOps
AWS
Marketing
Salesforce
LinkedIn
Apply
See all jobs
This is one of many
824,589 more open roles from verified company boards, updated every day.