995,522open jobs
59,374companies
165,411added this week
Browse all
Salary
≈ $113k – $240k per year (Estimated)
Location
Remote (United States, EST hours)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 4, 2026.

Overview
Company
Impact
Profile match
Wynd Labs is revolutionizing AI data exploration with its decentralized approach to web scraping. As the pioneer in data provisioning networks, Wynd Labs offers a more transparent and ethical process for acquiring data compared to its competitors. Their base layer, Grass, is becoming essential infrastructure for decentralized AI and facilitates the acquisition of various types of data, including financial insights, price scraping, audio transcripts, and more. Wynd Labs provides a comprehensive suite of data-driven solutions to meet the evolving needs of businesses and individuals. Whether you are looking to enhance your AI capabilities or gain valuable insights, Wynd Labs has the tools to help you succeed. Take control of your internet experience with Wynd Labs and unlock the power of AI-driven data solutions. Contact Wynd Labs today to learn more about their products and services.

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance.

This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team.

Who You Are:

  • Bachelor’s degree or equivalent work experience

  • Python (advanced) - strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs

  • Web scraping at scale - hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)

  • Distributed data pipelines - experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)

  • Data warehousing - practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses

  • Docker & Kubernetes - containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads

  • Linux & bare-metal ops - comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions

  • CI/CD for data workflows (GitHub Actions, ArgoCD)

  • Writing Scalable API

What You'll Be Doing:

  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.

  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.

  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.

  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity.

  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.

  • Participate in research and development projects to improve the Company’s data products and workflows.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better. We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
995,522 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
In your city
≈ $21k – $63k per year (Estimated) • Remote (likely EAEU) • Internship • Bachelor's Degree • Moscow
Python
SQL
Python
pySpark
AI/ML
Spark
MLFlow
PyTorch
Machine Learning
DevOps
Docker
Kubernetes
Grafana
Apply
Sr. Data Architect 4 months ago
≈ $127k – $252k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Atlanta
Databases
Snowflake
Databricks
Microsoft Fabric
AI/ML
dbt
Embeddings
AI Agents
LLM
LLMOps
Feature Store
GraphRAG
DevOps
Azure
CI/CD
AWS
FinOps
Analytics
Data Vault
Dimensional Modeling
Apply
$228k – $423k per year • Hybrid • Full-Time • 8+ years exp • Master's Degree • San Francisco • New York • Bellevue
Python
SQL
Databases
Snowflake
AI/ML
Spark
Reinforcement Learning
NLP
TensorFlow
Keras
PyTorch
Recommender Systems
Machine Learning
DevOps
GCP
AWS
Apply
≈ $109k – $213k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Atlanta
Python
Databases
Snowflake
Amazon Redshift
Teradata
DevOps
GCP
Azure
AWS
Unix
Analytics
ETL/ELT
Informatica
Apply
≈ $119k – $235k per year (Estimated) • Hybrid • Full-Time • 18+ years exp • Master's Degree • Atlanta
Python
SQL
Databases
Snowflake
AI/ML
Multimodal AI
Function Calling
Computer Vision
AI Agents
NLP
TensorFlow
PyTorch
Time Series Forecasting
Feature Store
Human-in-the-Loop
LLM Guardrails
Recommender Systems
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GCP
CI/CD
AWS
Apply
Engineer I 3 days ago
≈ $34k – $85k per year (Estimated) • In office • Bachelor's Degree • United States
Python
SQL
Visual Basic
DevOps
Windows
Apply
In office • Contractor • Singapore
Python
SQL
Python
pySpark
Databases
Databricks
Amazon Redshift
AI/ML
Spark
dbt
LLM
DevOps
CI/CD
Git
AWS
AWS Lambda
Amazon S3
Analytics
Tableau
ETL/ELT
Apply
≈ $8.5k – $16k per year (Estimated) • In office • Internship • Bachelor's Degree • Chennai
Python
Java
SQL
AI/ML
AutoGen
LangChain
AI Agents
CrewAI
LLM
Tool Use
DevOps
Azure
AWS
Docker
Analytics
ETL/ELT
Apply
In office • Internship • Chennai
Python
Java
Ruby
Ruby
Ruby on Rails
AI/ML
OpenCV
Computer Vision
TensorFlow
NumPy
OCR
Machine Learning
DevOps
Rest API
Apply
≈ $23k – $55k per year (Estimated) • In office • Full-Time • 10+ years exp • Taguig
Python
DevOps
Splunk
Terraform
Ansible
VMWare
CloudFormation
Azure
Git
AWS
Bitbucket
CentOS Stream
Linux
Windows
Unix
VPN
BGP
OSPF
MPLS
SOAP
Management
Jira
ServiceNow
Agile
Scrum
ITIL
Apply
≈ $97k – $195k per year (Estimated) • Equity • In office • Full-Time • United States
Java
Rust
AI/ML
Knowledge Graph
Apply
≈ $75k – $188k per year (Estimated) • Equity • In office • Full-Time • United States
Python
JavaScript
Java
Rust
C++
Node JS
Node JS
Puppeteer
AI/ML
Multimodal AI
NLP
LLM
Knowledge Graph
QA
Playwright
Chrome DevTools
Apply
≈ $136k – $251k per year (Estimated) • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • United States
Python
Go
Databases
Redis
AI/ML
Knowledge Graph
DevOps
CI/CD
Kubernetes
Apply
≈ $135k – $257k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
Python
AI/ML
Knowledge Distillation
NLP
LLM
Time Series Forecasting
OCR
Knowledge Graph
Model Distillation
Machine Learning
Apply
≈ $94k – $190k per year (Estimated) • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • United States
Go
TypeScript
AI/ML
Knowledge Graph
DevOps
Datadog
Management
Slack
QA
Playwright
Apply
See all jobs
This is one of many
995,522 more open roles from verified company boards, updated every day.