952,274open jobs
57,593companies
155,586added this week
Browse all
Salary
≈ $26k – $56k per year (Estimated)
Location
Hybrid (India)
Seniority
Staff · 12+ years exp

First seen by Alion on Aug 13, 2026.

Overview
Company
Impact
Profile match
Merit Group uses proprietary technology driven by AI to provide our clients with data-driven business intelligence solutions.

Experience : 12 - 18 Years


Location : PAN INDIA


Work Mode : Permanent Remote


Job Type : Full Time Employment


Notice : Immediate Joiners


Job Description :


Primary (70%) : Web scraping architecture anti-bot, distributed crawling, proxy/CAPTCHA strategy, compliance.


Secondary (30%) : Data engineering pipelines, lakehouse, orchestration, dbt.


Mandatory ownership of the technical solution and effort estimation for every new scraping proposal/RFP.


Senior profile : 9+ years overall, 5+ years in scraping, 4+ years in data engineering.


Technical Lead / Architect Web Scraping & Data Engineering :


Role Summary :


We are seeking a Technical lead for the design and delivery of large-scale web scraping and data extraction solutions, with strong supporting expertise in data engineering, the technical authority on all scraping initiatives from pre-sales solutioning and effort estimation through to architecture, build, and stabilization.


The primary mandate is web scraping : designing resilient crawlers, anti-bot strategies, distributed extraction systems, and compliance frameworks. The secondary mandate is data engineering : ensuring extracted data flows reliably into well-modelled, query-ready storage layers and downstream analytical or operational systems.


Mandatory Involvement Areas :


- Author or co-author the technical solution section of every new RFP / RFI / proposal document for scraping projects.


- Lead technical discovery calls with prospective clients to understand target sources, data SLAs, volume, and compliance constraints.


- Conduct target-site feasibility assessments anti-bot complexity, dynamic content, login walls, geo-restrictions, rate limits and document findings before commercials are committed.


- Own end-to-end effort estimation for scraping engagements : crawler build effort, infrastructure sizing, proxy and CAPTCHA cost projections, maintenance overhead, and contingency.


- Produce solution architecture diagrams, tech stack recommendations, and assumption logs as part of every proposal.


- Define and document SLAs, KPIs, and acceptance criteria proposed to the client.


- Participate in client orals, technical defence sessions, and commercial negotiations as the technical SPOC.


- Maintain an internal estimation knowledge base reusable estimation templates, complexity matrices, target-site classification, and historical actuals and continuously refine it after every project closure.


- Sign off on the technical feasibility and risk profile of every proposal before submission. No scraping proposal goes out without architect approval.


Core Responsibilities :


- Design end-to-end scraping solutions covering crawl orchestration, extraction, parsing, storage, and downstream data consumption.


- Define reference architectures for high-volume, high-velocity, and high-variety scraping use cases.


- Architect resilient systems handling JavaScript-heavy sites, CAPTCHAs, rate limiting, IP blocking, and frequent DOM changes.


- Evaluate and select frameworks, proxy networks, and anti-bot bypass strategies aligned with cost, performance, and compliance.


- Design data quality, deduplication, validation, and schema-evolution strategies for scraped data.


Data Engineering Architecture (Secondary) :


- Design downstream data pipelines that ingest, transform, and serve scraped data to analytics, ML, or operational consumers.


- Architect lakehouse / warehouse layers, define data modelling standards (dimensional, Data Vault, or hybrid), and govern schema evolution.


- Define ELT/ETL patterns, orchestration strategy, and SLA-backed data freshness commitments.


- Establish data quality, observability, lineage, and cataloguing practices across the platform.


- Recommend storage formats, partitioning, and indexing strategies for cost and query performance.


Delivery Leadership :


- Translate business requirements into technical specifications, sprint plans, and implementation roadmaps.


- Produce HLD, LLD, and Architecture Decision Records (ADRs) for each engagement.


- Provide hands-on guidance, perform code reviews, and mentor scraping and data engineers.


- Drive proof-of-concept (PoC) builds for complex or high-risk targets before full-scale rollout.


Compliance, Risk & Governance :


- Ensure all scraping work adheres to applicable laws and platform terms (GDPR, CCPA, DPDP, robots.txt, copyright).


- Define and enforce ethical scraping practices, request throttling, and PII handling guidelines.


- Conduct risk assessments for each target source and recommend mitigations.


Performance, Cost & Operations :


- Establish SLAs for crawl freshness, completeness, and accuracy.


- Design monitoring, alerting, and self-healing mechanisms for scraping pipelines.


- Optimise infrastructure cost (compute, proxies, storage) without compromising delivery KPIs.


Required Technical Skills Primary (Web Scraping) :


These are non-negotiable. The candidate must demonstrate deep, hands-on expertise in each of the following :


Scraping Frameworks & Tooling :


- Expert-level Python : Scrapy, BeautifulSoup, lxml, Requests, httpx, parsel.


- Headless browser automation : Playwright, Puppeteer, Selenium, Pyppeteer.


- Node.js scraping stack (where applicable) : Puppeteer, Cheerio, Crawlee.


Anti-Bot & Evasion Strategy :


- Proven experience bypassing Cloudflare, Akamai Bot Manager, DataDome, PerimeterX, Imperva, Kasada.


- CAPTCHA handling : reCAPTCHA v2/v3, hCaptcha, FunCaptcha, image / audio solvers; integration with 2Captcha, Anti-Captcha, CapSolver.


- TLS / JA3 / JA4 fingerprinting awareness; HTTP/2 fingerprint evasion; browser fingerprint spoofing.


- Stealth plugins, user-agent rotation, header normalisation, cookie / session management at scale.


Proxy & Network Infrastructure :


- Hands-on experience with rotating, residential, mobile, ISP, and datacenter proxies.


- Integration with providers : Bright Data, Oxylabs, Smartproxy, NetNut, IPRoyal, SOAX.


- Proxy pool design, health-checking, geo-targeting, sticky sessions, and cost optimisation.


Parsing & Extraction :


- XPath, CSS selectors, regex, JSON-LD, microdata, RDFa.


- Reverse-engineering of internal / mobile APIs, GraphQL endpoints, and XHR traffic.


- ML/LLM-assisted extraction for unstructured layouts (nice-to-have, increasingly expected).


Distributed Crawling & Orchestration :


- Scrapy-Redis, Scrapy Cluster, Frontera, Crawlee, or equivalent distributed crawling frameworks.


- Job scheduling and orchestration with Apache Airflow, Prefect, Dagster, or Celery.


- Queue-based architectures using Kafka, RabbitMQ, AWS SQS, GCP Pub/Sub.


Required Technical Skills Secondary (Data Engineering) :


Data Pipelines & Orchestration :


- ETL / ELT design patterns, idempotent pipelines, CDC (Change Data Capture), incremental loads.


- Apache Airflow, Prefect, Dagster, AWS Glue, Azure Data Factory, GCP Dataflow.


- Stream processing : Kafka Streams, Apache Flink, Spark Structured Streaming.


- Batch processing : Apache Spark (PySpark), Databricks, EMR, Dataproc.


Data Modelling & Storage :


- Dimensional modelling (Kimball), Data Vault 2.0, normalised vs. denormalised trade-offs.


- Data warehouses : Snowflake, BigQuery, Redshift, Synapse, Databricks SQL Warehouse.


- Data lakes / lakehouses : Delta Lake, Apache Iceberg, Apache Hudi on S3 / GCS / ADLS.


- OLTP databases : PostgreSQL, MySQL; NoSQL : MongoDB, DynamoDB, Cassandra, Elasticsearch, Redis.


- File formats : Parquet, Avro, ORC, JSON, CSV; partitioning, bucketing, compaction strategies.


Transformation & Quality :


- dbt (data build tool) for transformation, testing, and documentation.


- Data quality frameworks : Great Expectations, Soda, Deequ, custom validators.


- Data lineage and cataloguing : DataHub, OpenMetadata, Amundsen, Atlan, Collibra.


Cloud, DevOps & Observability :


- Strong on at least one of AWS, GCP, Azure (compute, storage, IAM, networking, serverless).


- Containerisation and orchestration : Docker, Kubernetes (EKS / GKE / AKS), ECS.


- Infrastructure as Code : Terraform, Pulumi, CloudFormation.


- CI/CD : GitHub Actions, GitLab CI, Jenkins, Argo CD.


- Observability : Prometheus, Grafana, ELK / OpenSearch, Datadog, Sentry, OpenTelemetry.


Experience & Qualifications :


- Bachelor's or Master's degree in Computer Science, Engineering, or related field.


- 10+ years of overall software engineering experience.


- 5+ years dedicated to large-scale web scraping / data extraction (primary).


- 3+ years of hands-on data engineering experience covering pipelines, warehouses / lakehouses, and orchestration (secondary).


- Proven track record of architecting scraping platforms processing millions of pages per day across diverse target sites.


- Prior experience as Solution Architect, Tech Lead, or Principal Engineer leading teams of 5+ engineers.


- Demonstrated experience in pre-sales / proposal authoring / RFP responses for scraping or data engineering engagements.


- Client-facing consulting experience strongly preferred.


Soft Skills :


- Excellent written and verbal communication; able to defend technical proposals to CXO-level audiences.


- Strong commercial acumen understands the cost / risk / quality trade-offs in estimation.


- Analytical, structured problem-solving mindset.


- Ownership-driven, comfortable being the single point of technical accountability.


- High standards for documentation, knowledge transfer, and reusability.


Expected Deliverables (RFP Scope) :


- Technical solution sections of all proposal documents.


- Target-site feasibility and risk assessment reports.


- Effort estimation models, assumption logs, and complexity matrices.


- Solution architecture diagrams and tech stack recommendations for proposals.


Delivery Phase :


- High-Level Design (HLD) and Low-Level Design (LLD) documents per initiative.


- Architecture Decision Records (ADRs).


- Reference implementations / PoCs for complex extraction scenarios.


- Code review records and engineering standards documentation.


- Operational runbooks, monitoring playbooks, and incident response procedures.


- Compliance and data governance documentation.


- Knowledge transfer sessions and final handover documentation.


Nice-to-Have :


- ML / NLP-based extraction, entity resolution, or LLM-assisted parsing experience.


- Exposure to GraphQL, gRPC, mobile API reverse-engineering, mitmproxy / Charles workflows.


- Domain experience : e-commerce price intelligence, market research, financial data, real estate, travel aggregation, hospitality.


- Open-source contributions to scraping or data engineering ecosystems.


- Familiarity with cross-jurisdictional legal frameworks for automated data collection.


- Experience with reverse ETL tools (Hightouch, Census) and feature stores.

Skills

Scrapy, Web Scraping, Data Engineering, Web Crawling, Puppeteer, Data Pipeline, Apache Airflow, Data Modeling, AWS, Snowflake DB

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
952,274 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
India
In office • Full-Time • Bengaluru
C#
C++
C++
DirectX
Mobile
Dependency Injection
DevOps
Bitbucket
Windows
Management
Scrum
Apply
Software Developer 10 hours ago
In office • Full-Time • Bengaluru
Python
C
Perl
C
Embedded C
DevOps
TCP/IP
Apply
≈ $51k – $110k per year (Estimated) • Hybrid • Full-Time
JavaScript
PHP
TypeScript
PHP
Laminas
Laravel
Databases
MySQL
MariaDB
ElasticSearch
AI/ML
Copilot
Claude
Frontend
Vue.js
DevOps
GitHub Actions
Logstash
Git
Docker
Linux
Management
Agile
Apply
≈ $129k – $238k per year (Estimated) • Hybrid • 7+ years exp • Bachelor's Degree
Python
Go
JavaScript
PowerShell
C#
Node JS
C#
.NET
Node JS
Commander.js
AI/ML
Copilot
Claude
ChatGPT
DevOps
Terraform
Ansible
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
AWS
Windows
Cybersecurity
Zero Trust
Least Privilege
Cryptography
Vault
Apply
Hybrid • 5+ years exp
JavaScript
Java
TypeScript
Java
Spring Boot
Databases
MySQL
PostgreSQL
AI/ML
Copilot
Cursor
Claude Code
LLM
RAG
Frontend
React.js
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Apply
$197k – $225k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • New York • McLean • Cambridge • Richmond
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Management
Agile
Apply
$179k – $205k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • San Jose • McLean • Cambridge • San Francisco • New York
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Management
Agile
Apply
Solution Architect 1 month ago
≈ $29k – $51k per year (Estimated) • Hybrid • 10+ years exp • India
Python
JavaScript
SQL
Node JS
Python
Scrapy
Beautiful Soup
Celery
pySpark
HTTPX
Node JS
Puppeteer
Cheerio
Databases
MySQL
PostgreSQL
Redis
Snowflake
Databricks
Apache Iceberg
Delta Lake
Cassandra
DynamoDB
RabbitMQ
ElasticSearch
Apache Kafka
OpenSearch
Google BigQuery
Amazon Redshift
Apache Hudi
BigQuery
AI/ML
Spark
Airflow
Dagster
dbt
Prefect
NLP
Great Expectations
Flink
LLM
Feature Store
Machine Learning
Frontend
GraphQL
DevOps
gRPC
Terraform
GCP
GitHub Actions
OpenTelemetry
CloudFormation
Datadog
Prometheus
Pulumi
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Cloudflare
Grafana
Self-Healing
Amazon EKS
Google GKE
Azure AKS
Akamai
SLI/SLO/SLA
Amazon S3
IAM
Amazon ECS
Apache HTTP Server
Cybersecurity
GDPR
Analytics
ETL/ELT
Azure Data Factory
AWS Glue
Data Vault
Dimensional Modeling
Collibra
QA
Selenium
Playwright
Sentry
Apply
≈ $17k – $37k per year (Estimated) • In office • 3+ years exp • India
SQL
C#
Perl
Databases
Databricks
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Kubernetes
Google GKE
Azure AKS
GitHub
VPN
Analytics
Azure Data Factory
Apply
≈ $20k – $59k per year (Estimated) • Remote (India) • Full-Time • 10+ years exp • India
Python
Go
Assembly
Python
tox
Assembly
Keystone Engine
DevOps
CI/CD
Kubernetes
Gerrit
OpenStack
Linux
Unix
DHCP
QA
Pytest
Apply
In office • 2+ years exp • Bachelor's Degree • India
Java
SQL
Databases
Apache Kafka
AI/ML
Copilot
Cursor
Claude Code
Function Calling
AI Agents
LLM
Agentic Workflows
Tool Use
DevOps
Rest API
AWS
API Gateway
Apply
≈ $16k – $41k per year (Estimated) • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • India
Python
JavaScript
TypeScript
Databases
MS SQL
Apache Kafka
AI/ML
Copilot
Windsurf
Unstructured.io
RAG
Devin
Frontend
Angular
DevOps
Rest API
Terraform
Ansible
Azure
CI/CD
Jenkins
Git
Docker
Kubernetes
TeamCity
Octopus Deploy
Management
Agile
Apply
≈ $15k – $40k per year (Estimated) • Remote (India) • 3+ years exp • India
Python
JavaScript
Node JS
Databases
PostgreSQL
Frontend
GraphQL
React.js
DevOps
CI/CD
AWS
Apply
Technical Lead 1 day ago
≈ $26k – $57k per year (Estimated) • In office • 6+ years exp • Bengaluru
Python
JavaScript
Java
C#
Node JS
C#
.NET
DevOps
Rest API
CI/CD
Git
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon ECS
API Gateway
Management
Agile
Apply
See all jobs
This is one of many
952,274 more open roles from verified company boards, updated every day.