986,892open jobs
58,977companies
161,835added this week
Browse all
Salary
≈ $29k – $51k per year (Estimated)
Location
Hybrid (India)
Seniority
Staff · 10+ years exp

First seen by Alion on Aug 13, 2026.

Overview
Company
Impact
Profile match
Merit Group uses proprietary technology driven by AI to provide our clients with data-driven business intelligence solutions.

Job Description :

Solution Architect - Web Scraping & Data Engineering

RFP Reference Document

Role Title : Solution Architect - Web Scraping (Primary) & Data Engineering (Secondary)

Engagement Type : Contract / RFP-based deployment

Reporting To : Program / Delivery Manager

Location : [Onsite / Hybrid / Remote - to be specified]

Primary Focus : Web Scraping & Data Extraction Architecture (~70%)

Secondary Focus : Data Engineering & Pipeline Architecture (~30%)

1. Role Summary :

We are seeking a senior Solution Architect to lead the design and delivery of large-scale web scraping and data extraction solutions, with strong supporting expertise in data engineering. The architect will be the technical authority on all scraping initiatives - from pre-sales solutioning and effort estimation through to architecture, build, and stabilization. The primary mandate is web scraping: designing resilient crawlers, anti-bot strategies, distributed extraction systems, and compliance frameworks. The secondary mandate is data engineering: ensuring extracted data flows reliably into well-modelled, query-ready storage layers and downstream analytical or operational systems.

2. Pre-Sales, Proposal & Estimation Responsibilities :

This is a non-negotiable part of the role. The Solution Architect is mandatorily involved in every new scraping opportunity from the proposal stage onward, and is jointly accountable with the sales / delivery leadership for the technical correctness of all proposals.

Mandatory Involvement Areas :

- Author or co-author the technical solution section of every new RFP / RFI / proposal document for scraping projects.

- Lead technical discovery calls with prospective clients to understand target sources, data SLAs, volume, and compliance constraints.

- Conduct target-site feasibility assessments - anti-bot complexity, dynamic content, login walls, geo-restrictions, rate limits - and document findings before commercials are committed.

- Own end-to-end effort estimation for scraping engagements: crawler build effort, infrastructure sizing, proxy and CAPTCHA cost projections, maintenance overhead, and contingency.

- Produce solution architecture diagrams, tech stack recommendations, and assumption logs as part of every proposal.

- Define and document SLAs, KPIs, and acceptance criteria proposed to the client.

- Participate in client orals, technical defence sessions, and commercial negotiations as the technical SPOC.

- Maintain an internal estimation knowledge base - reusable estimation templates, complexity matrices, target-site classification, and historical actuals - and continuously refine it after every project closure.

- Sign off on the technical feasibility and risk profile of every proposal before submission. No scraping proposal goes out without architect approval.

3. Core Responsibilities :

3.1 Scraping Architecture & Design (Primary) :

- Design end-to-end scraping solutions covering crawl orchestration, extraction, parsing, storage, and downstream data consumption.

- Define reference architectures for high-volume, high-velocity, and high-variety scraping use cases.

- Architect resilient systems handling JavaScript-heavy sites, CAPTCHAs, rate limiting, IP blocking, and frequent DOM changes.

- Evaluate and select frameworks, proxy networks, and anti-bot bypass strategies aligned with cost, performance, and compliance.

- Design data quality, deduplication, validation, and schema-evolution strategies for scraped data.

3.2 Data Engineering Architecture (Secondary) :

- Design downstream data pipelines that ingest, transform, and serve scraped data to analytics, ML, or operational consumers.

- Architect lakehouse / warehouse layers, define data modelling standards (dimensional, Data Vault, or hybrid), and govern schema evolution.

- Define ELT/ETL patterns, orchestration strategy, and SLA-backed data freshness commitments.

- Establish data quality, observability, lineage, and cataloguing practices across the platform.

- Recommend storage formats, partitioning, and indexing strategies for cost and query performance.

3.3 Delivery Leadership :

- Translate business requirements into technical specifications, sprint plans, and implementation roadmaps.

- Produce HLD, LLD, and Architecture Decision Records (ADRs) for each engagement.

- Provide hands-on guidance, perform code reviews, and mentor scraping and data engineers.

- Drive proof-of-concept (PoC) builds for complex or high-risk targets before full-scale rollout.

3.4 Compliance, Risk & Governance :

- Ensure all scraping work adheres to applicable laws and platform terms (GDPR, CCPA, DPDP, robots.txt, copyright).

- Define and enforce ethical scraping practices, request throttling, and PII handling guidelines.

- Conduct risk assessments for each target source and recommend mitigations.

3.5 Performance, Cost & Operations :

- Establish SLAs for crawl freshness, completeness, and accuracy.

- Design monitoring, alerting, and self-healing mechanisms for scraping pipelines.

- Optimise infrastructure cost (compute, proxies, storage) without compromising delivery KPIs.

4. Required Technical Skills - Primary (Web Scraping) :

These are non-negotiable. The candidate must demonstrate deep, hands-on expertise in each of the following:

4.1 Scraping Frameworks & Tooling :

- Expert-level Python: Scrapy, BeautifulSoup, lxml, Requests, httpx, parsel.

- Headless browser automation: Playwright, Puppeteer, Selenium, Pyppeteer.

- Node.js scraping stack (where applicable): Puppeteer, Cheerio, Crawlee.

4.2 Anti-Bot & Evasion Strategy :

- Proven experience bypassing Cloudflare, Akamai Bot Manager, DataDome, PerimeterX, Imperva, Kasada.

- CAPTCHA handling: reCAPTCHA v2/v3, hCaptcha, FunCaptcha, image / audio solvers; integration with 2Captcha, Anti-Captcha, CapSolver.

- TLS / JA3 / JA4 fingerprinting awareness; HTTP/2 fingerprint evasion; browser fingerprint spoofing.

- Stealth plugins, user-agent rotation, header normalisation, cookie / session management at scale.

4.3 Proxy & Network Infrastructure :

- Hands-on experience with rotating, residential, mobile, ISP, and datacenter proxies.

- Integration with providers: Bright Data, Oxylabs, Smartproxy, NetNut, IPRoyal, SOAX.

- Proxy pool design, health-checking, geo-targeting, sticky sessions, and cost optimisation.

4.4 Parsing & Extraction :

- XPath, CSS selectors, regex, JSON-LD, microdata, RDFa.

- Reverse-engineering of internal / mobile APIs, GraphQL endpoints, and XHR traffic.

- ML/LLM-assisted extraction for unstructured layouts (nice-to-have, increasingly expected).

4 Distributed Crawling & Orchestration :

- Scrapy-Redis, Scrapy Cluster, Frontera, Crawlee, or equivalent distributed crawling frameworks.

- Job scheduling and orchestration with Apache Airflow, Prefect, Dagster, or Celery.

- Queue-based architectures using Kafka, RabbitMQ, AWS SQS, GCP Pub/Sub.

Required Technical Skills - Secondary (Data Engineering) :

Working-to-strong proficiency expected. The architect must be able to design, review, and guide data engineering work without depending on a separate data architect.

5 Data Pipelines & Orchestration :

- ETL / ELT design patterns, idempotent pipelines, CDC (Change Data Capture), incremental loads.

- Apache Airflow, Prefect, Dagster, AWS Glue, Azure Data Factory, GCP Dataflow.

- Stream processing: Kafka Streams, Apache Flink, Spark Structured Streaming.

- Batch processing: Apache Spark (PySpark), Databricks, EMR, Dataproc.

5 Data Modelling & Storage :

- Dimensional modelling (Kimball), Data Vault 2.0, normalised vs. denormalised trade-offs.

- Data warehouses: Snowflake, BigQuery, Redshift, Synapse, Databricks SQL Warehouse.

- Data lakes / lakehouses: Delta Lake, Apache Iceberg, Apache Hudi on S3 / GCS / ADLS.

- OLTP databases: PostgreSQL, MySQL; NoSQL: MongoDB, DynamoDB, Cassandra, Elasticsearch, Redis.

- File formats: Parquet, Avro, ORC, JSON, CSV; partitioning, bucketing, compaction strategies.

5 Transformation & Quality :

- dbt (data build tool) for transformation, testing, and documentation.

- Data quality frameworks: Great Expectations, Soda, Deequ, custom validators.

- Data lineage and cataloguing: DataHub, OpenMetadata, Amundsen, Atlan, Collibra.

5. Cloud, DevOps & Observability :

- Strong on at least one of AWS, GCP, Azure (compute, storage, IAM, networking, serverless).

- Containerisation and orchestration: Docker, Kubernetes (EKS / GKE / AKS), ECS.

- Infrastructure as Code: Terraform, Pulumi, CloudFormation.

- CI/CD: GitHub Actions, GitLab CI, Jenkins, Argo CD.

- Observability: Prometheus, Grafana, ELK / OpenSearch, Datadog, Sentry, OpenTelemetry.

6. Experience & Qualifications :

- Bachelor's or Master's degree in Computer Science, Engineering, or related field.

- 10+ years of overall software engineering experience.

- 5+ years dedicated to large-scale web scraping / data extraction (primary).

- 3+ years of hands-on data engineering experience covering pipelines, warehouses / lakehouses, and orchestration (secondary).

- Proven track record of architecting scraping platforms processing millions of pages per day across diverse target sites.

- Prior experience as Solution Architect, Tech Lead, or Principal Engineer leading teams of 5+ engineers.

- Demonstrated experience in pre-sales / proposal authoring / RFP responses for scraping or data engineering engagements.

- Client-facing consulting experience strongly preferred.

7. Soft Skills :

- Excellent written and verbal communication; able to defend technical proposals to CXO-level audiences.

- Strong commercial acumen - understands the cost / risk / quality trade-offs in estimation.

- Analytical, structured problem-solving mindset.

- Ownership-driven, comfortable being the single point of technical accountability.

- High standards for documentation, knowledge transfer, and reusability.

8. Expected Deliverables (RFP Scope) :

The Solution Architect deployed under this RFP shall be responsible for producing, at minimum, the following artifacts:

Pre-Sales / Proposal Phase :

- Technical solution sections of all proposal documents.

- Target-site feasibility and risk assessment reports.

- Effort estimation models, assumption logs, and complexity matrices.

- Solution architecture diagrams and tech stack recommendations for proposals.

Delivery Phase :

- High-Level Design (HLD) and Low-Level Design (LLD) documents per initiative.

- Architecture Decision Records (ADRs).

- Reference implementations / PoCs for complex extraction scenarios.

- Code review records and engineering standards documentation.

- Operational runbooks, monitoring playbooks, and incident response procedures.

- Compliance and data governance documentation.

- Knowledge transfer sessions and final handover documentation.

9. Nice-to-Have :

- ML / NLP-based extraction, entity resolution, or LLM-assisted parsing experience.

- Exposure to GraphQL, gRPC, mobile API reverse-engineering, mitmproxy / Charles workflows.

- Domain experience: e-commerce price intelligence, market research, financial data, real estate, travel aggregation, hospitality.

- Open-source contributions to scraping or data engineering ecosystems.

- Familiarity with cross-jurisdictional legal frameworks for automated data collection.

- Experience with reverse ETL tools (Hightouch, Census) and feature stores.

10. Evaluation Criteria (RFP Response) :

Vendors proposing candidates against this role will be evaluated on:

1. Depth and recency of relevant scraping architecture experience (primary weight).

2. Strength of supporting data engineering experience (secondary weight).

3. Quality of past project case studies - scale, complexity, business outcomes.

4. Demonstrated capability in proposal authoring and technical estimation.

5. Technical depth in interviews, whiteboarding, and architecture discussions.

6. Communication and stakeholder management capability.

7. Commercial competitiveness of the rate card.

8. Availability and ramp-up timeline.

Skills

Web Scraping, Solution Architect, Technical Architect, Data Engineering, RFP, SLA, ETL Tools, Machine Learning, NLP, LLM, Scrapy

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
986,892 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Solutions
Similar stack
Same company
India
$145k – $175k per year • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Sunnyvale
Python
SQL
Apex
Apex
MuleSoft
Analytics
Tableau
Power BI
Microsoft Excel
Management
Agile
Apply
≈ $73k – $166k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Chandler
Apply
$106k – $167k per year • In office • Full-Time • 8+ years exp • United States
AI/ML
AI Agents
Apply
≈ $121k – $233k per year (Estimated) • Remote (United States) • Top Secret • Full-Time • 5+ years exp • Master's Degree
Python
AI/ML
Scikit-learn
Pandas
NumPy
DevOps
Git
Apply
≈ $108k – $217k per year (Estimated) • In office • Full-Time • 6+ years exp • Staines-upon-Thames
Management
ServiceNow
Agile
ITSM
Apply
$197k – $225k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • New York • McLean • Cambridge • Richmond
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Management
Agile
Apply
$179k – $205k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • San Jose • McLean • Cambridge • San Francisco • New York
Python
JavaScript
Rust
TypeScript
C#
Node JS
Scala
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Bitbucket
GitHub
Management
Agile
Apply
≈ $17k – $37k per year (Estimated) • In office • 3+ years exp • India
SQL
C#
Perl
Databases
Databricks
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Kubernetes
Google GKE
Azure AKS
GitHub
VPN
Analytics
Azure Data Factory
Apply
Technical Lead 2 months ago
≈ $26k – $56k per year (Estimated) • Hybrid • 12+ years exp • India
Python
JavaScript
SQL
Node JS
Python
Scrapy
Beautiful Soup
Celery
pySpark
HTTPX
Node JS
Puppeteer
Cheerio
Databases
MySQL
PostgreSQL
Redis
Snowflake
Databricks
Apache Iceberg
Delta Lake
Cassandra
DynamoDB
RabbitMQ
ElasticSearch
Apache Kafka
OpenSearch
Google BigQuery
Amazon Redshift
Apache Hudi
BigQuery
AI/ML
Spark
Airflow
Dagster
dbt
Prefect
NLP
Great Expectations
Flink
LLM
Feature Store
Frontend
GraphQL
DevOps
gRPC
Terraform
GCP
GitHub Actions
OpenTelemetry
CloudFormation
Datadog
Prometheus
Pulumi
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Cloudflare
Grafana
Self-Healing
Amazon EKS
Google GKE
Azure AKS
Akamai
SLI/SLO/SLA
Amazon S3
IAM
Amazon ECS
Apache HTTP Server
Cybersecurity
GDPR
Analytics
ETL/ELT
Azure Data Factory
AWS Glue
Data Vault
Dimensional Modeling
Collibra
QA
Selenium
Playwright
Sentry
Apply
≈ $28k – $50k per year (Estimated) • In office • 8+ years exp • India
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Management
Agile
Scrum
Apply
≈ $29k – $52k per year (Estimated) • In office • 10+ years exp • Kolkata
Python
AI/ML
Model Context Protocol
Prompt Engineering
Chain-of-Thought
AI Agents
Gemini
Google ADK
Human-in-the-Loop
Agentic Workflows
Multi-Agent Systems
DevOps
GCP
GitHub Actions
GitLab CI
CI/CD
Configuration Management
AIOps
IAM
Cybersecurity
Microsoft Entra ID
Management
Google Workspace
Gmail
SharePoint
Agile
Microsoft Office
Apply
Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Monterrey
Apply
In office • 2+ years exp • Bachelor's Degree • India
Python
Java
AI/ML
AI Agents
DevOps
GitHub Actions
Jenkins
Kubernetes
SaltStack
Akamai
Linux
TCP/IP
DNS
Cybersecurity
OWASP
QA
TestNG
Playwright
Pytest
Apply
Finance Controller 1 day ago
≈ $42k – $114k per year (Estimated) • Remote (United Kingdom) • Full-Time • Bachelor's Degree • United Kingdom • Portugal • India
Apply
See all jobs
This is one of many
986,892 more open roles from verified company boards, updated every day.