{"id":1285580,"url":"https://alion.io/job/meritgrouplimited-solution-architect","title":"Solution Architect","company":{"id":3809052,"name":"Meritgrouplimited","domain":"meritgrouplimited.com","url":"https://alion.io/company/meritgrouplimited","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Solutions","role_family":"Solutions","seniority":"staff","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":29000,"max_usd":51000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":35},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Akamai","optional":false},{"name":"Beautiful Soup","optional":false},{"name":"Cheerio","optional":false},{"name":"Cloudflare","optional":false},{"name":"Data Vault","optional":false},{"name":"ETL/ELT","optional":false},{"name":"GDPR","optional":false},{"name":"GraphQL","optional":false},{"name":"HTTPX","optional":false},{"name":"JavaScript","optional":false},{"name":"LLM","optional":false},{"name":"Node JS","optional":false},{"name":"Playwright","optional":false},{"name":"Puppeteer","optional":false},{"name":"Python","optional":false},{"name":"Scrapy","optional":false},{"name":"Selenium","optional":false},{"name":"Self-Healing","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Airflow","optional":true},{"name":"Amazon ECS","optional":true},{"name":"Amazon EKS","optional":true},{"name":"Amazon Redshift","optional":true},{"name":"Amazon S3","optional":true},{"name":"Apache HTTP Server","optional":true},{"name":"Apache Hudi","optional":true},{"name":"Apache Iceberg","optional":true},{"name":"Apache Kafka","optional":true},{"name":"ArgoCD","optional":true},{"name":"AWS","optional":true},{"name":"AWS Glue","optional":true},{"name":"Azure","optional":true},{"name":"Azure AKS","optional":true},{"name":"Azure Data Factory","optional":true},{"name":"BigQuery","optional":true},{"name":"Cassandra","optional":true},{"name":"Celery","optional":true},{"name":"CI/CD","optional":true},{"name":"CloudFormation","optional":true},{"name":"Collibra","optional":true},{"name":"Dagster","optional":true},{"name":"Databricks","optional":true},{"name":"Datadog","optional":true},{"name":"dbt","optional":true},{"name":"Delta Lake","optional":true},{"name":"Dimensional Modeling","optional":true},{"name":"Docker","optional":true},{"name":"DynamoDB","optional":true},{"name":"ElasticSearch","optional":true},{"name":"Feature Store","optional":true},{"name":"Flink","optional":true},{"name":"GCP","optional":true},{"name":"GitHub Actions","optional":true},{"name":"GitLab CI","optional":true},{"name":"Google BigQuery","optional":true},{"name":"Google GKE","optional":true},{"name":"Grafana","optional":true},{"name":"Great Expectations","optional":true},{"name":"gRPC","optional":true},{"name":"IAM","optional":true},{"name":"Jenkins","optional":true},{"name":"Kubernetes","optional":true},{"name":"Machine Learning","optional":true},{"name":"MySQL","optional":true},{"name":"NLP","optional":true},{"name":"OpenSearch","optional":true},{"name":"OpenTelemetry","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Prefect","optional":true},{"name":"Prometheus","optional":true},{"name":"Pulumi","optional":true},{"name":"pySpark","optional":true},{"name":"RabbitMQ","optional":true},{"name":"Redis","optional":true},{"name":"Sentry","optional":true},{"name":"Snowflake","optional":true},{"name":"Spark","optional":true},{"name":"SQL","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-08-13T10:35:50Z","employer_posted_date":null,"last_verified_at":"2026-08-13T10:35:50Z","board_verified":false,"closed_at":null,"days_open":54,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":54},"description":"Job Description :\n\nSolution Architect - Web Scraping & Data Engineering\n\nRFP Reference Document\n\nRole Title : Solution Architect - Web Scraping (Primary) & Data Engineering (Secondary)\n\nEngagement Type : Contract / RFP-based deployment\n\nReporting To : Program / Delivery Manager\n\nLocation : [Onsite / Hybrid / Remote - to be specified]\n\nPrimary Focus : Web Scraping & Data Extraction Architecture (~70%)\n\nSecondary Focus : Data Engineering & Pipeline Architecture (~30%)\n\n1. Role Summary :\n\nWe are seeking a senior Solution Architect to lead the design and delivery of large-scale web scraping and data extraction solutions, with strong supporting expertise in data engineering. The architect will be the technical authority on all scraping initiatives - from pre-sales solutioning and effort estimation through to architecture, build, and stabilization. The primary mandate is web scraping: designing resilient crawlers, anti-bot strategies, distributed extraction systems, and compliance frameworks. The secondary mandate is data engineering: ensuring extracted data flows reliably into well-modelled, query-ready storage layers and downstream analytical or operational systems.\n\n2. Pre-Sales, Proposal & Estimation Responsibilities :\n\nThis is a non-negotiable part of the role. The Solution Architect is mandatorily involved in every new scraping opportunity from the proposal stage onward, and is jointly accountable with the sales / delivery leadership for the technical correctness of all proposals.\n\nMandatory Involvement Areas :\n\n- Author or co-author the technical solution section of every new RFP / RFI / proposal document for scraping projects.\n\n- Lead technical discovery calls with prospective clients to understand target sources, data SLAs, volume, and compliance constraints.\n\n- Conduct target-site feasibility assessments - anti-bot complexity, dynamic content, login walls, geo-restrictions, rate limits - and document findings before commercials are committed.\n\n- Own end-to-end effort estimation for scraping engagements: crawler build effort, infrastructure sizing, proxy and CAPTCHA cost projections, maintenance overhead, and contingency.\n\n- Produce solution architecture diagrams, tech stack recommendations, and assumption logs as part of every proposal.\n\n- Define and document SLAs, KPIs, and acceptance criteria proposed to the client.\n\n- Participate in client orals, technical defence sessions, and commercial negotiations as the technical SPOC.\n\n- Maintain an internal estimation knowledge base - reusable estimation templates, complexity matrices, target-site classification, and historical actuals - and continuously refine it after every project closure.\n\n- Sign off on the technical feasibility and risk profile of every proposal before submission. No scraping proposal goes out without architect approval.\n\n3. Core Responsibilities :\n\n3.1 Scraping Architecture & Design (Primary) :\n\n- Design end-to-end scraping solutions covering crawl orchestration, extraction, parsing, storage, and downstream data consumption.\n\n- Define reference architectures for high-volume, high-velocity, and high-variety scraping use cases.\n\n- Architect resilient systems handling JavaScript-heavy sites, CAPTCHAs, rate limiting, IP blocking, and frequent DOM changes.\n\n- Evaluate and select frameworks, proxy networks, and anti-bot bypass strategies aligned with cost, performance, and compliance.\n\n- Design data quality, deduplication, validation, and schema-evolution strategies for scraped data.\n\n3.2 Data Engineering Architecture (Secondary) :\n\n- Design downstream data pipelines that ingest, transform, and serve scraped data to analytics, ML, or operational consumers.\n\n- Architect lakehouse / warehouse layers, define data modelling standards (dimensional, Data Vault, or hybrid), and govern schema evolution.\n\n- Define ELT/ETL patterns, orchestration strategy, and SLA-backed data freshness commitments.\n\n- Establish data quality, observability, lineage, and cataloguing practices across the platform.\n\n- Recommend storage formats, partitioning, and indexing strategies for cost and query performance.\n\n3.3 Delivery Leadership :\n\n- Translate business requirements into technical specifications, sprint plans, and implementation roadmaps.\n\n- Produce HLD, LLD, and Architecture Decision Records (ADRs) for each engagement.\n\n- Provide hands-on guidance, perform code reviews, and mentor scraping and data engineers.\n\n- Drive proof-of-concept (PoC) builds for complex or high-risk targets before full-scale rollout.\n\n3.4 Compliance, Risk & Governance :\n\n- Ensure all scraping work adheres to applicable laws and platform terms (GDPR, CCPA, DPDP, robots.txt, copyright).\n\n- Define and enforce ethical scraping practices, request throttling, and PII handling guidelines.\n\n- Conduct risk assessments for each target source and recommend mitigations.\n\n3.5 Performance, Cost & Operations :\n\n- Establish SLAs for crawl freshness, completeness, and accuracy.\n\n- Design monitoring, alerting, and self-healing mechanisms for scraping pipelines.\n\n- Optimise infrastructure cost (compute, proxies, storage) without compromising delivery KPIs.\n\n4. Required Technical Skills - Primary (Web Scraping) :\n\nThese are non-negotiable. The candidate must demonstrate deep, hands-on expertise in each of the following:\n\n4.1 Scraping Frameworks & Tooling :\n\n- Expert-level Python: Scrapy, BeautifulSoup, lxml, Requests, httpx, parsel.\n\n- Headless browser automation: Playwright, Puppeteer, Selenium, Pyppeteer.\n\n- Node.js scraping stack (where applicable): Puppeteer, Cheerio, Crawlee.\n\n4.2 Anti-Bot & Evasion Strategy :\n\n- Proven experience bypassing Cloudflare, Akamai Bot Manager, DataDome, PerimeterX, Imperva, Kasada.\n\n- CAPTCHA handling: reCAPTCHA v2/v3, hCaptcha, FunCaptcha, image / audio solvers; integration with 2Captcha, Anti-Captcha, CapSolver.\n\n- TLS / JA3 / JA4 fingerprinting awareness; HTTP/2 fingerprint evasion; browser fingerprint spoofing.\n\n- Stealth plugins, user-agent rotation, header normalisation, cookie / session management at scale.\n\n4.3 Proxy & Network Infrastructure :\n\n- Hands-on experience with rotating, residential, mobile, ISP, and datacenter proxies.\n\n- Integration with providers: Bright Data, Oxylabs, Smartproxy, NetNut, IPRoyal, SOAX.\n\n- Proxy pool design, health-checking, geo-targeting, sticky sessions, and cost optimisation.\n\n4.4 Parsing & Extraction :\n\n- XPath, CSS selectors, regex, JSON-LD, microdata, RDFa.\n\n- Reverse-engineering of internal / mobile APIs, GraphQL endpoints, and XHR traffic.\n\n- ML/LLM-assisted extraction for unstructured layouts (nice-to-have, increasingly expected).\n\n4 Distributed Crawling & Orchestration :\n\n- Scrapy-Redis, Scrapy Cluster, Frontera, Crawlee, or equivalent distributed crawling frameworks.\n\n- Job scheduling and orchestration with Apache Airflow, Prefect, Dagster, or Celery.\n\n- Queue-based architectures using Kafka, RabbitMQ, AWS SQS, GCP Pub/Sub.\n\nRequired Technical Skills - Secondary (Data Engineering) :\n\nWorking-to-strong proficiency expected. The architect must be able to design, review, and guide data engineering work without depending on a separate data architect.\n\n5 Data Pipelines & Orchestration :\n\n- ETL / ELT design patterns, idempotent pipelines, CDC (Change Data Capture), incremental loads.\n\n- Apache Airflow, Prefect, Dagster, AWS Glue, Azure Data Factory, GCP Dataflow.\n\n- Stream processing: Kafka Streams, Apache Flink, Spark Structured Streaming.\n\n- Batch processing: Apache Spark (PySpark), Databricks, EMR, Dataproc.\n\n5 Data Modelling & Storage :\n\n- Dimensional modelling (Kimball), Data Vault 2.0, normalised vs. denormalised trade-offs.\n\n- Data warehouses: Snowflake, BigQuery, Redshift, Synapse, Databricks SQL Warehouse.\n\n- Data lakes / lakehouses: Delta Lake, Apache Iceberg, Apache Hudi on S3 / GCS / ADLS.\n\n- OLTP databases: PostgreSQL, MySQL; NoSQL: MongoDB, DynamoDB, Cassandra, Elasticsearch, Redis.\n\n- File formats: Parquet, Avro, ORC, JSON, CSV; partitioning, bucketing, compaction strategies.\n\n5 Transformation & Quality :\n\n- dbt (data build tool) for transformation, testing, and documentation.\n\n- Data quality frameworks: Great Expectations, Soda, Deequ, custom validators.\n\n- Data lineage and cataloguing: DataHub, OpenMetadata, Amundsen, Atlan, Collibra.\n\n5. Cloud, DevOps & Observability :\n\n- Strong on at least one of AWS, GCP, Azure (compute, storage, IAM, networking, serverless).\n\n- Containerisation and orchestration: Docker, Kubernetes (EKS / GKE / AKS), ECS.\n\n- Infrastructure as Code: Terraform, Pulumi, CloudFormation.\n\n- CI/CD: GitHub Actions, GitLab CI, Jenkins, Argo CD.\n\n- Observability: Prometheus, Grafana, ELK / OpenSearch, Datadog, Sentry, OpenTelemetry.\n\n6. Experience & Qualifications :\n\n- Bachelor's or Master's degree in Computer Science, Engineering, or related field.\n\n- 10+ years of overall software engineering experience.\n\n- 5+ years dedicated to large-scale web scraping / data extraction (primary).\n\n- 3+ years of hands-on data engineering experience covering pipelines, warehouses / lakehouses, and orchestration (secondary).\n\n- Proven track record of architecting scraping platforms processing millions of pages per day across diverse target sites.\n\n- Prior experience as Solution Architect, Tech Lead, or Principal Engineer leading teams of 5+ engineers.\n\n- Demonstrated experience in pre-sales / proposal authoring / RFP responses for scraping or data engineering engagements.\n\n- Client-facing consulting experience strongly preferred.\n\n7. Soft Skills :\n\n- Excellent written and verbal communication; able to defend technical proposals to CXO-level audiences.\n\n- Strong commercial acumen - understands the cost / risk / quality trade-offs in estimation.\n\n- Analytical, structured problem-solving mindset.\n\n- Ownership-driven, comfortable being the single point of technical accountability.\n\n- High standards for documentation, knowledge transfer, and reusability.\n\n8. Expected Deliverables (RFP Scope) :\n\nThe Solution Architect deployed under this RFP shall be responsible for producing, at minimum, the following artifacts:\n\nPre-Sales / Proposal Phase :\n\n- Technical solution sections of all proposal documents.\n\n- Target-site feasibility and risk assessment reports.\n\n- Effort estimation models, assumption logs, and complexity matrices.\n\n- Solution architecture diagrams and tech stack recommendations for proposals.\n\nDelivery Phase :\n\n- High-Level Design (HLD) and Low-Level Design (LLD) documents per initiative.\n\n- Architecture Decision Records (ADRs).\n\n- Reference implementations / PoCs for complex extraction scenarios.\n\n- Code review records and engineering standards documentation.\n\n- Operational runbooks, monitoring playbooks, and incident response procedures.\n\n- Compliance and data governance documentation.\n\n- Knowledge transfer sessions and final handover documentation.\n\n9. Nice-to-Have :\n\n- ML / NLP-based extraction, entity resolution, or LLM-assisted parsing experience.\n\n- Exposure to GraphQL, gRPC, mobile API reverse-engineering, mitmproxy / Charles workflows.\n\n- Domain experience: e-commerce price intelligence, market research, financial data, real estate, travel aggregation, hospitality.\n\n- Open-source contributions to scraping or data engineering ecosystems.\n\n- Familiarity with cross-jurisdictional legal frameworks for automated data collection.\n\n- Experience with reverse ETL tools (Hightouch, Census) and feature stores.\n\n10. Evaluation Criteria (RFP Response) :\n\nVendors proposing candidates against this role will be evaluated on:\n\n1. Depth and recency of relevant scraping architecture experience (primary weight).\n\n2. Strength of supporting data engineering experience (secondary weight).\n\n3. Quality of past project case studies - scale, complexity, business outcomes.\n\n4. Demonstrated capability in proposal authoring and technical estimation.\n\n5. Technical depth in interviews, whiteboarding, and architecture discussions.\n\n6. Communication and stakeholder management capability.\n\n7. Commercial compe...","description_format":"text","description_chars":12203,"description_truncated":true,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["BI & Dashboard Software","Data Entry & Processing","B2B Company & Contact Data"],"lifecycle":[{"event":"open","at":"2026-09-26T04:00:00Z"}],"visa":[],"liveness":{"score":10,"band":"cold","label":"Long shot","p_open":0.4,"p_active":0.555,"p_room":0.45,"age_days":53,"expected_fill_days":30,"reasons":["seen:53","win:tail"],"computed_at":"2026-10-06T05:45:30Z"},"pay":null,"html_url":"https://alion.io/job/meritgrouplimited-solution-architect","json_url":"https://alion.io/job/meritgrouplimited-solution-architect.json","meta":{"generated_at":"2026-10-07T01:33:14Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2509,"day_limit":5000,"remaining_today":2491,"minute_limit":60,"resets_at":"2026-10-08T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3809052},"rest":"https://alion.io/mcp/rest/get_company?id=3809052"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fmeritgrouplimited-solution-architect"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fmeritgrouplimited-solution-architect"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fmeritgrouplimited-solution-architect"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/meritgrouplimited-solution-architect\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fmeritgrouplimited-solution-architect"}]}