The Web Scraping Club
Subscribe
Sign in
Home
News
The Lab
Advertise on TWSC
Proxy pricing benchmark
Consulting
Archive
About
Prepare Web Content and Documents for LLM Ingestion Using MarkItDown
Markdown is the language of AI. Learn how to convert different data formats into it!
READ THE LATEST
Most Popular
View all
THE LAB #1: Scraping data from an app
Sep 4, 2022
•
Pierluigi Vinciguerra
16
4
THE LAB #3: Scraping Cloudflare protected websites
Sep 27, 2022
•
Pierluigi Vinciguerra
10
1
THE LAB #73: How to Bypass Cloudflare in 2025
Jan 23, 2025
•
Pierluigi Vinciguerra
18
4
1
Ten years of web scraping: a personal perspective about selling web data
Mar 24, 2024
•
Pierluigi Vinciguerra
19
1
Platinum Partner
NetNut: The Fastest & Most Reliable Proxy Network for Web Scrapers
Meet the Platinum Partner of The Web Scraping Club
Feb 6, 2025
•
Pierluigi Vinciguerra
1
The Great Web Unblocker Benchmark - Cloudflare Edition
Testing different web unblockers against Indeed.com
Sep 22, 2024
•
Pierluigi Vinciguerra
6
The Web Unblocker Cost Benchmark
Price comparison between the most well-known web unblockers on the market
Dec 24, 2023
•
Pierluigi Vinciguerra
2
1
Recent posts
View all
Tool of the week #1: Bright Data Scraper Studio, From Prompt to Production Scraper in Minutes
Explore how Bright Data's Scraper Studio bridges the gap between AI-assisted building and full IDE-level control without locking you into either.
Aug 4
•
Federico Trotta
3
2
Why Your RAG Pipelines Are Only as Good as Your Data Acquisition Layer
A deep dive into what RAG actually requires, where web search integration falls short, and why scraping is the only retrieval layer that scales.
Aug 2
•
Federico Trotta
4
1
THE LAB #112: Intercepting a Website's Internal API Calls
How I pull a store's catalog from its internal API, by hand and in code.
Jul 30
•
Pierluigi Vinciguerra
1
See all
AI
View all
Why Your RAG Pipelines Are Only as Good as Your Data Acquisition Layer
A deep dive into what RAG actually requires, where web search integration falls short, and why scraping is the only retrieval layer that scales.
Aug 2
•
Federico Trotta
4
1
Build an AI Agent for Scraping and Analyzing Research Papers
Let’s build an AI agent in Python for research paper scraping and analysis.
Dec 14, 2025
•
Antonello Zanini
1
Using NLP for Entity Extraction From Scraped Data
From theory to practice: how to extract entities from textual scraped data using NLP
Dec 7, 2025
•
Federico Trotta
2
1
Using AI to Detect Patterns in Scraped Data
A practical guide on finding patterns in scraped data with advanced techniques
Nov 23, 2025
•
Federico Trotta
2
2
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts