List Crawling is a guiding blog about crawling, crawlers, and the lists they read. The site is published at listcrawling.org. It explains how a crawler finds pages, how a list crawl stays inside one index, and what changed in the tools and rules that affect the next job.
The writing is for analysts, researchers, and engineers who need a clear picture before they run anything. A guide is a standing explanation: what a crawler is, how pagination continues, how to tell a finished list from a page-one file, and how robots.txt differs from a contract. An update is dated. Vendors change defaults, sites move lists behind scripts, and a court or a standard can clarify a rule. Those notes say what changed, who it affects, and whether an older guide still holds.
What we mean by list crawling
A list page is an index, not the record. Category grids, job boards, city directories, and search-results pages stack cards that share a shape. List crawling treats that shape as the unit of work: collect every item link once, then read the record behind each link. A second, narrower meaning also appears in SEO and migration work. There, a list crawl is a fetch of URLs you already have, with no link following. Guides on this site label which job they mean.
What this site will not do
List Crawling does not teach login bypasses, challenge-page defeat, or private contact harvesting. Public business facts are the ordinary subject: a shop name, a published price, a job title on an open board. If a publisher offers an API or a CSV, that file is the better source. A crawler is the fallback.
How to use the site
Read the definition guide if crawler is still a fuzzy word. Read the list-crawl guide before you design a job. Read the pagination note if a file came back short.
Read the limits note before you store a column you did not mean to collect. Then check updates for the tool you are about to touch.
