Skip to content

Getting started

This page takes you from installing netweir to a crawl that writes a file of items. It takes about ten minutes. Each step builds on the one before.

netweir needs Python 3.10 or later, on Linux, macOS or Windows.

pip install netweir

The wheels include everything netweir needs, BoringSSL and the HTML parser among it, so there’s nothing else to install.

import netweir
page = netweir.get("https://books.toscrape.com/")
print(page.status) # 200
print(page.css("h1::text").get()) # All products

netweir.get sends the request the way Chrome would and waits for the answer. The page it returns holds the response (status, url, headers) and the parsed document.

Query the page with CSS, XPath or Beautiful Soup’s methods. They all work on the same page, so use whichever you know:

prices = page.css(".price_color::text").getall()
titles = page.xpath("//article//h3/a/@title").getall()
next_link = page.find("li", class_="next").find("a")["href"]

::text gives an element’s own text and ::attr(name) an attribute’s value. Without them, get() gives the element’s HTML. See Selecting for everything a query can do.

A spider fetches pages for you and follows links. Put this in books.py:

import netweir
class Books(netweir.Spider):
start_urls = ["https://books.toscrape.com/"]
async def parse(self, page):
for book in page.css("article.product_pod"):
yield {
"title": book.css("h3 a::attr(title)").get(),
"price": book.css(".price_color::text").get(),
}
if next_page := page.css("li.next a::attr(href)").get():
yield page.follow(next_page)

parse runs on every page the spider fetches. Each dict it yields is an item; each page.follow(...) is another page to fetch.

netweir crawl books.py -o books.jsonl

netweir fetches all 50 pages of the site and writes 1,000 items to books.jsonl, one JSON object per line. It reads the site’s robots.txt first and keeps a polite pace, so the crawl takes about 20 seconds. Name the file books.csv or books.parquet to get that format instead.

  • Crawling: requests, callbacks, pipelines, exporters and the command line.
  • Declarative spiders: describe items and links, and let Rust do the crawl without running Python on each page.
  • Self-healing: block recovery, crash-safe resume and selectors that survive a redesign.
  • Settings: every setting and its default.