Forms, logins and APIs
Most of the web you scrape is GET requests, but not all of it. A search box posts its query, a login posts a username and password, and a page’s own JavaScript often fetches its data from a JSON API with a POST. netweir sends all of these, and sends each the way the browser in your profile would.
Two kinds of POST
Section titled “Two kinds of POST”A browser sends a POST in one of two ways, and a server can tell them apart:
- Submitting a form is a navigation: the browser leaves the page and
loads the answer. It sends
Origin, aReferer,Sec-Fetch-Mode: navigateand, in Chrome,Cache-Control: max-age=0. - A script’s request (
fetch()in the page’s JavaScript) isn’t. It sendsAccept: */*andSec-Fetch-Mode: cors, leaves out the headers a navigation carries, and orders the rest differently.
form= sends the first kind; json= and body= send the second. Chrome
154, Firefox 156 and Safari 27 were recorded doing both, and following a
link, and loading the page a login redirects to; netweir’s tests check its
requests match them byte for byte, header order included.
Sending data
Section titled “Sending data”import netweir
# A form submission: URL-encoded, as a browser posts a form.page = netweir.post("https://example.com/search", form={"q": "rust", "page": 2})
# A script's request with JSON, as a page's fetch() sends it.page = netweir.post("https://example.com/api/search", json={"q": "rust"})
# Any body; without a Content-Type it's sent as text, as fetch() sends a string.page = netweir.post("https://example.com/upload", body=b"a,b\n1,2", headers={"Content-Type": "text/csv"})
# Any method.page = netweir.request("PUT", "https://example.com/api/items/7", json={"price": 9})Pass one of form=, json= and body=, with a method that carries a body:
a GET or HEAD with one is a ValueError. referer= names the page the
request comes from, which sets Referer, Origin and Sec-Fetch-Site as
the browser would; without it, a form or a script’s request comes from the
root of the target’s site. A GET with referer= looks like following a
link on that page:
page = netweir.get("https://example.com/item/7", referer="https://example.com/list")Redirects follow the browsers’ rules. A 303, or a 301 or 302 after a POST, turns into a GET without a body, which is how a login usually answers; a 307 or 308 repeats the request, body and all.
Forms on a page
Section titled “Forms on a page”page.form() reads a form the way the browser would submit it: its method,
where it goes, how it’s encoded, and the fields in order. That means the
named, enabled controls; checked checkboxes and radio buttons; each select’s
selected options (or its first); textareas; and the submit button pressed.
Hidden fields come along, which is where sites keep their CSRF tokens.
async with netweir.Client() as client: login = await client.get("https://example.com/login") form = login.form(data={"username": "ada", "password": "s3cret"}) home = await client.submit(form)Use a Client for anything with a session: it keeps the cookies the login
page set, and the form usually needs them. netweir.submit(form) works too,
but starts without any cookies.
Pick the form with a CSS selector or an XPath, or by formid=,
formname= or formnumber= (the first form otherwise). data= replaces
fields of the same name and adds new ones; a list value sends the name more
than once. click= names the submit button pressed; it’s the first one
unless you say otherwise, and click=False presses none.
form = page.form("form.search", data={"q": "maps", "sort": "price"}, click="go")form = page.form(formname="lookup")form.method, form.action, form.enctype, form.fieldsA GET form puts its fields in the query string. A multipart/form-data
form is sent as multipart, with a boundary shaped as Chrome’s; file uploads
aren’t supported yet.
In a crawl
Section titled “In a crawl”Request takes the same arguments, and its method is POST when there’s a
form, JSON or a body and GET otherwise. Request.from_form(page, ...)
submits a form on a page, as Scrapy’s FormRequest.from_response does:
class Search(netweir.Spider): start_urls = ["https://example.com/search"]
def parse(self, page): for term in ["rust", "python"]: yield netweir.Request.from_form(page, data={"q": term}, callback=self.results)
def results(self, page): for row in page.css(".result"): yield {"title": row.css("h2::text").get()}
class Api(netweir.Spider): def start(self): yield netweir.Request("https://example.com/api/items", method="POST", json={"page": 1})page.follow() sends links as followed from the page they’re on, with its
Referer.
The duplicate filter tells requests apart by their body as well as their URL, so two searches for different terms are two requests. A checkpoint keeps each request’s method and body, and a resumed crawl sends them again as they were.
Two things are different for a POST, because sending one twice can do something twice (place an order, post a comment):
- A POST or PATCH that gets a server error (500, 502, 504 and the like), or
loses its connection after it was sent, isn’t retried. Pass
retry_post=Trueon the request if repeating it is safe. A 429, a 503 withRetry-After, or a block page from bot protection says the site never acted on it, so it’s retried as any other. - It’s never handed to Chrome. With
browser="on_block", a blocked POST goes to youron_blockhandler instead, andbrowser=Trueon a POST is aValueError.