netweir.Settings holds every option a crawl has, each with a default that
is safe to start from. Set them on a spider:
class Books ( netweir . Spider ):
settings = netweir . Settings ( concurrency = 16 , max_depth = 3 )
or from the shell with -s name=value (-s max_depth=none for None,
and commas between the proxies of -s proxies=...). A misspelt name is an
error, not ignored. Times are in seconds.
Setting
Default
Does
profile
"chrome"
the browser every request looks like: "chrome", "firefox" or "safari" (the newest of each), or a version such as "safari-27-macos"
proxy
None
an http://, https:// or socks5:// proxy for every request
proxies
()
proxies a site moves through each time it blocks a session
timeout
30.0
the limit for one whole request
max_response_size
67108864
the most bytes of a response body (64 MiB), counted after decompression, so a small compressed answer that unpacks to gigabytes is stopped too; a larger one fails with FetchError(kind="too_large"), and None reads any size
Setting
Default
Does
concurrency
64
requests in flight across all sites
per_domain
8
requests in flight to one site
obey_robots
True
read robots.txt and stay out of what it disallows
robots_agent
"netweir"
the name matched against robots.txt’s User-agent lines
obey_tdmrep
True
skip pages whose owners reserve text and data mining rights
throttle
True
adapt each site’s delay to how fast it answers
start_delay
1.0
each site’s delay before its first answer is timed
min_delay
0.0
the shortest delay the throttle sets
max_delay
60.0
the longest delay, from the throttle, blocks or Crawl-delay
target_concurrency
1.0
requests the throttle aims to have in flight to one site
Setting
Default
Does
max_depth
None
links from a start page beyond which requests are dropped
max_pages_per_domain
None
requests accepted for one site beyond which more are dropped
max_items
None
stop the crawl once it has delivered this many items
max_pages
None
send no more than this many requests, a retry counting as one, then stop
max_errors
None
stop the crawl after this many errors in callbacks and pipelines
max_time
None
stop the crawl after this many seconds
Setting
Default
Does
retries
3
further tries after a server error, a network error, throttling or a block
backoff_base
1.0
retry n waits a random time up to backoff_base × 2ⁿ
backoff_max
60.0
the most any retry waits
breaker_window
50
how many of a site’s last responses the circuit breaker looks at
breaker_ratio
0.3
the share of those that were blocks above which the site pauses
breaker_pause
300.0
how long the site pauses
Setting
Default
Does
browser
"off"
which requests go through Chrome: "off" (those with browser=True), "on_block" (also any still blocked after its retries) or "always"
browser_pages
4
Chrome pages open at once
Setting
Default
Does
checkpoint
None
a directory to keep the crawl’s state in, so it resumes after a crash
track_threshold
0.75
how similar an element must be to count as a tracked selector’s element
fail_fast
False
stop the crawl at the first exception in a callback or pipeline
workers
1
processes to run callbacks in; see Crawling