A 'robots.txt' Parser and 'Webbot'/'Spider'/'Crawler' Permissions Checker

Provides functions to download and parse 'robots.txt' files. Ultimately the package makes it easy to check if bots (spiders, crawler, scrapers, ...) are allowed to access specific resources on a domain.


Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("robotstxt")

0.7.15 by Pedro Baltazar, 2 years ago


https://docs.ropensci.org/robotstxt/, https://github.com/ropensci/robotstxt


Report a bug at https://github.com/ropensci/robotstxt/issues


Browse source code at https://github.com/cran/robotstxt


Authors: Pedro Baltazar [aut, cre] , Peter Meissner [aut] , Kun Ren [aut, cph] (Author and copyright holder of list_merge.R.) , Oliver Keys [ctb] (original release code review) , Rich Fitz John [ctb] (original release code review)


Documentation:   PDF Manual  


MIT + file LICENSE license


Imports stringr, httr, spiderbar, future.apply, magrittr, utils

Suggests knitr, rmarkdown, dplyr, testthat, covr, curl


Imported by ralger.

Suggested by newsanchor, spiderbar, vosonSML, webchem.


See at CRAN