Join the Telegram channel to stay updated, report bugs, or request custom data mining scripts.
PyPaperBot is a Python tool for downloading scientific papers and bibtex using Google Scholar, Crossref, SciHub, and SciDB. The tool tries to download papers from different sources such as PDF provided by Scholar, Scholar related links, and Scihub. PyPaperbot is also able to download the bibtex of each paper.
- Download papers given a query
- Download papers given paper's DOIs
- Download papers given a Google Scholar link
- Generate Bibtex of the downloaded paper
- Filter downloaded paper by year, journal and citations number
Use pip to install from pypi:
pip install PyPaperBotIf on windows you get an error saying error: Microsoft Visual C++ 14.0 is required.. try to install Microsoft C++ Build Tools or Visual Studio
Since numpy cannot be directly installed....
pkg install wget
wget https://its-pointless.github.io/setup-pointless-repo.sh
pkg install numpy
export CFLAGS="-Wno-deprecated-declarations -Wno-unreachable-code"
pip install pandas
and
pip install PyPaperbot
PyPaperBot arguments:
You can use only one of the arguments in the following groups
- --query, --doi-file, and --doi
- --max-dwn-year and and max-dwn-cites
One of the arguments --scholar-pages, --query , and --file is mandatory The arguments --scholar-pages is mandatory when using *--query * The argument --dwn-dir is mandatory
The argument --journal-filter require the path of a CSV containing a list of journal name paired with a boolean which indicates whether or not to consider that journal (0: don't consider /1: consider) Example
The argument --doi-file require the path of a txt file containing the list of paper's DOIs to download organized with one DOI per line Example
Use the --proxy argument at the end of all other arguments and specify the protocol to be used. See the examples to understand how to use the option.
If access to SciHub is blocked in your country, consider using a free VPN service like ProtonVPN Also, you can use proxy option above.
Download a maximum of 30 papers from the first 3 pages given a query and starting from 2018 using the mirror https://sci-hub.do:
python -m PyPaperBot --query="Machine learning" --scholar-pages=3 --min-year=2018 --dwn-dir="C:\User\example\papers" --scihub-mirror="https://sci-hub.do"Download papers from pages 4 to 7 (7th included) given a query and skip words:
python -m PyPaperBot --query="Machine learning" --scholar-pages=4-7 --dwn-dir="C:\User\example\papers" --skip-words="ai,decision tree,bot"Download a paper given the DOI:
python -m PyPaperBot --doi="10.0086/s41037-711-0132-1" --dwn-dir="C:\User\example\papers" -use-doi-as-filename`Download papers given a file containing the DOIs:
python -m PyPaperBot --doi-file="C:\User\example\papers\file.txt" --dwn-dir="C:\User\example\papers"`If it doesn't work, try to use py instead of python i.e.
py -m PyPaperBot --doi="10.0086/s41037-711-0132-1" --dwn-dir="C:\User\example\papers"`Search papers that cite another (find ID in scholar address bar when you search citations):
python -m PyPaperBot --cites=3120460092236365926Using proxy
python -m PyPaperBot --query=rheumatoid+arthritis --scholar-pages=1 --scholar-results=7 --dwn-dir=/download --proxy="http://1.1.1.1::8080,https://8.8.8.8::8080"
python -m PyPaperBot --query=rheumatoid+arthritis --scholar-pages=1 --scholar-results=7 --dwn-dir=/download -single-proxy=http://1.1.1.1::8080
In termux, you can directly use PyPaperBot followed by arguments...
Feel free to contribute to this project by proposing any change, fix, and enhancement on the dev branch
- Tests
- Code documentation
- General improvements
This application is for educational purposes only. I do not take responsibility for what you choose to do with this application.
If you like this project, you can give me a cup of coffee :)

