Search for and download PDF file links from a webpage.
pip install pdf_hunterAfter installing, the pdf-hunter executable is made available in your path.
By default, running pdf-hunter with a webpage URL will print all discovered absolute PDF links to standard output:
pdf-hunter "https://example.com/books-list"Output:
https://example.com/books/guide-to-python.pdf
https://example.com/books/advanced-algorithms.pdf
Pass the -d (or --download) flag to download all discovered PDFs to your current directory:
pdf-hunter "https://example.com/books-list" -dUse the -o (or --output-dir) option to specify a target directory for the downloaded files:
pdf-hunter "https://example.com/books-list" -d -o /path/to/downloadsYou can also use pdf-hunter programmatically in your Python scripts.
import pdf_hunter
url = "https://github.com/EbookFoundation/free-programming-books/blob/main/books/free-programming-books-langs.md"pdf_urls = pdf_hunter.get_pdf_urls(url)
print(pdf_urls[:3])Output:
[
"https://www.cs.uni.edu/~mccormic/4740/guide-c2ada.pdf",
"http://www.adapower.com/pdfs/AdaDistilled07-27-2003.pdf",
"https://www.adacore.com/uploads/books/pdf/Ada_for_the_C_or_Java_Developer-cc.pdf",
]import os
pdf_url = pdf_urls[0]
file_name = pdf_hunter.get_pdf_name(pdf_url)
# Download to a specific directory
pdf_hunter.download_file(pdf_url, folder_path=os.getcwd())
print(os.path.isfile(file_name)) # Truepdf_hunter.download_pdf_files(url, folder_path=os.getcwd())