In this live tutorial we show how to install dependencies and get started scraping an ecommerce page using Python libraries including Selenium, BeautifulSoup, and Diffbot's Crawlbot and Automatic Extraction APIs.
While all three extraction methods work for our given target, we consider some of the pros and cons of these different approaches and also talk about web extraction best practices.
Links In Video:
Example Scrape Page: https://www.iherb.com/c/Manuka-Honey
Install Python: https://www.python.org/downloads/
Install Jupyter Notebook: https://jupyter.org/install
Install Selenium: https://pypi.org/project/selenium/
Webdrivers for Selenium: https://selenium-python.readthedocs.i...
Selenium Selectors: https://selenium-python.readthedocs.i...
Installing BeautifulSoup: https://www.crummy.com/software/Beaut...
BeautifulSoup FindAll(): https://www.crummy.com/software/Beaut...
BeautifulSoup Types of Objects: https://www.crummy.com/software/Beaut...
Diffbot Free Trial: https://www.diffbot.com/get-started/
Diffbot Python Boilerplate: https://github.com/diffbot/diffbot-py...
Diffbot Crawlbot API Docs: https://docs.diffbot.com/docs/en/api-...
Finding Your Diffbot Token: https://app.diffbot.com/diffbot-users...
Diffbot Crawlbot Interface: https://app.diffbot.com/crawls/