Scraping the web for data

Data Harvesting

Article from Issue 233/2020

Author(s): Marco Fioretti

Web scraping lets you automatically download and extract data from websites to build your own database. With a simple scraping script, you can harvest information from the web.

If you are looking to collect data from the Internet for a personal database, your first stop is often a Google search. However, a search for mortgage rates can (in theory) return dozens of pages full of relevant images and data, as well as a lot of irrelevant content. You could visit every web page pulled up by your search and cut and paste the relevant data into your database. Or you could use a web scraper to automatically download and extract raw data from the web pages and reformat it into a table, graph, or spreadsheet on your computer.

Not just big data professionals, but also small business owners, teachers, students, shoppers, or just curious people can use web scraping to do all manner of tasks from researching a PhD thesis to creating a database of local doctors to comparing prices for online shopping. Unless you need to do really complicated stuff with super-optimized performance, web scraping is relatively easy. In this article, I'll show you how web scraping works with some practical examples that use the open source tool Beautiful Soup.

Caveats

Web scraping does have its limits. First, you have to start off with a well-crafted search engine query; web scraping can't replace the initial search. To protect their business and observe legal constraints, search engines deploy anti-scraping features; overcoming them is not worth the time of the occasional web scraper. Instead, Web scraping shines (and is irreplaceable) after you have completed your web search.

[...]

Use Express-Checkout link below to read the full article (PDF).

Buy this article as PDF

Express-Checkout as PDF

Price $2.95
(incl. VAT)

Buy Linux Magazine

SINGLE ISSUES

Print Issues

Digital Issues

SUBSCRIPTIONS

Print Subs

Digisubs

TABLET & SMARTPHONE APPS

US / Canada

UK / Australia

Support Our Work

Linux Magazine content is made possible with support from readers like you. Please consider contributing when you’ve found an article to be beneficial.

News

OpenMandriva Lx 6.0 Available for Installation

Linux , OpenMandriva , Plasma

The latest release of OpenMandriva has arrived with a new kernel, an updated Plasma desktop, and a server edition.
TrueNAS 25.04 Arrives with Thousands of Changes

Linux , Storage , TrueNAS

One of the most popular Linux-based NAS solutions has rolled out the latest edition, based on Ubuntu 25.04.
Fedora 42 Available with Two New Spins

Fedora , Gnome , Plasma

The latest release from the Fedora Project includes the usual updates, a new kernel, an official KDE Plasma spin, and a new System76 spin.
So Long, ArcoLinux

Linux , open source , Operating Systems

The ArcoLinux distribution is the latest Linux distribution to shut down.
What Open Source Pros Look for in a Job Role

FOSS , open source

Learn what professionals in technical and non-technical roles say is most important when seeking a new position.
Asahi Linux Runs into Issues with M4 Support

Linux , open source

Due to Apple Silicon changes, the Asahi Linux project is at odds with adding support for the M4 chips.
Plasma 6.3.4 Now Available

KDE , Linux , Plasma

Although not a major release, Plasma 6.3.4 does fix some bugs and offer a subtle change for the Plasma sidebar.
Linux Kernel 6.15 First Release Candidate Now Available

Kernel , Linux

Linux Torvalds has announced that the release candidate for the final release of the Linux 6.15 series is now available.
Akamai Will Host kernel.org

Kernel , Linux , Security

The organization dedicated to cloud-based solutions has agreed to host kernel.org to deliver long-term stability for the development team.
Linux Kernel 6.14 Released

Kernel , Linux , Rust

The latest Linux kernel has arrived with extra Rust support and more.

Scraping the web for data

Data Harvesting

Caveats

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters

Support Our Work

News

OpenMandriva Lx 6.0 Available for Installation

TrueNAS 25.04 Arrives with Thousands of Changes

Fedora 42 Available with Two New Spins

So Long, ArcoLinux

What Open Source Pros Look for in a Job Role

Asahi Linux Runs into Issues with M4 Support

Plasma 6.3.4 Now Available

Linux Kernel 6.15 First Release Candidate Now Available

Akamai Will Host kernel.org

Linux Kernel 6.14 Released

Scraping the web for data

Data Harvesting

Caveats

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters Find Linux and Open Source Jobs Subscribe to our ADMIN Newsletters

Support Our Work

News

Tag Cloud

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters