Notice: This page requires JavaScript to function properly.
Please enable JavaScript in your browser settings or update your browser.
Lære First Steps | HTML Files and DevTools
Web Scraping with Python

Stryg for at vise menuen

book
First Steps

The web page is comprised of HTML.

HTML is the markup language for creating web pages.

Let’s get the HTML file of the web page!

We can work with the data of sites by their URLs. To open website URLs in your Python programs, use the function urlopen() from the module urllib.request and define the URL you want to open as a string variable:

1234
from urllib.request import urlopen url = "https://codefinity-content-media.s3.eu-west-1.amazonaws.com/18a4e428-1a0f-44c2-a8ad-244cd9c7985e/page.html" page = urlopen(url)
copy

It looks simple, but when we want to print the variable page to see what is going on in the HTML file, we get:

python

It returns HTTPResponse object. To parse it, use the .read() method, which returns a sequence of bytes, and then the function decode("utf-8") to decode the data from bytes to string

1234
bytes = page.read() html = bytes.decode("utf-8") print(html)
copy

We can also use methods consequentially: page.read().decode("utf-8").

Here we open the URL we will work within this course!

Opgave

Swipe to start coding

Write the missing code to get the HTML structure from the page which interests you.

  1. Import module urlopen to open URLs from your code.
  2. Open the URL. Assign the result to the variable page.
  3. Get a sequence of bytes using the method .read(). Assign the result to the variable bytes.
  4. Decode bytes to string using the method .decode(). Assign the result to the variable html.

Løsning

Switch to desktopSkift til skrivebord for at øve i den virkelige verdenFortsæt der, hvor du er, med en af nedenstående muligheder
Var alt klart?

Hvordan kan vi forbedre det?

Tak for dine kommentarer!

Sektion 1. Kapitel 2
Vi beklager, at noget gik galt. Hvad skete der?

Spørg AI

expand
ChatGPT

Spørg om hvad som helst eller prøv et af de foreslåede spørgsmål for at starte vores chat

book
First Steps

The web page is comprised of HTML.

HTML is the markup language for creating web pages.

Let’s get the HTML file of the web page!

We can work with the data of sites by their URLs. To open website URLs in your Python programs, use the function urlopen() from the module urllib.request and define the URL you want to open as a string variable:

1234
from urllib.request import urlopen url = "https://codefinity-content-media.s3.eu-west-1.amazonaws.com/18a4e428-1a0f-44c2-a8ad-244cd9c7985e/page.html" page = urlopen(url)
copy

It looks simple, but when we want to print the variable page to see what is going on in the HTML file, we get:

python

It returns HTTPResponse object. To parse it, use the .read() method, which returns a sequence of bytes, and then the function decode("utf-8") to decode the data from bytes to string

1234
bytes = page.read() html = bytes.decode("utf-8") print(html)
copy

We can also use methods consequentially: page.read().decode("utf-8").

Here we open the URL we will work within this course!

Opgave

Swipe to start coding

Write the missing code to get the HTML structure from the page which interests you.

  1. Import module urlopen to open URLs from your code.
  2. Open the URL. Assign the result to the variable page.
  3. Get a sequence of bytes using the method .read(). Assign the result to the variable bytes.
  4. Decode bytes to string using the method .decode(). Assign the result to the variable html.

Løsning

Switch to desktopSkift til skrivebord for at øve i den virkelige verdenFortsæt der, hvor du er, med en af nedenstående muligheder
Var alt klart?

Hvordan kan vi forbedre det?

Tak for dine kommentarer!

Sektion 1. Kapitel 2
Switch to desktopSkift til skrivebord for at øve i den virkelige verdenFortsæt der, hvor du er, med en af nedenstående muligheder
Vi beklager, at noget gik galt. Hvad skete der?
some-alt