Clean and Convert

We have decided to create a complete chapter on the topic of text cleaning and preprocessing. As you may imagine, complete texts cannot be directly fed into an ML model. For this reason, we will apply specific preprocessing techniques.

The first step will be to remove punctuation from our column to reduce noise in our data. We will do this using regex (regular expression matching operations).

Then we will vectorize our text. Refer to the picture below for more information. Essentially, we will represent words, sentences, or even larger units of text as vectors.

Task

Swipe to start coding

Remove punctuaction with regex by using the appropriate method to replace the given pattern with an empty string.
Vectorize the texts of the articles.

Solution

Mark tasks as Completed

Switch to desktop for real-world practiceContinue from where you are using one of the options below

Everything was clear?

Thanks for your feedback!

Section 1. Chapter 4

Ask AI

Ask anything or try one of the suggested questions to begin our chat

Course Content

Identifying Fake News

Introduction True News and Fake News Data Preprocessing Clean and Convert Initial Model Fit Decision Tree Comparison Fake News Tool (Bonus)