Notice: This page requires JavaScript to function properly.
Please enable JavaScript in your browser settings or update your browser.
学ぶ Poor Data Presentation | Preprocessing Data: Part I
Data Manipulation using pandas

bookPoor Data Presentation

メニューを表示するにはスワイプしてください

One of the reasons that can cause types inconsistency may be poor data presentation. For instance, values of weight column may have also the measurment unit (like, 25kg, 14lb). In this case Python will understand these values as strings.

Let's see what is wrong with the values in the columns we considered to have wrong type.

1234567
# Importing the library import pandas as pd # Reading the file df = pd.read_csv('https://codefinity-content-media.s3.eu-west-1.amazonaws.com/f2947b09-5f0d-4ad9-992f-ec0b87cd4b3f/data.csv') # Output values of 'problematic' columns print(df.loc[:,['totinch', 'morgh', 'valueh', 'grosrth', 'omphtotinch']])
copy

We found the root of the problem. All the columns but 'totinch' use dots . as indicator for missing values, while values in the 'totinch' column use commas , as the decimal separator. This may happen due to data origin, for instance. This problem can be solved by replacing commas with dots, and converting to float type.

Note, if you try to convert existing values into numeric type, then the error ValueError will be raised.

すべて明確でしたか?

どのように改善できますか?

フィードバックありがとうございます!

セクション 1.  3

AIに質問する

expand

AIに質問する

ChatGPT

何でも質問するか、提案された質問の1つを試してチャットを始めてください

セクション 1.  3
some-alt