如何用pandas get_dummies函数消除键错误

问题描述 投票:1回答:2

当我运行pandas get_dummies()函数时,它会返回一个keyerror,指出我的所有列都不存在。以下代码使用受版权保护的数据,我引用它:UCI机器学习库的成人数据集引用了Dua,D。和Graff,C。(2019)。 UCI机器学习库[http://archive.ics.uci.edu/ml]。加利福尼亚州欧文市:加州大学信息与计算机科学学院。

我不确定该尝试什么。

age, workclass, fnlwgt, education, education-num, marital-status, occupation, forces, relationship, race, sex, capital-gain, capital-loss, hours-per-week, native-country,
39, State-gov, 77516, Bachelors, 13, Never-married, Adm-clerical, Not-in-family, White, Male, 2174, 0, 40, United-States, <=50K
50, Self-emp-not-inc, 83311, Bachelors, 13, Married-civ-spouse, Exec-managerial, Husband, White, Male, 0, 0, 13, United-States, <=50K
38, Private, 215646, HS-grad, 9, Divorced, Handlers-cleaners, Not-in-family, White, Male, 0, 0, 40, United-States, <=50K
53, Private, 234721, 11th, 7, Married-civ-spouse, Handlers-cleaners, Husband, Black, Male, 0, 0, 40, United-States, <=50K
28, Private, 338409, Bachelors, 13, Married-civ-spouse, Prof-specialty, Wife, Black, Female, 0, 0, 40, Cuba, <=50K
37, Private, 284582, Masters, 14, Married-civ-spouse, Exec-managerial, Wife, White, Female, 0, 0, 40, United-States, <=50K
49, Private, 160187, 9th, 5, Married-spouse-absent, Other-service, Not-in-family, Black, Female, 0, 0, 16, Jamaica, <=50K
52, Self-emp-not-inc, 209642, HS-grad, 9, Married-civ-spouse, Exec-managerial, Husband, White, Male, 0, 0, 45, United-States, >50K
#import modules
import pandas as pd

#define functions
def open_infile():
    d = pd.read_csv('adult.data.txt', sep = ',')
    return d

def onehot_encode(data):
    data = pd.get_dummies(data, columns = ['workclass', 'education', 'marital-status', 'occupation', 'forces',
                                         'relationship', 'race', 'sex', 'native-country'])
    return data
##########gather data##########
#opoen infile
data = open_infile()
print(len(data))

##########process data##########
#one-hot encode categorical columns
onehot_encode(data)
print(data.head())
Traceback (most recent call last):
  File "C:/Users/Hezekiah/PycharmProjects/Artificial Intelligence 0/Chapter 1 Application Adult.py", line 20, in <module>
    onehot_encode(data)
  File "C:/Users/Hezekiah/PycharmProjects/Artificial Intelligence 0/Chapter 1 Application Adult.py", line 11, in onehot_encode
    'relationship', 'race', 'sex', 'native-country'])
  File "C:\Users\Hezekiah\PycharmProjects\Artificial Intelligence 0\venv\lib\site-packages\pandas\core\reshape\reshape.py", line 812, in get_dummies
    data_to_encode = data[columns]
  File "C:\Users\Hezekiah\PycharmProjects\Artificial Intelligence 0\venv\lib\site-packages\pandas\core\frame.py", line 2934, in __getitem__
    raise_missing=True)
  File "C:\Users\Hezekiah\PycharmProjects\Artificial Intelligence 0\venv\lib\site-packages\pandas\core\indexing.py", line 1354, in _convert_to_indexer
    return self._get_listlike_indexer(obj, axis, **kwargs)[1]
  File "C:\Users\Hezekiah\PycharmProjects\Artificial Intelligence 0\venv\lib\site-packages\pandas\core\indexing.py", line 1161, in _get_listlike_indexer
    raise_missing=raise_missing)
  File "C:\Users\Hezekiah\PycharmProjects\Artificial Intelligence 0\venv\lib\site-packages\pandas\core\indexing.py", line 1246, in _validate_read_indexer
    key=key, axis=self.obj._get_axis_name(axis)))
KeyError: "None of [Index(['workclass', 'education', 'marital-status', 'occupation', 'forces',\n       'relationship', 'race', 'sex', 'native-country'],\n      dtype='object')] are in the [columns]"

我希望pandas get_dummies()函数将所有分类属性转换为数字属性,但是pycharm会返回一个keyerror,它告诉我没有任何列存在,显然它们确实存在。

python-3.x pandas keyerror
2个回答
1
投票

列名称中的尾随空格存在问题,解决方法是使用str.strip

data.columns = data.columns.str.strip()

或者用strip列出理解:

data.columns = [x.strip() for x in data.columns]

1
投票

你的主要问题是将adult.namesadult.data文件合并时的数据你提到的网站数据中没有强制列。如果你正确合并数据,你也不会得到这个error

即使你正在使用这个专栏来制作假人。

© www.soinside.com 2019 - 2024. All rights reserved.