使用pandas读取csv时设置列类型

use*_*815 9 python csv dictionary types pandas

尝试使用以下格式将csv文件读入pandas数据帧

dp = pd.read_csv('products.csv', header = 0,  dtype = {'name': str,'review': str,
                                                      'rating': int,'word_count': dict}, engine = 'c')
print dp.shape
for col in dp.columns:
    print 'column', col,':', type(col[0])
print type(dp['rating'][0])
dp.head(3)
Run Code Online (Sandbox Code Playgroud)

这是输出:

(183531, 4)
column name : <type 'str'>
column review : <type 'str'>
column rating : <type 'str'>
column word_count : <type 'str'>
<type 'numpy.int64'>
Run Code Online (Sandbox Code Playgroud)

在此输入图像描述

我可以理解,大熊猫可能会发现很难将字典的字符串表示转换为字典并给出这个和这个.但是如何将"rating"列的内容同时为str和numpy.int64 ???

顺便说一下,不指定引擎或标题的调整不会改变任何东西.

感谢致敬

Col*_*vel 6

在你的循环中,你正在做:

for col in dp.columns:
    print 'column', col,':', type(col[0])
Run Code Online (Sandbox Code Playgroud)

并且您str在任何地方都正确地看到了输出,因为col[0]是列名称的第一个字母,它是一个字符串。

例如,如果您运行此循环:

for col in dp.columns:
    print 'column', col,':', col[0]
Run Code Online (Sandbox Code Playgroud)

您将看到打印出每个列名称的字符串的第一个字母 - 这col[0]就是。

您的循环只迭代列名,而不是系列数据。

您真正想要的是在循环中检查每列数据的类型(不是其标题或部分标题)。

所以这样做是为了获取列数据的类型(非标题数据):

for col in dp.columns:
    print 'column', col,':', type(dp[col][0])
Run Code Online (Sandbox Code Playgroud)

这与您在rating单独打印列类型时所做的类似。


Mik*_*ler 6

使用:

dp.info()
Run Code Online (Sandbox Code Playgroud)

查看列的数据类型。dp.columns指的是列标题名称,它是字符串。