从csv-file读取数据并转换为正确的数据类型

wew*_*ewa 18 python csv python-2.5

我有以下问题.我写了一个二维列表,其中每列具有不同的类型(bool,str,int,list),到csv文件.现在我想再次从csv文件中读出数据.但我读到的每个单元都被解释为一个字符串.

如何自动将读入数据转换为正确的类型?或者更好:是否有可能告诉csv-reader每列的正确数据类型?

示例数据(如csv文件中):

IsActive,Type,Price,States
True,Cellphone,34,"[1, 2]"
,FlatTv,3.5,[2]
False,Screen,100.23,"[5, 1]"
True,Notebook, 50,[1]
Run Code Online (Sandbox Code Playgroud)

cor*_*opy 13

正如文档所述,CSV阅读器不执行自动数据转换.您有QUOTE_NONNUMERIC格式选项,但这只会将所有非引用字段转换为浮点数.这与其他csv读者非常相似.

我不相信Python的csv模块对这种情况有任何帮助.正如其他人已经指出的那样,literal_eval()是一个更好的选择.

以下工作和转换:

  • 字符串
  • INT
  • 彩车
  • 名单
  • 字典

您也可以将它用于布尔值和NoneType,尽管这些必须相应地格式化literal_eval()才能通过.LibreOffice Calc以大写字母显示布尔值,而在Python布尔值为大写时.此外,你必须用None(不含引号)替换空字符串

我正在为mongodb写一个进口商来做这一切.以下是我到目前为止编写的代码的一部分.

[注意:我的csv使用tab作为字段分隔符.您可能还想添加一些异常处理]

def getFieldnames(csvFile):
    """
    Read the first row and store values in a tuple
    """
    with open(csvFile) as csvfile:
        firstRow = csvfile.readlines(1)
        fieldnames = tuple(firstRow[0].strip('\n').split("\t"))
    return fieldnames

def writeCursor(csvFile, fieldnames):
    """
    Convert csv rows into an array of dictionaries
    All data types are automatically checked and converted
    """
    cursor = []  # Placeholder for the dictionaries/documents
    with open(csvFile) as csvFile:
        for row in islice(csvFile, 1, None):
            values = list(row.strip('\n').split("\t"))
            for i, value in enumerate(values):
                nValue = ast.literal_eval(value)
                values[i] = nValue
            cursor.append(dict(zip(fieldnames, values)))
    return cursor
Run Code Online (Sandbox Code Playgroud)

  • 好的解决方案。所需模块:csv、ast 和 itertools。 (2认同)

小智 7

你必须映射你的行:

data = """True,foo,1,2.3,baz
False,bar,7,9.8,qux"""

reader = csv.reader(StringIO.StringIO(data), delimiter=",")
parsed = (({'True':True}.get(row[0], False),
           row[1],
           int(row[2]),
           float(row[3]),
           row[4])
          for row in reader)
for row in parsed:
    print row
Run Code Online (Sandbox Code Playgroud)

结果是

(True, 'foo', 1, 2.3, 'baz')
(False, 'bar', 7, 9.8, 'qux')
Run Code Online (Sandbox Code Playgroud)

  • 由于OP在他们的例子中有一个bool列.对于`row [0]`假设只有"True"为"True",你可以使用`{'True':True} .get(row [0],False)` (3认同)

mar*_*eau 7

我知道这是一个相当老的问题,标记为,但这里的答案适用于 Python 3.6+,使用该语言的最新版本的人们可能会感兴趣。

它利用了typing.NamedTuplePython 3.5 中添加的内置类。从文档中可能不明显的是,每个字段的“类型”可以是一个函数。

示例使用代码还使用了所谓的f-string文字,这些文字直到 Python 3.6 才被添加,但不需要使用它们来进行核心数据类型转换。

#!/usr/bin/env python3.6
import ast
import csv
from typing import NamedTuple


class Record(NamedTuple):
    """ Define the fields and their types in a record. """
    IsActive: bool
    Type: str
    Price: float
    States: ast.literal_eval  # Handles string represenation of literals.

    @classmethod
    def _transform(cls: 'Record', dict_: dict) -> dict:
        """ Convert string values in given dictionary to corresponding Record
            field type.
        """
        return {name: cls.__annotations__[name](value)
                    for name, value in dict_.items()}


filename = 'test_transform.csv'

with open(filename, newline='') as file:
    for i, row in enumerate(csv.DictReader(file)):
        row = Record._transform(row)
        print(f'row {i}: {row}')
Run Code Online (Sandbox Code Playgroud)

输出:

#!/usr/bin/env python3.6
import ast
import csv
from typing import NamedTuple


class Record(NamedTuple):
    """ Define the fields and their types in a record. """
    IsActive: bool
    Type: str
    Price: float
    States: ast.literal_eval  # Handles string represenation of literals.

    @classmethod
    def _transform(cls: 'Record', dict_: dict) -> dict:
        """ Convert string values in given dictionary to corresponding Record
            field type.
        """
        return {name: cls.__annotations__[name](value)
                    for name, value in dict_.items()}


filename = 'test_transform.csv'

with open(filename, newline='') as file:
    for i, row in enumerate(csv.DictReader(file)):
        row = Record._transform(row)
        print(f'row {i}: {row}')
Run Code Online (Sandbox Code Playgroud)

由于实现的方式,通过创建一个只包含泛型类方法的基类来概括这一点并不简单typing.NamedTuple。

为了避免这个问题,在 Python 3.7+ 中,dataclasses.dataclass可以使用 a 代替,因为它们没有继承问题——所以创建一个可以重用的通用基类很简单:

#!/usr/bin/env python3.7
import ast
import csv
from dataclasses import dataclass, fields
from typing import Type, TypeVar

T = TypeVar('T', bound='GenericRecord')

class GenericRecord:
    """ Generic base class for transforming dataclasses. """
    @classmethod
    def _transform(cls: Type[T], dict_: dict) -> dict:
        """ Convert string values in given dictionary to corresponding type. """
        return {field.name: field.type(dict_[field.name])
                    for field in fields(cls)}


@dataclass
class CSV_Record(GenericRecord):
    """ Define the fields and their types in a record.
        Field names must match column names in CSV file header.
    """
    IsActive: bool
    Type: str
    Price: float
    States: ast.literal_eval  # Handles string represenation of literals.


filename = 'test_transform.csv'

with open(filename, newline='') as file:
    for i, row in enumerate(csv.DictReader(file)):
        row = CSV_Record._transform(row)
        print(f'row {i}: {row}')
Run Code Online (Sandbox Code Playgroud)

从某种意义上说,使用哪个并不是很重要,因为从未创建过类的实例——使用一个实例只是在记录数据结构中指定和保存字段名称及其类型的定义的一种干净方式。

ATypedDict被添加到typingPython 3.8 中的模块中,它也可用于提供类型信息,但必须以稍微不同的方式使用,因为它实际上并没有像NamedTuple和dataclassesdo那样定义新类型——所以它需要有一个独立的转换功能:

#!/usr/bin/env python3.8
import ast
import csv
from dataclasses import dataclass, fields
from typing import TypedDict


def transform(dict_, typed_dict) -> dict:
    """ Convert values in given dictionary to corresponding types in TypedDict . """
    fields = typed_dict.__annotations__
    return {name: fields[name](value) for name, value in dict_.items()}


class CSV_Record_Types(TypedDict):
    """ Define the fields and their types in a record.
        Field names must match column names in CSV file header.
    """
    IsActive: bool
    Type: str
    Price: float
    States: ast.literal_eval


filename = 'test_transform.csv'

with open(filename, newline='') as file:
    for i, row in enumerate(csv.DictReader(file), 1):
        row = transform(row, CSV_Record_Types)
        print(f'row {i}: {row}')

Run Code Online (Sandbox Code Playgroud)