将Bitstring(1和0的字符串)转换为numpy数组

beg*_*er_ 6 python numpy bitstring pandas

我有一个包含1列的pandas Dataframe,其中包含一串位,例如.'100100101'.我想将此字符串转换为numpy数组.

我怎样才能做到这一点?

编辑:

运用

features = df.bit.apply(lambda x: np.array(list(map(int,list(x)))))
#...
model.fit(features, lables)
Run Code Online (Sandbox Code Playgroud)

导致错误model.fit:

ValueError: setting an array element with a sequence.
Run Code Online (Sandbox Code Playgroud)

由于有明确的答案,我想出的解决方案适用于我的案例:

for bitString in input_table['Bitstring'].values:
    bits = np.array(map(int, list(bitString)))
    featureList.append(bits)
features = np.array(featureList)
#....
model.fit(features, lables)
Run Code Online (Sandbox Code Playgroud)

jed*_*rds 14

对于字符串s = "100100101",您可以至少以两种不同的方式将其转换为numpy数组.

第一个是使用numpy的fromstring方法.这有点尴尬,因为你必须指定数据类型并减去元素的"基数"值.

import numpy as np

s = "100100101"
a = np.fromstring(s,'u1') - ord('0')

print a  # [1 0 0 1 0 0 1 0 1]
Run Code Online (Sandbox Code Playgroud)

'u1'数据类型在哪里,ord('0')用于从每个元素中减去"基数"值.

第二种方法是将每个字符串元素转换为整数(因为字符串是可迭代的),然后将该列表传递到np.array:

import numpy as np

s = "100100101"
b = np.array(map(int, s))

print b  # [1 0 0 1 0 0 1 0 1]
Run Code Online (Sandbox Code Playgroud)

然后

# To see its a numpy array:
print type(a)  # <type 'numpy.ndarray'>
print a[0]     # 1
print a[1]     # 0
# ...
Run Code Online (Sandbox Code Playgroud)

注意,随着输入字符串的长度s增加,第二种方法比第一种方法显着更差.对于小字符串,它很接近,但考虑timeit90个字符的字符串的结果(我刚才使用s * 10):

fromstring: 49.283392424 s
map/array:   2.154540959 s
Run Code Online (Sandbox Code Playgroud)

(这是使用默认timeit.repeat参数,最少3次运行,每次运行计算运行1M字符串 - >数组转换的时间)

  • 注意``np.array(map(int,s))`就足够了 - 首先不需要构建`list` ...而且,它不是很直观,但是`np.fromstring(s,'i1) ') - 48`快了大约50%...... (3认同)
  • @JonClements 我不认为这在 Python 3.x 中仍然成立。现在 map 返回一个映射对象(迭代器),您必须将其包装在 `list` 中或使用 `np.fromiter(map(int, s))` (2认同)