beg*_*er_ 6 python numpy bitstring pandas
我有一个包含1列的pandas Dataframe,其中包含一串位,例如.'100100101'.我想将此字符串转换为numpy数组.
我怎样才能做到这一点?
编辑:
运用
features = df.bit.apply(lambda x: np.array(list(map(int,list(x)))))
#...
model.fit(features, lables)
Run Code Online (Sandbox Code Playgroud)
导致错误model.fit:
ValueError: setting an array element with a sequence.
Run Code Online (Sandbox Code Playgroud)
由于有明确的答案,我想出的解决方案适用于我的案例:
for bitString in input_table['Bitstring'].values:
bits = np.array(map(int, list(bitString)))
featureList.append(bits)
features = np.array(featureList)
#....
model.fit(features, lables)
Run Code Online (Sandbox Code Playgroud)
jed*_*rds 14
对于字符串s = "100100101",您可以至少以两种不同的方式将其转换为numpy数组.
第一个是使用numpy的fromstring方法.这有点尴尬,因为你必须指定数据类型并减去元素的"基数"值.
import numpy as np
s = "100100101"
a = np.fromstring(s,'u1') - ord('0')
print a # [1 0 0 1 0 0 1 0 1]
Run Code Online (Sandbox Code Playgroud)
'u1'数据类型在哪里,ord('0')用于从每个元素中减去"基数"值.
第二种方法是将每个字符串元素转换为整数(因为字符串是可迭代的),然后将该列表传递到np.array:
import numpy as np
s = "100100101"
b = np.array(map(int, s))
print b # [1 0 0 1 0 0 1 0 1]
Run Code Online (Sandbox Code Playgroud)
然后
# To see its a numpy array:
print type(a) # <type 'numpy.ndarray'>
print a[0] # 1
print a[1] # 0
# ...
Run Code Online (Sandbox Code Playgroud)
注意,随着输入字符串的长度s增加,第二种方法比第一种方法显着更差.对于小字符串,它很接近,但考虑timeit90个字符的字符串的结果(我刚才使用s * 10):
fromstring: 49.283392424 s
map/array: 2.154540959 s
Run Code Online (Sandbox Code Playgroud)
(这是使用默认timeit.repeat参数,最少3次运行,每次运行计算运行1M字符串 - >数组转换的时间)