Tom*_*bet 5 python surf neural-network keras tensorflow
我正在尝试将 Keras 与 TensorFlow 结合使用,以根据我从多张图像中获得的 SURF 特征来训练网络。我将所有这些功能存储在一个包含以下列的 CSV 文件中:
[ID, Code, PointX, PointY, Desc1, ..., Desc64]
Run Code Online (Sandbox Code Playgroud)
“ID”列是我存储所有值时由熊猫创建的自动增量索引。“代码”列是点的标签,这只是我通过将实际代码(它是一个字符串)与一个数字配对得到的一个数字。“PointX/Y”是在给定类的图像中找到的点的坐标,“Desc#”是该点对应描述符的浮点值。
CSV 文件包含在所有 20.000 个图像中找到的所有关键点和描述符。这使我的磁盘总大小接近 60GB,显然我无法放入内存中。
我一直在尝试使用 Pandas 批量加载文件,然后将所有值放在一个 numpy 数组中,然后拟合我的模型(只有 3 层的序列模型)。我使用以下代码来做到这一点:
chunksize = 10 ** 6
for chunk in pd.read_csv("surf_kps.csv", chunksize=chunksize):
dataset_chunk = chunk.to_numpy(dtype=np.float32, copy=False)
print(dataset_chunk)
# Divide dataset in data and labels
X = dataset_chunk[:,9:]
Y = dataset_chunk[:,1]
# Train model
model.fit(x=X,y=Y,batch_size=200,epochs=20)
# Evaluate model
scores = model.evaluate(X, Y)
print("\n%s: %.2f%%" % (model.metrics_names[1], scores[1]*100))
Run Code Online (Sandbox Code Playgroud)
这在加载第一个块时没问题,但是当循环获取另一个块时,准确性和损失卡在 0 上。
我试图加载所有这些信息的方式是错误的吗?
提前致谢!
- - - 编辑 - - -
好的,现在我制作了一个简单的生成器,如下所示:
def read_csv(filename):
with open(filename, 'r') as f:
for line in f.readlines():
record = line.rstrip().split(',')
features = [np.float32(n) for n in record[9:73]]
label = int(record[1])
print("features: ",type(features[0]), " ", type(label))
yield np.array(features), label
Run Code Online (Sandbox Code Playgroud)
并使用 fit_generator :
tf_ds = read_csv("mini_surf_kps.csv")
model.fit_generator(tf_ds,steps_per_epoch=1000,epochs=20)
Run Code Online (Sandbox Code Playgroud)
我不知道为什么,但在第一个纪元开始之前我一直收到错误消息:
ValueError: Error when checking input: expected dense_input to have shape (64,) but got array with shape (1,)
Run Code Online (Sandbox Code Playgroud)
模型的第一层有input_dim=64,生成的特征数组的形状也是 64。
| 归档时间: |
|
| 查看次数: |
5853 次 |
| 最近记录: |