使用pandas在HDF中存储包含字符串的数据帧时的谜团

Art*_* B. 7 python hdf5 pandas

这是万圣节大熊猫和HDF的幽灵:

df = pandas.DataFrame([['a','b'] for i in range(1,1000)])
store = pandas.HDFStore('test.h5')
store['x'] = df
store.close()
Run Code Online (Sandbox Code Playgroud)

然后

ls -l test.h5
-rw-r--r-- 1 arthur arthur 1072080 Oct 26 10:50 test.h5
Run Code Online (Sandbox Code Playgroud)

1.1M?有点陡峭但为什么不呢.这里的事情变得非常怪异

store = pandas.HDFStore('test.h5') #open it again
store['x'] = df #do the same thing as before!
store.close()
Run Code Online (Sandbox Code Playgroud)

然后

ls -l test.h5
-rw-r--r-- 1 arthur arthur 2122768 Oct 26 10:52 test.h5
Run Code Online (Sandbox Code Playgroud)

您现在已进入Twilight区域.毋庸置疑,操作后商店无法区分,但每次迭代都会使文件更加肥胖.

似乎只有在涉及到字符串时才会发生.在我提交错误报告之前,我想知道我在这里遗漏了什么......

Art*_* B. 4

看来这可能是原因:http ://www.hdfgroup.org/hdf5-quest.html#del

这是 HDF5 的一个大问题,wtf。