如何从 Google 的 AudioSet 中提取音频嵌入(特征)?

Jan*_*ska 1 python protocol-buffers tensorflow

I\xe2\x80\x99m 谈论https://research.google.com/audioset/download.html上提供的音频特征数据集上提供的音频特征数据集,作为由帧级音频 tfrecord 组成的 tar.gz 存档。

\n\n

从 tfrecord 文件中提取其他所有内容都可以正常工作(我可以提取键:video_id、start_time_seconds、end_time_seconds、标签),但训练所需的实际嵌入似乎根本不存在。当我迭代数据集中任何 tfrecord 文件的内容时,仅打印四个键 video_id、start_time_seconds、end_time_seconds 和 labels。

\n\n

这是我正在使用的代码:

\n\n
import tensorflow as tf\nimport numpy as np\n\ndef readTfRecordSamples(tfrecords_filename):\n\n    record_iterator = tf.python_io.tf_record_iterator(path=tfrecords_filename)\n\n    for string_record in record_iterator:\n        example = tf.train.Example()\n        example.ParseFromString(string_record)\n        print(example)  # this prints the abovementioned 4 keys but NOT audio_embeddings\n\n        # the first label can be then parsed like this:\n        label = (example.features.feature[\'labels\'].int64_list.value[0])\n        print(\'label 1: \' + str(label))\n\n        # this, however, does not work:\n        #audio_embedding = (example.features.feature[\'audio_embedding\'].bytes_list.value[0])\n\nreadTfRecordSamples(\'embeddings/01.tfrecord\')\n
Run Code Online (Sandbox Code Playgroud)\n\n

有什么技巧可以提取 128 维嵌入吗?\n或者它们真的不在此数据集中吗?

\n

Jan*_*ska 5

解决了,tfrecord 文件需要作为序列示例而不是示例来读取。如果行以上代码有效

example = tf.train.Example()
Run Code Online (Sandbox Code Playgroud)

被替换为

example = tf.train.SequenceExample()
Run Code Online (Sandbox Code Playgroud)

然后只需运行即可查看嵌入和所有其他内容

print(example)
Run Code Online (Sandbox Code Playgroud)

  • 谢谢,这非常有帮助。您还能澄清一下如何获取实际数字而不是字节字符串吗? (2认同)