Google数据存储区 - Blob或文本

Key*_*yur 6 google-app-engine google-cloud-datastore

坚持在谷歌数据存储大串的2种可行方法TextBlob数据类型.

从存储消费的角度来看,推荐哪两个?从protobuf序列化和反序列化的角度来看同样的问题.

Dav*_*ill 4

两者之间没有显着的性能差异 - 只需使用最适合您的数据的一种即可。 BlobProperty应该用于存储二进制数据(例如,str 对象),而TextProperty应该用于存储任何文本数据(例如,unicodestr对象)。请注意,如果将 a 存储str在 a 中TextProperty,则它必须仅包含 ASCII 字节(小于十六进制 80 或十进制 128)(与 不同BlobProperty)。

正如您在源代码UnindexedProperty中看到的那样,这两个属性均派生自。

下面是一个示例应用程序,它演示了这些 ASCII 或 UTF-8 字符串的存储开销没有差异:

import struct

from google.appengine.ext import db, webapp
from google.appengine.ext.webapp.util import run_wsgi_app

class TestB(db.Model):
    v = db.BlobProperty(required=False)

class TestT(db.Model):
    v = db.TextProperty(required=False)

class MainPage(webapp.RequestHandler):
    def get(self):
        self.response.headers['Content-Type'] = 'text/plain'

        # try simple ASCII data and a bytestring with non-ASCII bytes
        ascii_str = ''.join([struct.pack('>B', i) for i in xrange(128)])
        arbitrary_str = ''.join([struct.pack('>2B', 0xC2, 0x80+i) for i in xrange(64)])
        u = unicode(arbitrary_str, 'utf-8')

        t = [TestT(v=ascii_str), TestT(v=ascii_str*1000), TestT(v=u*1000)]
        b = [TestB(v=ascii_str), TestB(v=ascii_str*1000), TestB(v=arbitrary_str*1000)]

        # demonstrate error cases
        try:
            err = TestT(v=arbitrary_str)
            assert False, "should have caused an error: can't store non-ascii bytes in a Text"
        except UnicodeDecodeError:
            pass
        try:
            err = TestB(v=u)
            assert False, "should have caused an error: can't store unicode in a Blob"
        except db.BadValueError:
            pass

        # determine the serialized size of each model (note: no keys assigned)
        fEncodedSz = lambda o : len(db.model_to_protobuf(o).Encode())
        sz_t = tuple([fEncodedSz(x) for x in t])
        sz_b = tuple([fEncodedSz(x) for x in b])

        # output the results
        self.response.out.write("text:   1=>%dB  2=>%dB  3=>%dB\n" % sz_t)
        self.response.out.write("blob:   1=>%dB  2=>%dB  3=>%dB\n" % sz_b)

application = webapp.WSGIApplication([('/', MainPage)])
def main(): run_wsgi_app(application)
if __name__ == '__main__': main()
Run Code Online (Sandbox Code Playgroud)

这是输出:

text:   1=>172B  2=>128047B  3=>128047B
blob:   1=>172B  2=>128047B  3=>128047B
Run Code Online (Sandbox Code Playgroud)

  • 我不知道 Text 属性只能包含 ASCII 字节。这种认识回答了我的问题。谢谢。 (2认同)
  • 这不是真的 - 文本属性存储 unicode。但是,如果您将字节(“原始”)字符串(类型“str”)分配给文本属性,它将尝试转换为 unicode,它使用系统默认编码,即 ASCII。如果您想这样做,则需要显式解码字符串。 (2认同)