如何以1000块的形式阅读集合?

Dam*_*mir 9 python mongodb pymongo

我需要在Python代码中读取MongoDB的整个集合(集合名称为"test").我尝试过

    self.__connection__ = Connection('localhost',27017)
    dbh = self.__connection__['test_db']            
    collection = dbh['test']
Run Code Online (Sandbox Code Playgroud)

如何通过1000读取块中的集合(以避免内存溢出,因为集合可能非常大)?

Mic*_*ear 6

我同意雷蒙的意见,但是你提到了1000批,他的回答并没有真正涵盖.您可以在光标上设置批量大小:

cursor.batch_size(1000);
Run Code Online (Sandbox Code Playgroud)

您还可以跳过记录,例如:

cursor.skip(4000);
Run Code Online (Sandbox Code Playgroud)

这是你在找什么?这实际上是一种分页模式.但是,如果您只是想避免内存耗尽,那么您实际上不需要设置批量大小或跳过.


ale*_*kov 6

受到@Rafael Valero + 修复他代码中最后一个块错误并使其更通用的启发,我创建了生成器函数以通过查询和投影遍历 mongo 集合:

def iterate_by_chunks(collection, chunksize=1, start_from=0, query={}, projection={}):
   chunks = range(start_from, collection.find(query).count(), int(chunksize))
   num_chunks = len(chunks)
   for i in range(1,num_chunks+1):
      if i < num_chunks:
          yield collection.find(query, projection=projection)[chunks[i-1]:chunks[i]]
      else:
          yield collection.find(query, projection=projection)[chunks[i-1]:chunks.stop]
Run Code Online (Sandbox Code Playgroud)

例如,您首先创建一个这样的迭代器:

mess_chunk_iter = iterate_by_chunks(db_local.conversation_messages, 200, 0, query={}, projection=projection)
Run Code Online (Sandbox Code Playgroud)

然后分块迭代:

chunk_n=0
total_docs=0
for docs in mess_chunk_iter:
   chunk_n=chunk_n+1        
   chunk_len = 0
   for d in docs:
      chunk_len=chunk_len+1
      total_docs=total_docs+1
   print(f'chunk #: {chunk_n}, chunk_len: {chunk_len}')
print("total docs iterated: ", total_docs)

chunk #: 1, chunk_len: 400
chunk #: 2, chunk_len: 400
chunk #: 3, chunk_len: 400
chunk #: 4, chunk_len: 400
chunk #: 5, chunk_len: 400
chunk #: 6, chunk_len: 400
chunk #: 7, chunk_len: 281
total docs iterated:  2681
Run Code Online (Sandbox Code Playgroud)


Rem*_*iet 5

使用游标.游标有一个"batchSize"变量,用于控制在执行查询后每批实际发送到客户端的文档数.您不必触摸此设置,因为默认设置很好,并且在大多数驱动程序中隐藏了调用"getmore"命令的复杂性.我不熟悉pymongo,但它的工作原理如下:

cursor = db.col.find() // Get everything!

while(cursor.hasNext()) {
    /* This will use the documents already fetched and if it runs out of documents in it's local batch it will fetch another X of them from the server (where X is batchSize). */
    document = cursor.next();

    // Do your magic here
}
Run Code Online (Sandbox Code Playgroud)

  • 你是如何用Python做到的? (3认同)