使用Python并行进行多个API调用(IPython)

use*_*289 8 python api parallel-processing

我在我的本地机器(Mac)上使用Python(IPython和Canopy)和RESTful内容API.

我有一个由3000个唯一ID组成的数组,用于从API中提取数据,并且一次只能使用一个ID调用API.

我希望能以某种方式同时制作3组1000个电话以加快速度.

这样做的最佳方式是什么?

在此先感谢您的帮助!

min*_*nrk 18

如果没有关于您正在做什么的更多信息,很难肯定地说,但简单的线程方法可能有意义.

假设您有一个处理单个ID的简单函数:

import requests

url_t = "http://localhost:8000/records/%i"

def process_id(id):
    """process a single ID"""
    # fetch the data
    r = requests.get(url_t % id)
    # parse the JSON reply
    data = r.json()
    # and update some data with PUT
    requests.put(url_t % id, data=data)
    return data
Run Code Online (Sandbox Code Playgroud)

您可以将其扩展为处理一系列ID的简单函数:

def process_range(id_range, store=None):
    """process a number of ids, storing the results in a dict"""
    if store is None:
        store = {}
    for id in id_range:
        store[id] = process_id(id)
    return store
Run Code Online (Sandbox Code Playgroud)

最后,您可以相当轻松地将子范围映射到线程上,以允许一些请求并发:

from threading import Thread

def threaded_process_range(nthreads, id_range):
    """process the id range in a specified number of threads"""
    store = {}
    threads = []
    # create the threads
    for i in range(nthreads):
        ids = id_range[i::nthreads]
        t = Thread(target=process_range, args=(ids,store))
        threads.append(t)

    # start the threads
    [ t.start() for t in threads ]
    # wait for the threads to finish
    [ t.join() for t in threads ]
    return store
Run Code Online (Sandbox Code Playgroud)

IPython笔记本中的完整示例:http://nbviewer.ipython.org/5732094

如果您的个人任务花费的时间更广泛,您可能需要使用ThreadPool,它将一次分配一个作业(如果个别任务非常小,通常会更慢,但在异质情况下保证更好的平衡).

  • 这意味着迈步.当你指定一个切片时,有三个数字:`start:stop:stride`.所以`1 :: 3`表示每个第三个元素,从1开始,即`[1,4,7,...]`.这只是对列表进行同等分区的简单方法. (2认同)
  • 所以双冒号只是意味着未指定停止点,并且默认为“结束”。 (2认同)