当stream = True但数据并不总是流入时,如何退出Python请求?

Zac*_*lor 6 python qthread python-requests

我正在使用请求在网页上发出获取请求,其中当现实世界中发生事件时添加新数据。只要窗口打开,我就想继续获取这些数据,因此我设置数据stream = True,然后在数据流入时逐行迭代。

page = requests.get(url, headers=headers, stream=True)
# Process the LiveLog data until stopped from exterior source
for html_line in page.iter_lines(chunk_size=1):
    # Do other work here
Run Code Online (Sandbox Code Playgroud)

我对这部分没有问题,但是当涉及到退出这个循环时我遇到了问题。通过查看其他 StackOverflow 线程,我了解到我无法捕获任何信号,因为我的 for 循环被阻塞。相反,我尝试使用以下代码,该代码确实有效,但有一个大问题。

if QThread.currentThread().isInterruptionRequested():
    break
Run Code Online (Sandbox Code Playgroud)

这段代码将使我摆脱循环,但我发现 for 循环迭代的唯一时间是在将新数据引入 get 时,而在我的情况下,这不是连续的。我可以在几分钟或更长时间内没有任何新数据,并且不想在再次执行循环以检查是否请求中断之前必须等待这些新数据落地。

如何在用户操作后立即退出循环?

小智 4

您可以尝试 aiohttp 库https://github.com/aio-libs/aiohttp,特别是https://aiohttp.readthedocs.io/en/stable/streams.html#asynchronous-iteration-support。它看起来像:

import asyncio
import aiohttp

async def main():
    url = 'https://httpbin.org/stream/20'
    chunk_size = 1024
    async with aiohttp.ClientSession() as session:
        async with session.get(url) as resp:
            while True:
                data = await resp.content.readline():
                print(data) # do work here

if __name__ == "__main__":
    asyncio.run(main())
Run Code Online (Sandbox Code Playgroud)

值得注意的是,这resp.content是一个StreamReader,因此您可以使用其他可用的方法https://aiohttp.readthedocs.io/en/stable/streams.html#aiohttp.StreamReader