我试图用anaconda python 3.5从url中获取数据
import requests
url ='http://eutils.ncbi.nlm.nih.gov/entrez/eutils/einfo.fcgi'
r = requests.get(url)
r.content
Run Code Online (Sandbox Code Playgroud)
可以在浏览器中打开网址而不会出现问题...
但我收到一个错误(对于这个网址和我尝试的任何其他网址):
-------------------------------------------------- ------------------------ TypeError Traceback(最近一次调用最后一次)C:\ Anaconda3\lib\site-packages\requests\packages\urllib3\connectionpool _make_request中的.py(self,conn,method,url,timeout,**httplib_request_kw)375尝试:#Python 2.7,使用HTTP响应的缓冲 - > 376 httplib_response = conn.getresponse(buffering = True)377除TypeError外:## Python 2.6及更早版本
TypeError:getresponse()得到一个意外的关键字参数'buffering'
在处理上述异常期间,发生了另一个异常:
RemoteDisconnected Traceback(最近一次调用最后一次)在urlopen中的C:\ Anaconda3\lib\site-packages\requests\packages\urllib3\connectionpool.py(self,method,url,body,headers,retries,redirect,assert_same_host,timeout,pool_timeout ,release_conn,**response_kw)558 timeout = timeout_obj, - > 559 body = body,headers = headers)560
在_make_request(self,conn,method,url,timeout,**httplib_request_kw)中的C:\ Anaconda3\lib\site-packages\requests\packages\urllib3\connectionpool.py 377除TypeError外:#Python 2.6及更早版本 - > 378 httplib_response = conn.getresponse()379除了(SocketTimeout,BaseSSLError,SocketError)为e:
C:\ Anaconda3\lib\http\client.py在getresponse(self)1173
尝试: - > 1174 response.begin()1175除了ConnectionError:C:\ Anaconda3\lib\http\client.py in begin(self)281,True: - > 282 version,status,reason = self._read_status()283 if status!= CONTINUE:
C:\ Anaconda3\lib\http\client.py在_read_status(self)250#中发送有效回复. - …
我使用请求和线程编写了一个可暂停的多线程下载器,但是下载在恢复后无法完成,长话短说,由于特殊的网络条件,连接通常会在需要刷新连接的下载过程中终止。
您可以在我之前的问题中查看代码:
我观察到下载恢复后可以超过 100% 并且不会停止(至少我没有看到它们停止),mmap 索引将超出范围并出现大量错误消息......
我终于发现这是因为先前请求的幽灵,导致服务器错误地从上次连接发送了未下载的额外数据。
这是我的解决方案:
s = requests.session()
r = s.get(
url, headers={'connection': 'close', 'range': 'bytes={0}-{1}'.format(start, end)}, stream=True)
Run Code Online (Sandbox Code Playgroud)
r.close()
s.close()
del r
del s
Run Code Online (Sandbox Code Playgroud)
在我的测试中,我发现requests有两个名为session的属性,一个是Titlecase,一个是小写,小写的是一个函数,另一个是一个类构造函数,它们都创建了一个requests.sessions.Session对象,有没有他们之间的区别?
如何将 keep-alive 设置为 False?
这里找到的方法不再有效:
In [39]: s = requests.session()
...: s.config['keep_alive'] = False
---------------------------------------------------------------------------
AttributeError Traceback (most recent call last)
<ipython-input-39-497f569a91ba> in <module>
1 s = requests.session()
----> 2 s.config['keep_alive'] = False
AttributeError: 'Session' object has no attribute 'config'
Run Code Online (Sandbox Code Playgroud)
这里的方法不会抛出错误:
s = requests.session()
s.keep_alive = False …Run Code Online (Sandbox Code Playgroud) requests.exceptions.ConnectionError: ('Connection aborted.', error(99, 'Cannot assign requested address'))
Run Code Online (Sandbox Code Playgroud)
当使用python请求库运行多个进程并将post函数调用到返回非常快(<10ms)的API时,我收到此错误.
拨打正在运行的进程数会产生延迟效果,但只有拨入1进程才能解决问题.这不是一个解决方案,但确实表明有限的资源是罪魁祸首.
我正在使用 python requests 库请求具有此类代码的 API:
api_request = requests.get(f"http://data.api.org/search?q=example&ontologies=BFO&roots_only=true",
headers={'Authorization': 'apikey token=' + 'be03c61f-2ab8'})
api_result = api_request.json()
collection = api_result["collection"]
...
Run Code Online (Sandbox Code Playgroud)
当我不请求大量内容时,此代码工作正常,但否则我会收到错误。奇怪的是我每次请求很多内容都得不到。错误消息如下:
Traceback (most recent call last):
File "/home/nobu/.local/lib/python3.6/site-packages/urllib3/connection.py", line 160, in _new_conn
(self._dns_host, self.port), self.timeout, **extra_kw
File "/home/nobu/.local/lib/python3.6/site-packages/urllib3/util/connection.py", line 61, in create_connection
for res in socket.getaddrinfo(host, port, family, socket.SOCK_STREAM):
File "/usr/lib/python3.6/socket.py", line 745, in getaddrinfo
for res in _socket.getaddrinfo(host, port, family, type, proto, flags):
socket.gaierror: [Errno -3] Temporary failure in name resolution
During handling of the above exception, another exception …Run Code Online (Sandbox Code Playgroud) 我在 AWS Fargate 上使用 Amazon ECS,我的实例可以访问互联网,但连接在350秒后断开。平均而言,在 100 次中,我的服务收到ConnectionResetError: [Errno 104] Connection Reset by Peer Error 的次数大约有 5 次。我发现了一些建议来解决我的服务器端代码上的该问题,请参阅此处和此处
原因
如果使用 NAT 网关的连接空闲 350 秒或更长时间,连接就会超时。
当连接超时时,NAT 网关会向 NAT 网关后面尝试继续连接的任何资源返回 RST 数据包(它不会发送 FIN 数据包)。
解决方案
为了防止连接被丢弃,您可以通过该连接启动更多流量。或者,您可以在实例上启用 TCP keepalive,其值小于 350 秒。
现有代码:
url = "url to call http"
params = {
"year": year,
"month": month
}
response = self.session.get(url, params=params)
Run Code Online (Sandbox Code Playgroud)
为了解决这个问题,我目前正在使用使用tenacity 的创可贴重试逻辑解决方案,
@retry(
retry=(
retry_if_not_exception_type(
HTTPError
) # specific: requests.exceptions.ConnectionError
),
reraise=True,
wait=wait_fixed(2),
stop=stop_after_attempt(5),
) …Run Code Online (Sandbox Code Playgroud) python amazon-web-services amazon-ecs python-requests aws-fargate
我在我的盒子里部署了一个Web服务.我想用各种输入检查这项服务的结果.这是我正在使用的代码:
import sys
import httplib
import urllib
apUrl = "someUrl:somePort"
fileName = sys.argv[1]
conn = httplib.HTTPConnection(apUrl)
titlesFile = open(fileName, 'r')
try:
for title in titlesFile:
title = title.strip()
params = urllib.urlencode({'search': 'abcd', 'text': title})
conn.request("POST", "/somePath/", params)
response = conn.getresponse()
data = response.read().strip()
print data+"\t"+title
conn.close()
finally:
titlesFile.close()
Run Code Online (Sandbox Code Playgroud)
打印相同行数后,此代码出错(28233).错误信息:
Traceback (most recent call last):
File "testService.py", line 19, in ?
conn.request("POST", "/somePath/", params)
File "/usr/lib/python2.4/httplib.py", line 810, in request
self._send_request(method, url, body, headers)
File "/usr/lib/python2.4/httplib.py", line 833, in _send_request
self.endheaders() …Run Code Online (Sandbox Code Playgroud) python ×5
python-3.x ×3
amazon-ecs ×1
anaconda ×1
api ×1
aws-fargate ×1
connection ×1
http ×1
php ×1