ken*_*hfm 5 python cpu sleep while-loop
我试图找出为什么以下非常简单、最小的示例在我的 i7-5500U CPU、Windows 10 计算机上占用了大约 33% 的 CPU 使用率:
import time
import numpy as np
import scipy.linalg
import cProfile
class CPUTest:
def __init__(self):
self.running = True
def compute_stuff(self):
dims = 150
A = np.random.random((dims, dims))
B = scipy.linalg.inv(np.dot(A.T, A))
def run(self):
prev_time = time.time()
start_time = prev_time
while self.running:
time.sleep(0.3)
st = time.time()
self.compute_stuff()
et = time.time()
print 'Time for the whole iteration, inc. sleep: %.3f (ms), whereas the processing segment took %.3f (ms): ' % ((st - prev_time) * 1000, (et - st) * 1000)
prev_time = st
if st - start_time > 10.0:
break
t = CPUTest()
t.run()
# cProfile.run('t.run()')
Run Code Online (Sandbox Code Playgroud)
compute_stuff 函数只需要 2 毫秒,其余时间程序都在休眠。由于 sleep 不应该使用 CPU,因此该程序理论上应该只使用 0.6% 的 CPU 使用率运行,但目前它占用了大约 30%。
我试过一个分析器,它确认程序在 10 秒中处于睡眠状态 9.79 秒。
有人可以提示一下为什么python会这样吗?什么是减少 CPU 使用率的替代方法。
非常感谢!
编辑
总而言之,程序在 97% 以上的时间处于睡眠状态,而我仍然得到 33% 的 CPU 使用率。我想在不牺牲计算频率的情况下减少 CPU 使用率。
在这里您可以找到程序输出的示例:
Time for the whole iteration, inc. sleep: 302.000 (ms), whereas the processing segment took 1.000 (ms):
Time for the whole iteration, inc. sleep: 301.000 (ms), whereas the processing segment took 2.000 (ms):
Time for the whole iteration, inc. sleep: 303.000 (ms), whereas the processing segment took 3.000 (ms):
Time for the whole iteration, inc. sleep: 303.000 (ms), whereas the processing segment took 2.000 (ms):
Time for the whole iteration, inc. sleep: 302.000 (ms), whereas the processing segment took 1.000 (ms):
Time for the whole iteration, inc. sleep: 302.000 (ms), whereas the processing segment took 2.000 (ms):
Time for the whole iteration, inc. sleep: 302.000 (ms), whereas the processing segment took 2.000 (ms):
Time for the whole iteration, inc. sleep: 303.000 (ms), whereas the processing segment took 1.000 (ms):
Time for the whole iteration, inc. sleep: 301.000 (ms), whereas the processing segment took 2.000 (ms):
Time for the whole iteration, inc. sleep: 303.000 (ms), whereas the processing segment took 1.000 (ms):
Run Code Online (Sandbox Code Playgroud)
这是分析器的输出:
Ordered by: standard name
ncalls tottime percall cumtime percall filename:lineno(function)
1 0.000 0.000 10.050 10.050 <string>:1(<module>)
1 0.019 0.019 0.021 0.021 __init__.py:133(<module>)
1 0.067 0.067 0.119 0.119 __init__.py:205(<module>)
1 0.000 0.000 0.000 0.000 _components.py:1(<module>)
1 0.000 0.000 0.000 0.000 _laplacian.py:3(<module>)
49 0.000 0.000 0.000 0.000 _methods.py:37(_any)
49 0.000 0.000 0.001 0.000 _methods.py:40(_all)
49 0.011 0.000 0.137 0.003 _util.py:141(_asarray_validated)
1 0.001 0.001 0.001 0.001 _validation.py:1(<module>)
1 0.000 0.000 0.000 0.000 _version.py:114(_compare)
1 0.000 0.000 0.000 0.000 _version.py:148(__gt__)
2 0.000 0.000 0.000 0.000 _version.py:55(__init__)
1 0.000 0.000 0.000 0.000 _version.py:78(_compare_version)
1 0.008 0.008 0.009 0.009 base.py:1(<module>)
1 0.000 0.000 0.000 0.000 base.py:15(SparseWarning)
1 0.000 0.000 0.000 0.000 base.py:19(SparseFormatWarning)
1 0.000 0.000 0.000 0.000 base.py:23(SparseEfficiencyWarning)
1 0.000 0.000 0.000 0.000 base.py:61(spmatrix)
49 0.000 0.000 0.000 0.000 base.py:887(isspmatrix)
49 0.043 0.001 0.185 0.004 basic.py:619(inv)
49 0.000 0.000 0.001 0.000 blas.py:177(find_best_blas_type)
49 0.001 0.000 0.002 0.000 blas.py:223(_get_funcs)
1 0.000 0.000 0.000 0.000 bsr.py:1(<module>)
1 0.000 0.000 0.000 0.000 bsr.py:22(bsr_matrix)
1 0.012 0.012 0.012 0.012 compressed.py:1(<module>)
1 0.000 0.000 0.000 0.000 compressed.py:21(_cs_matrix)
1 0.000 0.000 0.000 0.000 construct.py:2(<module>)
1 0.000 0.000 0.000 0.000 coo.py:1(<module>)
1 0.000 0.000 0.000 0.000 coo.py:21(coo_matrix)
49 0.000 0.000 0.000 0.000 core.py:5960(isMaskedArray)
49 0.001 0.000 0.242 0.005 cpuTests.py:10(compute_stuff)
1 0.013 0.013 10.050 10.050 cpuTests.py:15(run)
1 0.000 0.000 0.000 0.000 csc.py:1(<module>)
1 0.000 0.000 0.000 0.000 csc.py:19(csc_matrix)
1 0.008 0.008 0.020 0.020 csr.py:1(<module>)
1 0.000 0.000 0.000 0.000 csr.py:21(csr_matrix)
18 0.000 0.000 0.000 0.000 data.py:106(_create_method)
1 0.000 0.000 0.000 0.000 data.py:121(_minmax_mixin)
1 0.000 0.000 0.000 0.000 data.py:22(_data_matrix)
1 0.000 0.000 0.000 0.000 data.py:7(<module>)
1 0.000 0.000 0.000 0.000 dia.py:1(<module>)
1 0.000 0.000 0.000 0.000 dia.py:17(dia_matrix)
1 0.000 0.000 0.000 0.000 dok.py:1(<module>)
1 0.000 0.000 0.000 0.000 dok.py:29(dok_matrix)
1 0.000 0.000 0.000 0.000 extract.py:2(<module>)
49 0.000 0.000 0.001 0.000 fromnumeric.py:1887(any)
49 0.005 0.000 0.006 0.000 function_base.py:605(asarray_chkfinite)
49 0.000 0.000 0.000 0.000 getlimits.py:245(__init__)
49 0.000 0.000 0.000 0.000 getlimits.py:270(max)
49 0.000 0.000 0.002 0.000 lapack.py:405(get_lapack_funcs)
49 0.002 0.000 0.003 0.000 lapack.py:447(_compute_lwork)
1 0.000 0.000 0.000 0.000 lil.py:19(lil_matrix)
1 0.002 0.002 0.002 0.002 lil.py:2(<module>)
49 0.000 0.000 0.000 0.000 misc.py:169(_datacopied)
3 0.000 0.000 0.000 0.000 nosetester.py:181(__init__)
3 0.000 0.000 0.000 0.000 ntpath.py:174(split)
3 0.000 0.000 0.000 0.000 ntpath.py:213(dirname)
3 0.000 0.000 0.000 0.000 ntpath.py:96(splitdrive)
49 0.000 0.000 0.000 0.000 numeric.py:406(asarray)
49 0.000 0.000 0.000 0.000 numeric.py:476(asanyarray)
98 0.000 0.000 0.000 0.000 numerictypes.py:942(_can_coerce_all)
49 0.000 0.000 0.000 0.000 numerictypes.py:964(find_common_type)
5 0.000 0.000 0.000 0.000 re.py:138(match)
2 0.000 0.000 0.000 0.000 re.py:143(search)
7 0.000 0.000 0.000 0.000 re.py:230(_compile)
1 0.000 0.000 0.000 0.000 sputils.py:2(<module>)
1 0.000 0.000 0.000 0.000 sputils.py:227(IndexMixin)
3 0.000 0.000 0.000 0.000 sre_compile.py:228(_compile_charset)
3 0.000 0.000 0.000 0.000 sre_compile.py:256(_optimize_charset)
3 0.000 0.000 0.000 0.000 sre_compile.py:433(_compile_info)
6 0.000 0.000 0.000 0.000 sre_compile.py:546(isstring)
3 0.000 0.000 0.000 0.000 sre_compile.py:552(_code)
3 0.000 0.000 0.000 0.000 sre_compile.py:567(compile)
3 0.000 0.000 0.000 0.000 sre_compile.py:64(_compile)
7 0.000 0.000 0.000 0.000 sre_parse.py:149(append)
3 0.000 0.000 0.000 0.000 sre_parse.py:151(getwidth)
3 0.000 0.000 0.000 0.000 sre_parse.py:189(__init__)
16 0.000 0.000 0.000 0.000 sre_parse.py:193(__next)
3 0.000 0.000 0.000 0.000 sre_parse.py:206(match)
13 0.000 0.000 0.000 0.000 sre_parse.py:212(get)
3 0.000 0.000 0.000 0.000 sre_parse.py:268(_escape)
3 0.000 0.000 0.000 0.000 sre_parse.py:317(_parse_sub)
3 0.000 0.000 0.000 0.000 sre_parse.py:395(_parse)
3 0.000 0.000 0.000 0.000 sre_parse.py:67(__init__)
3 0.000 0.000 0.000 0.000 sre_parse.py:706(parse)
3 0.000 0.000 0.000 0.000 sre_parse.py:92(__init__)
1 0.000 0.000 0.000 0.000 utils.py:117(deprecate)
1 0.000 0.000 0.000 0.000 utils.py:51(_set_function_name)
1 0.000 0.000 0.000 0.000 utils.py:68(__init__)
1 0.000 0.000 0.000 0.000 utils.py:73(__call__)
3 0.000 0.000 0.000 0.000 {_sre.compile}
1 0.000 0.000 0.000 0.000 {dir}
343 0.000 0.000 0.000 0.000 {getattr}
3 0.000 0.000 0.000 0.000 {hasattr}
158 0.000 0.000 0.000 0.000 {isinstance}
270 0.000 0.000 0.000 0.000 {len}
49 0.000 0.000 0.001 0.000 {method 'all' of 'numpy.ndarray' objects}
49 0.000 0.000 0.000 0.000 {method 'any' of 'numpy.ndarray' objects}
211 0.000 0.000 0.000 0.000 {method 'append' of 'list' objects}
49 0.000 0.000 0.000 0.000 {method 'astype' of 'numpy.ndarray' objects}
1 0.000 0.000 0.000 0.000 {method 'disable' of '_lsprof.Profiler' objects}
5 0.000 0.000 0.000 0.000 {method 'end' of '_sre.SRE_Match' objects}
6 0.000 0.000 0.000 0.000 {method 'extend' of 'list' objects}
3 0.000 0.000 0.000 0.000 {method 'find' of 'bytearray' objects}
205 0.000 0.000 0.000 0.000 {method 'get' of 'dict' objects}
2 0.000 0.000 0.000 0.000 {method 'group' of '_sre.SRE_Match' objects}
49 0.000 0.000 0.000 0.000 {method 'index' of 'list' objects}
3 0.000 0.000 0.000 0.000 {method 'items' of 'dict' objects}
1 0.000 0.000 0.000 0.000 {method 'join' of 'str' objects}
5 0.000 0.000 0.000 0.000 {method 'match' of '_sre.SRE_Pattern' objects}
49 0.021 0.000 0.021 0.000 {method 'random_sample' of 'mtrand.RandomState' objects}
98 0.001 0.000 0.001 0.000 {method 'reduce' of 'numpy.ufunc' objects}
3 0.000 0.000 0.000 0.000 {method 'replace' of 'str' objects}
2 0.000 0.000 0.000 0.000 {method 'search' of '_sre.SRE_Pattern' objects}
2 0.000 0.000 0.000 0.000 {method 'split' of 'str' objects}
60 0.000 0.000 0.000 0.000 {method 'startswith' of 'str' objects}
1 0.000 0.000 0.000 0.000 {method 'update' of 'dict' objects}
6 0.000 0.000 0.000 0.000 {min}
147 0.000 0.000 0.000 0.000 {numpy.core.multiarray.array}
49 0.036 0.001 0.036 0.001 {numpy.core.multiarray.dot}
4 0.000 0.000 0.000 0.000 {ord}
18 0.000 0.000 0.000 0.000 {setattr}
3 0.000 0.000 0.000 0.000 {sys._getframe}
49 9.794 0.200 9.794 0.200 {time.sleep}
99 0.000 0.000 0.000 0.000 {time.time}
Run Code Online (Sandbox Code Playgroud)
第二次编辑
我已经实现了等效的 C++ 版本(如下)。C++ 版本确实具有我期望的行为:它只使用了0.3% 到 0.5%的 CPU 使用率!
#include <iostream>
#include <chrono>
#include <random>
#include <thread>
// Tune this values to get a computation lasting from 2 to 10ms
#define DIMS 50
#define MULTS 20
/*
This function will compute MULTS times matrix multiplications of transposed(A)*A
We simply want to waste enough time doing computations (tuned to waste between 2ms and 10ms)
*/
double compute_stuff(double A[][DIMS], double B[][DIMS]) {
double res = 0.0;
for (int k = 0; k < MULTS; k++) {
for (int i = 0; i < DIMS; i++) {
for (int j = 0; j < DIMS; j++) {
B[i][j] = 0.0;
for (int l = 0; l < DIMS; l++) {
B[i][j] += A[l][j] * A[j][l];
}
}
}
// We store the result from the matrix B
for (int i = 0; i < DIMS; i++) {
for (int j = 0; j < DIMS; j++) {
A[i][j] = B[i][j];
}
}
}
for (int i = 0; i < DIMS; i++) {
for (int j = 0; j < DIMS; j++) {
res += A[i][j];
}
}
return res;
}
int main() {
std::cout << "Running main" << std::endl;
double A[DIMS][DIMS]; // Data buffer for a random matrix
double B[DIMS][DIMS]; // Data buffer for intermediate computations
std::default_random_engine generator;
std::normal_distribution<double> distribution(0.0, 1.0);
for (int i = 0; i < DIMS; i++) {
for (int j = 0; j < DIMS; j++) {
A[i][j] = distribution(generator);
}
}
bool keep_running = true;
auto prev_time = std::chrono::high_resolution_clock::now();
auto start_time = prev_time;
while (keep_running)
{
std::this_thread::sleep_for(std::chrono::milliseconds(300));
auto st = std::chrono::high_resolution_clock::now();
double res = compute_stuff(A, B);
auto et = std::chrono::high_resolution_clock::now();
auto iteration_time = std::chrono::duration_cast<std::chrono::milliseconds>(st - prev_time).count();
auto computation_time = std::chrono::duration_cast<std::chrono::milliseconds>(et - st).count();
auto elapsed_time = std::chrono::duration_cast<std::chrono::milliseconds>(et - start_time).count();
std::cout << "Time for the whole iteration, inc. sleep:" << iteration_time << " (ms), whereas the processing segment took " << computation_time << "(ms)" << std::endl;
keep_running = elapsed_time < 10 * 1000;
prev_time = st;
}
}
Run Code Online (Sandbox Code Playgroud)
在这里您还可以看到 C++ 等效程序的输出:
Time for the whole iteration, inc. sleep:314 (ms), whereas the processing segment took 7(ms)
Time for the whole iteration, inc. sleep:317 (ms), whereas the processing segment took 7(ms)
Time for the whole iteration, inc. sleep:316 (ms), whereas the processing segment took 8(ms)
Time for the whole iteration, inc. sleep:316 (ms), whereas the processing segment took 7(ms)
Time for the whole iteration, inc. sleep:314 (ms), whereas the processing segment took 10(ms)
Run Code Online (Sandbox Code Playgroud)
似乎有一些特定于 python 的事情正在发生。已在 3 台机器(linux 和 Windows)中确认了相同的行为
小智 7
当我为游戏编写程序时,我发现了这个问题。
我意识到,即使我创建了一个 while 无限循环,只打印一条 hello world 消息,我的程序的 CPU 使用率仍然是 30%。
所以我在 while 循环的开始和结束时使用 time.sleep(0.05) 。
我的问题解决了。只需在循环中使用 sleep 即可。我认为可以做到。
我认为你正在衡量不同的事物,这会引起一些混乱。
首先,切换上下文需要成本;如果您有批处理作业,那么让系统决定何时切换到其他任务可能比您自己插入睡眠更好。每次您的进程休眠时,它都会花费一些时间调用系统以重新安排时间并设置再次唤醒的警报,然后在警报触发后恢复。
传统上,任务管理器使用的 CPU 使用指示也是不精确的。它们的目的是找出哪个程序使系统保持忙碌,并指示调度程序正在处理什么。例如,一个常见的迹象是系统中有一个空闲进程需要花费大量时间;该进程只是为了保持一致性,因此对于调度程序来说,无事可做时进入睡眠状态并不是特殊情况。
CPU 速度本身现在是可变的。如果您的程序经常很少睡眠,许多计算机会减慢速度以匹配它,该功能旨在使播放视频等工作不需要在运行和睡眠模式之间切换,而这本身需要一些时间。特别是,一旦进入睡眠状态,再次启动需要时间,这使基于时间的调度(睡眠和超时)变得复杂并延迟反应。这意味着 CPU 百分比只能与负载高度相似的情况下的另一个进行比较。
您的系统可能有一些其他任务在后台运行,这些任务很少需要 CPU 时间。当睡眠时间较短时,这些任务可能会被分配到同一个处理器上,但如果该任务睡眠时间较长,则通常会在另一个处理器上运行。由于您的程序仅需要一个处理器容量的一小部分,因此百分比差异很大。
我们看到的另一个方面是时间测量仅以毫秒为单位。当工作切片被指示为一到三毫秒之间时,我们有一个非常大的相对量化误差。这些切片太小,无法使用该系统上的任务管理器或 time.time() 进行可靠测量。
考虑到所有这些额外的变量,我们真正知道的是,睡眠时间越多,程序的开销就越大。像 unix time(1) 这样的工具将通过划分 wall 花费的时间(实时经过的时间)、用户(运行程序本身所花费的时间)和系统(处理程序调用所花费的时间,包括管理时间)来指示特定任务的分布。睡眠的开销,但不是实际睡眠的时间)。
这些睡眠的目标是什么?通过设置线程优先级不是更好吗?