如何用/dev/random的输出填充一定量的RAM?

Alf*_*.37 2 random bash shell ram

以下命令行可用于分配任意数量的内存,例如 1GB。内存区域被零填充。

</dev/zero head -c $((1024**3)) | tail
Run Code Online (Sandbox Code Playgroud)

然而,如何用随机数类似地填充给定数量的内存,例如 1GB?

与预期相反,以下两个变体不会导致工作内存的可选择分配大小,而是在终端中输出数百个随机字符:

</dev/random head -c $((1024**3)) | tail
Run Code Online (Sandbox Code Playgroud)

下面的看起来不像是一个解决方案,因为它操作数据(remove/0)并且只保存了几百个字符的变量,而不是1GB:

output=$(sudo head -c $((1024**3)) /dev/random | tr -d '\0' | tail)
Run Code Online (Sandbox Code Playgroud)

关于以下可能是一种解决方案的想法,但看起来 var 上只有几百个字符而不是 1GB:

output=""
while IFS= read -r -d '' substring; do
 output+="$substring"
done < <(sudo head -c $((1024**3)) /dev/random | tail)
Run Code Online (Sandbox Code Playgroud)

mar*_*rkp 7

为了便于讨论:

  • 我们将fill RAM通过将一长串数据分配给一个bash变量
  • 我们将限制自己的数据量为 1 GB (1024^3)(增加大小是可行的,但在某些时候我们还需要考虑操作系统设置,特别是ARG_MAX它决定可以分配给的最大数据量一个变量)
  • 由于 OP 似乎主要对填充内存感兴趣,我们将基于a)重复字符(以单个数字0为例)和b)随机数据来查看几个选项
  • 根据操作系统、操作系统设置、我们保存的数据类型以及我们如何填充变量,会出现一些警告/问题;其中一些问题将在答案的后一部分中讨论
  • 我们将测试的系统:Ubuntu 22.04, bash v.5.1.16(1), i7-1260P,64 GB RAM

总体结果:

Solution                     Time     Max Memory(*)   Notes
---------------------------------------------------------------------------------------------------------------------------
Appending `bash` variables   12 secs      1.3 GB      reduce memory overhead (smaller "t") at cost of taking longer to run
/dev/urandom (100%)          18 secs      2.0 GB
/dev/urandom (appending)     15 secs      1.3 GB      16x copies of 64 MB of random data

(*) - ignoring the additional 3.5-4 GB of memory required for: echo "${#x} / ${x:0:9} / ${x: -10}"
Run Code Online (Sandbox Code Playgroud)

解决方案:附加bash变量

我发现(到目前为止)最快的解决方案也限制了过多的内存使用:

$ cat fill_ram.1
#!/bin/bash

sleep 2                                           # allow time to grab initial memory usage

printf -v t "%0*d" $(( 64*1024*1024 )) 0          # populate variable "t" with 64 million 0's

echo "t(length) = ${#t}"                          # verify length of "t"

unset x

time for ((i=1;i<=16;i++)); do x+="${t}"; done    # populate variable "x" with 1 billion 0's

sleep 5

echo "x(length) = ${#x} / x(head) = ${x:0:9} / x(tail) = ${x: -10}"

sleep 2                                           # allow time to grab final memory usage
Run Code Online (Sandbox Code Playgroud)

进行试驾:

$ ./fill_ram.1
t(length) = 67108864

real    0m11.185s
user    0m6.448s
sys     0m4.735s

x(length) = 1073741824 / x(head) = 0000000000 / x(tail) = 0000000000
Run Code Online (Sandbox Code Playgroud)

从这个输出中我们可以看到该变量x包含 1 GB (1073741824) 的 0。

当脚本运行时,我在单独的控制台中运行以下命令来捕获总内存使用情况:

$ while true; do ps -eo size,command --sort -size | grep "[f]ill_ram"; sleep 0.1; done

  436 /bin/bash ./fill_ram.1              # initial memory reading repeats for 2 secs until guts of script start
  436 /bin/bash ./fill_ram.1 
  436 /bin/bash ./fill_ram.1 
... snip ... 
65976 /bin/bash ./fill_ram.1 
197052 /bin/bash ./fill_ram.1 
328128 /bin/bash ./fill_ram.1 
328128 /bin/bash ./fill_ram.1 
262592 /bin/bash ./fill_ram.1 
... snip ...
1245636 /bin/bash ./fill_ram.1 
1311172 /bin/bash ./fill_ram.1 
1311172 /bin/bash ./fill_ram.1 
1114560 /bin/bash ./fill_ram.1 
1114556 /bin/bash ./fill_ram.1 
1114556 /bin/bash ./fill_ram.1 
1114556 /bin/bash ./fill_ram.1            # final memory reading repeats for 5 secs until the last "echo" ...
... snip ...
3211844 /bin/bash ./fill_ram.1            # additional memory for echo / ${#x} / ${x:0:9} / ${x: -10}
... snip ...
5309000 /bin/bash ./fill_ram.1            # additional memory for echo / ${#x} / ${x:0:9} / ${x: -10}
... snip ...
3211844 /bin/bash ./fill_ram.1            # additional memory for echo / ${#x} / ${x:0:9} / ${x: -10}
... snip ...
5309000 /bin/bash ./fill_ram.1            # additional memory for echo / ${#x} / ${x:0:9} / ${x: -10}
3211844 /bin/bash ./fill_ram.1
1114688 /bin/bash ./fill_ram.1            # final memory reading for 2 secs, after the last "echo" and before the script ends
1114688 /bin/bash ./fill_ram.1
1114688 /bin/bash ./fill_ram.1
Run Code Online (Sandbox Code Playgroud)

有趣的是以下细节:

  • 内存使用量上下跳跃,直到接近结束时达到稳定状态
  • 最大内存(不含echo)实际上最高为 1280 MB (1311172),然后稳定在 1088 MB (1114556)
  • 请记住,针对 1 GB 变量(例如 )的后续操作echo ${#x} / ${x:0:9) / ${x: -10}如何进一步增加内存使用量(在本例中高达 5 GB),并且在内存有限的系统上,您可以考虑生成OOM errors
  • 最终/总内存使用量1088 MB分为64 MBfor 变量t和1024 MBfor 变量x
  • 每次内存“上升”然后“下降”时,下降量是 64 MB(变量的大小t)的倍数,并且代表操作的副作用x+="${t}",即短暂bash地复制变量t(这在side effect其他解决方案我们稍后会讨论)
  • 选择使用 64 MB 的变量t来平衡a)我们必须执行循环的次数bash/for,b)填充变量的(相对)较短的时间x,而c)限制变量的额外内存使用t(以及内存side effect使用情况)

解决方案:/dev/urandom (100%)

虽然之前的解决方案包括x用 10 亿个 0 填充变量,但这里有一种x用 10 亿个随机字符填充变量的方法:

  • bash变量无法存储\0字符,因此我们需要跳过该字符
  • 在测试过程中,我发现消耗来自的所有/dev/urandom数据往往会比我想要的 10 亿个字符多得多;我没有尝试head -c“正确”工作,而是选择将操作限制为可打印字符

我最终运行的代码:

$ cat fill_ram.2
#!/bin/bash

sleep 2

time x=$( tr -dc '[[:print:]]' < /dev/urandom | head -c $(( 1024**3 )) )

sleep 5

echo "x(length) = ${#x} / x(head) = ${x:0:9} / x(tail) = ${x: -10}"

sleep 2
Run Code Online (Sandbox Code Playgroud)

进行试驾:

$ ./fill_ram.2

real    0m17.607s
user    0m14.148s
sys     0m8.441s

x(length) = 1073741824 / x(head) = _d\rg(}Ah / x(tail) = A&Jk>7AkM!
Run Code Online (Sandbox Code Playgroud)

从这个输出中我们看到变量x包含 1 GB (1073741824) 的随机数据。

ps | grep本次运行的输出:

  436 /bin/bash ./fill_ram.2              # initial memory reading repeats for 2 secs until guts of script start
  436 /bin/bash ./fill_ram.2
  436 /bin/bash ./fill_ram.2
... snip ... 
  568 /bin/bash ./fill_ram.2              # parent script
 8228 /bin/bash ./fill_ram.2              # tr|head subprocess
  568 /bin/bash ./fill_ram.2              # parent script
16228 /bin/bash ./fill_ram.2              # tr|head subprocess
  568 /bin/bash ./fill_ram.2              # parent script
23440 /bin/bash ./fill_ram.2              # tr|head subprocess
... snip ...
  568   /bin/bash ./fill_ram.2            # parent script
1049196 /bin/bash ./fill_ram.2            # tr|head about to exit and data is copied to variable "x" in the parent script
2097776 /bin/bash ./fill_ram.2            # this runs for a few seconds ...
2097776 /bin/bash ./fill_ram.2
2097776 /bin/bash ./fill_ram.2
... snip ...
2097776 /bin/bash ./fill_ram.2
1049196 /bin/bash ./fill_ram.2            # until the copy finishes;
1049196 /bin/bash ./fill_ram.2            # the tr|head exits and
1049196 /bin/bash ./fill_ram.2            # we're left with just the
1049196 /bin/bash ./fill_ram.2            # contents of variable "x"
1049196 /bin/bash ./fill_ram.2
... snip ...
3146352 /bin/bash ./fill_ram.2            # then several secs of memory increases to ...
5243508 /bin/bash ./fill_ram.2            # support the last 
"echo"
... snip ...
1049196 /bin/bash ./fill_ram.2            # and finally back to just
1049196 /bin/bash ./fill_ram.2            # the memory for variable "x"
1049196 /bin/bash ./fill_ram.2            # until the script ends
Run Code Online (Sandbox Code Playgroud)

有趣的是以下细节:

  • tr|head随着子进程生成 10 亿个可打印字符,内存使用量稳步增加
  • 当tr|head子进程退出时,生成的 1 GB 数据将被复制到变量中x,导致总内存使用量跳至 2 GB
  • 复制完成后,内存使用量将回落至 1 GB
  • 正如我们在第一个解决方案中看到的,最后一个echo操作使内存使用量激增至 5 GB;同样,在内存紧张的情况下,这些变量引用 ( ${#x} / ${x:0:9) / ${x: -10}) 可能会导致OOM errors

解决方案:/dev/urandom(追加)

前两个解决方案的合并:

  • 通过子进程t使用 1 MB 可打印的随机字符填充变量tr|head
  • 使用bash/for循环来构建变量x

该解决方案的代码:

$ cat fill_ram.3
#!/bin/bash

time t=$( tr -dc '[[:print:]]' < /dev/urandom | head -c $(( 64*1024*1024 )) ) 

echo "t(length) = ${#t}"

unset x

time for ((i=1;i<=16;i++)); do x+="${t}"; done

echo "x(length) = ${#x} / x(head) = ${x:0:9} / x(tail) = ${x: -10}"
Run Code Online (Sandbox Code Playgroud)

进行试驾:

real    0m1.146s                   # tr|head => variable "t"
user    0m0.995s
sys     0m0.516s

t(length) = 67108864

real    0m13.343s                  # for loop => variable "x"
user    0m7.892s
sys     0m4.681s

x(length) = 1073741824 / x(head) = bQZzP/y,K / x(tail) = w3gL<RP_(H
Run Code Online (Sandbox Code Playgroud)

与/dev/urandom (100%)解决方案一样,我们看到变量x包含 1 GB (1073741824) 的随机数据。

我不会再用一堆输出来打扰(无聊?)你ps/memory;结果与解决方案相同Appending bash variables...最大内存达到 1.3 GB,然后稳定在1088 MB(64 MB对于变量t和1024 MB变量x)。


其他详细信息,排名不分先后...

cygwin 与 Linux

我最初开始在这两种环境中工作 -cygwin并且linux (Ubuntu 22.04)都运行bash v.5.x.

这个cygwin环境远没有 Ubuntu 那么“稳定”,所以我最终不得不在cygwin. “/dev/urandom”组件是最挑剔的。

我最终选择关注环境linux。

其余的杂乱细节包括测试中的一些细节cygwin。额外的cygwin零碎内容可能可以从答案的编辑历史记录中抢救出来。

/dev/random 与 /dev/urandom

我们将使用/dev/urandom而不是/dev/random; 请参阅以下内容了解更多详细信息:

保留/丢弃哪些字符?

bash变量无法存储\0字符,因此填充变量的初始尝试x包括以下内容:

x=$( tr -d '\0' </dev/urandom | head -c $(( 1024**3 )) )
Run Code Online (Sandbox Code Playgroud)

在我的 Linux 环境中,这会导致x在进程挂起(并且必须被终止)之前填充 2 GB 的数据。(请参阅ARG_MAX下面的更多细节)。

在我的cygwin环境中,这导致进程四处走动(几分钟后被终止),而内存使用量到处跳跃(277 MB、653 MB、327 MB、603 MB、221 MB,......)。

我很快发现我可以通过将数据限制为可打印字符来解决这些问题,因此:

x=$( tr -dc '[[:print:]]' < /dev/urandom | head -c $(( 1024**3 )) )    # linux
Run Code Online (Sandbox Code Playgroud)

对于cygwin环境来说,这是不太可接受的,因为bash进程仍然四处走动,而内存使用量(仍然)到处跳跃,所以cygwin我决定将初始数据大小减少到 1 MB ( 1024**2):

x=$( tr -dc '[[:print:]]' < /dev/urandom | head -c $(( 1024**2 )) )    # cygwin
Run Code Online (Sandbox Code Playgroud)

然后x通过循环将内容扩展1024倍bash/while:

i=1024; while ((i>1)); do x+="$x"; ((i/=2)); done
Run Code Online (Sandbox Code Playgroud)

我们可以捕捉多少个角色?

x正如前面提到的,一旦字节数达到x2 GB,变量的初始 linux 数量就会挂起;事实证明,这是操作系统配置的限制,特别是,我遇到了ARG_MAX操作系统配置的限制:

$ getconf ARG_MAX
2097152
Run Code Online (Sandbox Code Playgroud)

虽然可以增加此设置,但在本例中这是一个没有实际意义的问题,因为我只想在变量中保存 1 GB 的数据x。

在这种情况下cygwin,谁知道...

$ getconf ARG_MAX
undefined
Run Code Online (Sandbox Code Playgroud)

我没有进一步研究这个问题,因为我能够x通过最终的解决方案在变量中保存 1 GB 的数据(可打印字符,从 urandom 中获取 1MB,手动扩展x1024 倍)。

如果您在将 1 GB 数据填充到 ( bash) 变量中时遇到问题,或者希望超过 1 GB 但遇到问题,有几个切线可供研究ARG_MAX:对于初学者,请参阅以下内容:

如何验证已使用的内存量?

如何测量内存使用情况取决于操作系统、操作系统中可用的工具以及您正在使用的软件(bash在本例中)。

了解变量中数据的大小(bash / ${#x};其他软件可能有可以提供此信息的分析工具)应该足以证明内存正在使用,但如果您想在操作系统级别验证它,那么有很多解决这个问题的方法。

在我的 Linux 环境中,我选择了ps -o,而在我的cygwin环境中,我选择了简单/懒惰的路线,只是从 Windows 的任务管理器中获取了一些快照(并且因为cygwin's ps实现不支持该-o选项)。

有关如何测量内存的其他想法:

填充超过 1 GB 内存

如果你想填满超过1GB的内存怎么办?

使用${#x}==1073741824你会认为使用像以下之一这样简单的东西是安全的:

x+="${x}"          # double the size of variable "x"

y="${x}"           # copy variable "x" into variable "y"
Run Code Online (Sandbox Code Playgroud)

但在这两种情况下,我都会遇到以下情况OOM error:

./testme: xrealloc: cannot allocate 18446744071562068096 bytes
Run Code Online (Sandbox Code Playgroud)

老实说,我(到目前为止)还没有能够对这一事件找到令人满意的解释。即使side effect memory(即复制变量x)增加一倍或三倍,我们也只是谈论高达 3-5 GB 的总内存使用量……不足以OOM error在具有 64 GB RAM 的机器上引起问题,并且远不及所引用的 18 quintillion (?) 字节OOM error。

与此同时,我已经验证了两种附加解决方案确实允许通过简单地增加执行循环的次数来创建更大的变量(例如,2 GB、3 GB)for。

显然(?)在某些时候,大量数据可能会产生ARG_MAX值,但随后您可以考虑填充另一个变量。