bash tail在一个实时日志文件中,计算具有相同日期/时间的uniq行

zap*_*app 5 bash logging tail uniq

我正在寻找一种在实时日志文件上拖尾的好方法,并显示具有相同日期/时间的行数.

目前这是有效的:

 tail -F /var/logs/request.log | [cut the date-time] | uniq -c
Run Code Online (Sandbox Code Playgroud)

但性能不够好.延迟超过一分钟,并且每次以少量线路输出.

任何的想法?

Flo*_*ris 10

您的问题很可能与系统中的缓冲有关,而不是您的代码行本质上的任何错误.我能够创建一个可以重现它的测试场景 - 然后让它消失.我希望它也适合你.

这是我的测试场景.首先,我写了一个短脚本,每100毫秒(大约)将时间写入一个文件 - 这是我的"日志文件",它生成足够的数据,uniq -c每秒钟应该给我一个有趣的输出:

#!/bin/ksh
while :
do
  echo The time is `date` >> a.txt
  sleep 0.1
done
Run Code Online (Sandbox Code Playgroud)

(注意 - 我必须使用ksh哪个能够做到亚秒级sleep)

在另一个窗口中,我输入

tail -f a.txt | uniq -c
Run Code Online (Sandbox Code Playgroud)

果然,您每秒都会出现以下输出:

   9 The time is Thu Dec 12 21:01:05 EST 2013
  10 The time is Thu Dec 12 21:01:06 EST 2013
  10 The time is Thu Dec 12 21:01:07 EST 2013
   9 The time is Thu Dec 12 21:01:08 EST 2013
  10 The time is Thu Dec 12 21:01:09 EST 2013
   9 The time is Thu Dec 12 21:01:10 EST 2013
  10 The time is Thu Dec 12 21:01:11 EST 2013
  10 The time is Thu Dec 12 21:01:12 EST 2013
Run Code Online (Sandbox Code Playgroud)

等没有延误.重要的是要注意 - 我没有试图减少时间.接下来,我做到了

tail -f a.txt | cut -f7 -d' ' | uniq -c
Run Code Online (Sandbox Code Playgroud)

你的问题再现了 - 它会"挂起"很长一段时间(直到缓冲区中有4k个字符,然后它会立刻呕吐出来).

在线搜索(/sf/answers/1177648461/)告诉我一个名为stdbuf的实用程序.在该参考文献中,它特别提到了您的场景,并提供了以下解决方法(复述以匹配上面的场景):

tail -f a.txt | stdbuf -oL cut -f7 -d' ' | uniq -c
Run Code Online (Sandbox Code Playgroud)

这将是伟大的...除了我的机器(Mac OS)上不存在此实用程序 - 它特定于GNU coreutils.这让我无法测试 - 虽然它可能是一个很好的解决方案.

永远不要害怕 - 我根据socat命令找到了以下解决方法(我实际上几乎无法理解,但我改编自https://unix.stackexchange.com/a/25377给出的答案).

创建一个名为的小文件tailcut.sh(这是上面链接中的"long_running_command"):

#!/bin/ksh
tail -f a.txt | cut -f7 -d' '
Run Code Online (Sandbox Code Playgroud)

用它赋予它执行权限chmod 755 tailcut.sh.然后发出以下命令:

socat EXEC:./tailcut.sh,pty,ctty STDIO | uniq -c
Run Code Online (Sandbox Code Playgroud)

嘿presto - 你的块状输出不再是块状.在socat从脚本输出发送直奔下水管,和uniq可以做的事情.

  • stdbuf实现了它。超级答案。 (2认同)