无法在 for 循环中获取 awk 输出

jet*_*bil 2 shell bash awk text-processing

我正在尝试创建一个脚本来检查网站上的单词。我有几个要检查,所以我试图通过另一个文件输入它们。

该文件称为“testurls”。在文件中,我先列出关键字,然后列出 URL。它们用分号分隔。

Example Domains;www.example.com
Google;www.google.com
Run Code Online (Sandbox Code Playgroud)

这是脚本:

#!/bin/bash
clear

# Call list of keywords and urls
DATA=`cat testurls`

for keyurl in $DATA
do
    keyword=`awk -F ";" '{print $1}' $keyurl`
    url=`awk -F ";" '{print $2}' $keyurl`
    curl -silent $url | grep '$keyword' > /dev/null
 if [ $? != 0 ]; then
    # Fail
        echo "Did not find $keyword on $url"
    else
    # Pass
        echo $url "Okay"
fi
done
Run Code Online (Sandbox Code Playgroud)

输出是:

awk: cannot open Example (No such file or directory)
awk: cannot open Example (No such file or directory)
curl: no URL specified!
curl: try 'curl --help' or 'curl --manual' for more information
Did not find  on
awk: cannot open Domains;www.example.com (No such file or directory)
awk: cannot open Domains;www.example.com (No such file or directory)
curl: no URL specified!
curl: try 'curl --help' or 'curl --manual' for more information
Did not find  on
awk: cannot open Google;www.google.com (No such file or directory)
awk: cannot open Google;www.google.com (No such file or directory)
curl: no URL specified!
curl: try 'curl --help' or 'curl --manual' for more information
Did not find  on
Run Code Online (Sandbox Code Playgroud)

我已经破解了很多年了。非常欢迎任何帮助。

Gil*_*il' 6

您的脚本有几个问题。我列出了我找到的那些,但我还没有测试过,可能还有其他的。

for keyurl in $DATA; do …$DATA在每个空格处拆分,而不是在每个换行符处拆分。所以在第一次迭代中,$DATA将是Example; 然后Domains;www.example.com,依此类推。此外,每个值都经过通配符扩展,因此如果*关键字中有 a,根据当前目录中存在的文件,您可能会看到时髦的结果。

您正在尝试处理换行符分隔的数据。一个简单的方法是

while read -r keyurl; do
  …
done <testurls
Run Code Online (Sandbox Code Playgroud)

这会去除每一行的缩进,这在这里可能不是一件坏事。(IFS= read -r keyurl如果您想keyurl准确地包含每一行,请使用。)

您对 的调用awk不起作用,因为您$keyurl作为文件名传递。您需要将其作为输入传递。当您使用它时,始终在变量替换周围使用双引号(否则 shell 会对它们的值执行一些扩展)。我还建议使用$(…)代替`…`;它们是相同的,除了`…`当你想引用里面的东西时很难使用,而 的语法$(…)是直观的。

keyword=`echo "$keyurl" | awk -F ";" '{print $1}'`
url=`echo "$keyurl" | awk -F ";" '{print $2}'`
Run Code Online (Sandbox Code Playgroud)

有一种更好的方法可以在第一个分号处拆分变量:使用 shell 的内置构造从字符串中去除前缀​​或后缀。

keyword=${keyurl%%;*} url=${keyurl#*;}
Run Code Online (Sandbox Code Playgroud)

但是由于您的数据来自read内置数据并且分隔符是单个字符,因此您可以利用该IFS功能并在阅读时直接拆分输入。

while IFS=';' read -r keyword url; do …
Run Code Online (Sandbox Code Playgroud)

来到 curl 和 grep 调用,请注意您正在寻找文字 text $keyword,因为您使用了单引号。使用双引号;请注意,关键字将被解释为基本的正则表达式。如果您希望将关键字解释为文字字符串,请将-F选项传递给grep. 您还应该放在-e模式之前,以防关键字以字符开头-(否则关键字将被解释为 grep 的选项)。最后关于 grep 的话题,它的-q选项相当于>/dev/null. 还要记住$url.周围的双引号。

curl -silent "$url" | grep -Fqe "$keyword"
Run Code Online (Sandbox Code Playgroud)

您可以if [ $? != 0 ]; then通过将命令直接放在那里来缩短该部分。

if curl -silent "$url" | grep -Fqe "$keyword"; then
Run Code Online (Sandbox Code Playgroud)

总之;

while IFS=';' read -r keyword url; do
  if curl -silent "$url" | grep -Fqe "$keyword"; then
    echo "Did not find $keyword on $url"
  else
    echo $url "Okay"
  fi
done
Run Code Online (Sandbox Code Playgroud)