dim*_*mid 3 select json filtering identifier jq
我有一个具有以下格式的 JSON 文件:
[
{
"id": "00001",
"attr": {
"a": "foo",
"b": "bar",
...
}
},
{
"id": "00002",
"attr": {
...
},
...
},
...
]
Run Code Online (Sandbox Code Playgroud)
和一个带有 id 列表的文本文件,每行一个。我想jq仅用于过滤文本文件中提及其 ID 的记录。即如果列表包含“00001”,则只应打印第一个。
请注意,我不能简单地grep因为每个记录可能具有任意数量的属性和子属性。
基本上有两种方法可以进行:
两者都是可行的,但在这里我们说明(2),因为它导致了一个简单但有效的解决方案。
假设 JSON 文件名为 in.json,id 列表位于名为 ids.txt 的文件中,如下所示:
00001
00010
Run Code Online (Sandbox Code Playgroud)
请注意,此文件没有引号。如果是这样,则可以显着简化以下内容,如后记所示。
诀窍是将 ids.txt 转换为 JSON 数组。有了上述关于引号的假设,这可以通过以下方式完成:
jq -R . ids.txt | jq -s .
Run Code Online (Sandbox Code Playgroud)
假设有一个合理的外壳,现在有一个简单的解决方案:
jq --argjson ids "$(jq -R . ids.txt | jq -s .)" '
map( select( .id as $id | $ids | index($id) ))' in.json
Run Code Online (Sandbox Code Playgroud)
假设您的 jq 具有any/2,那么可以通过定义获得一个更简单、更有效的解决方案:
def isin($a): . as $in | any($a[]; $in == .);
Run Code Online (Sandbox Code Playgroud)
所需的 jq 过滤器就是:
map( select( .id | isin($ids) ) )
Run Code Online (Sandbox Code Playgroud)
如果将这两行 jq 放入名为 select.jq 的文件中,则所需的咒语很简单:
jq --argjson ids "$(jq -R . ids.txt | jq -s)" -f select.jq in.json
Run Code Online (Sandbox Code Playgroud)
如果索引文件包含有效的 JSON 文本流(例如,带引号的字符串)并且您的 jq 支持该--slurpfile选项,则调用可以进一步简化为:
jq --slurpfile ids ids.txt -f select.jq in.json
Run Code Online (Sandbox Code Playgroud)
或者,如果您希望将所有内容都作为单行:
jq --slurpfile ids ids.txt 'map(select(.id as $id|any($ids[];$id==.)))' in.json
Run Code Online (Sandbox Code Playgroud)