使用自定义模式排序

Ser*_*kin 5 sort

有没有办法使用自定义模式输出文件的内容?

例如,有一个myfile包含以下内容的文件:

a
d
b
c
Run Code Online (Sandbox Code Playgroud)

..如何使用以下模式对其进行排序:首先打印以“b”开头的行,然后打印以“d”开头的行,然后按正常字母顺序打印行,因此预期输出为:

b
d
a
c
Run Code Online (Sandbox Code Playgroud)

Gil*_*il' 14

当您需要对数据进行超出sort能力范围的排序时,一种常见的方法是对数据进行预处理以添加排序键,然后进行排序,最后删除额外的排序键。例如,在这里,添加一个0if 一行以 开头b,一个1if 一行以 开头d,2否则添加一个。

sed -e 's/^b/0&/' -e t -e 's/^d/1&/' -e 't' -e 's/^/2/' |
sort |
sed 's/^.//'
Run Code Online (Sandbox Code Playgroud)

请注意,这会对所有b和d行进行排序。如果您希望这些行按原始顺序排列,那么最简单的方法是将您希望保留为 unsorted 的行分开。但是,您可以将原始行处理为排序键,nl但这里更复杂。(\t如果您的 sed 不理解该语法,请在整个过程中替换为文字制表符。)

nl -ba -nln |
sed 's/^[0-9]* *\t\([bd]\)/\1\t&/; t; s/^[0-9]* *\t/z\t0\t/' |
sort -k1,1 -k2,2n |
sed 's/^[^\t]*\t[^\t]*\t//'
Run Code Online (Sandbox Code Playgroud)

或者,使用 Perl、Python 或 Ruby 等语言可以轻松指定自定义排序函数。

perl -e 'print sort {($b =~ /^[bd]/) - ($a =~ /^[bd]/) ||
                     $a cmp $b} <>'
python -c 'import sys; sys.stdout.write(sorted(sys.stdin.readlines(), key=lambda s: (0 if s[0]=="b" else 1 if s[0]=="d" else 2), s))'
Run Code Online (Sandbox Code Playgroud)

或者,如果您想按原始顺序保留b和d行:

perl -e 'while (<>) {push @{/^b/ ? \@b : /^d/ ? \@d : \@other}, $_}
         print @b, @d, sort @other'
python -c 'import sys
b = []; d = []; other = []
for line in sys.stdin.readlines():
    if line[0]=="b": b += line
    elif line[0]=="d": d += line
    else: other += line
other.sort()
sys.stdout.writelines(b); sys.stdout.writelines(d); sys.stdout.writelines(other)'
Run Code Online (Sandbox Code Playgroud)


Ant*_*hon 5

您需要使用的不仅仅是sort命令。首先grep是b行,然后是d行,然后对没有b或d结尾的任何内容进行排序。

grep '^b' myfile > outfile
grep '^d' myfile >> outfile
grep -v '^b' myfile | grep -v '^d' | sort >> outfile
cat outfile
Run Code Online (Sandbox Code Playgroud)

将导致:

b
d
a
c
Run Code Online (Sandbox Code Playgroud)

这是假设该行以“模式”b和d如果是这样的整线内侧的图案或东西,你可以离开了插入符号(^)

单行等效项是:

(grep '^b' myfile ; grep '^d' myfile ; grep -v '^b' myfile | grep -v '^d' | sort)
Run Code Online (Sandbox Code Playgroud)