有没有办法使用自定义模式输出文件的内容?
例如,有一个myfile包含以下内容的文件:
a
d
b
c
Run Code Online (Sandbox Code Playgroud)
..如何使用以下模式对其进行排序:首先打印以“b”开头的行,然后打印以“d”开头的行,然后按正常字母顺序打印行,因此预期输出为:
b
d
a
c
Run Code Online (Sandbox Code Playgroud)
Gil*_*il' 14
当您需要对数据进行超出sort能力范围的排序时,一种常见的方法是对数据进行预处理以添加排序键,然后进行排序,最后删除额外的排序键。例如,在这里,添加一个0if 一行以 开头b,一个1if 一行以 开头d,2否则添加一个。
sed -e 's/^b/0&/' -e t -e 's/^d/1&/' -e 't' -e 's/^/2/' |
sort |
sed 's/^.//'
Run Code Online (Sandbox Code Playgroud)
请注意,这会对所有b和d行进行排序。如果您希望这些行按原始顺序排列,那么最简单的方法是将您希望保留为 unsorted 的行分开。但是,您可以将原始行处理为排序键,nl但这里更复杂。(\t如果您的 sed 不理解该语法,请在整个过程中替换为文字制表符。)
nl -ba -nln |
sed 's/^[0-9]* *\t\([bd]\)/\1\t&/; t; s/^[0-9]* *\t/z\t0\t/' |
sort -k1,1 -k2,2n |
sed 's/^[^\t]*\t[^\t]*\t//'
Run Code Online (Sandbox Code Playgroud)
或者,使用 Perl、Python 或 Ruby 等语言可以轻松指定自定义排序函数。
perl -e 'print sort {($b =~ /^[bd]/) - ($a =~ /^[bd]/) ||
$a cmp $b} <>'
python -c 'import sys; sys.stdout.write(sorted(sys.stdin.readlines(), key=lambda s: (0 if s[0]=="b" else 1 if s[0]=="d" else 2), s))'
Run Code Online (Sandbox Code Playgroud)
或者,如果您想按原始顺序保留b和d行:
perl -e 'while (<>) {push @{/^b/ ? \@b : /^d/ ? \@d : \@other}, $_}
print @b, @d, sort @other'
python -c 'import sys
b = []; d = []; other = []
for line in sys.stdin.readlines():
if line[0]=="b": b += line
elif line[0]=="d": d += line
else: other += line
other.sort()
sys.stdout.writelines(b); sys.stdout.writelines(d); sys.stdout.writelines(other)'
Run Code Online (Sandbox Code Playgroud)
您需要使用的不仅仅是sort命令。首先grep是b行,然后是d行,然后对没有b或d结尾的任何内容进行排序。
grep '^b' myfile > outfile
grep '^d' myfile >> outfile
grep -v '^b' myfile | grep -v '^d' | sort >> outfile
cat outfile
Run Code Online (Sandbox Code Playgroud)
将导致:
b
d
a
c
Run Code Online (Sandbox Code Playgroud)
这是假设该行以“模式”b和d如果是这样的整线内侧的图案或东西,你可以离开了插入符号(^)
单行等效项是:
(grep '^b' myfile ; grep '^d' myfile ; grep -v '^b' myfile | grep -v '^d' | sort)
Run Code Online (Sandbox Code Playgroud)