dfc*_*dfc 24 coreutils text-processing history
我正在浏览 coreutils 中包含的文件列表,并且我能够想出一个示例,说明我如何亲自使用除 ptx 之外提供的所有命令。你能举出一两个(或三个)使用 ptx 的例子吗?用例越多样化越好。
$ apropos ptx
ptx(1) - produce a permuted index of file contents
Run Code Online (Sandbox Code Playgroud)
Jos*_* R. 11
显然,它在过去被用来索引 Unix 参考手册。
在下面的参考资料中,维基百科文章解释了置换索引是什么(也称为 KWIC,或“上下文中的关键字”)并以神秘结尾:
由许多带有自己的描述性标题的短章节组成的书籍,最显着的是手册页的集合,通常以排列的索引部分结束,使读者可以通过标题中的任何单词轻松找到一个章节。这种做法已不再普遍。
更多的搜索揭示了参考文献中剩余的文章,这些文章解释了更多关于 Unix 手册页如何使用置换索引的信息。他们处理的主要问题似乎是手册页没有连续编号。
从我收集到的信息来看,使用置换索引的做法现在已经过时了。
参考
bis*_*hop 10
@Joseph R. 接受的历史答案很好,但让我们看看如何使用它。
ptx从文本生成一个置换词索引(“ptx”)。一个例子最容易理解:
$ cat input
a
b
c
$ ptx -A -w 25 input
:1: a b c
:2: a b c
:3: a b c
^^^^ ^ ^^^^-words to the input's right
| +-here is the actual input
+-words to the input's left
Run Code Online (Sandbox Code Playgroud)
在右侧,您会看到输入中的不同单词以及围绕它们的左右单词上下文。第一个字是“a”。它出现在第一行,其右边是“b”和“c”。第二个词是“b”,出现在第二行,左边是“a”,右边是“c”。最后,“c”出现在第三行,然后是“a”和“b”。
使用它,您可以找到文本中任何单词的行号和周围单词。这听起来很像grep,嗯?不同之处在于ptx理解文本的结构,以单词和句子的逻辑单位。这使得ptx处理英文文本时的上下文输出比 grep 更相关。
让我们比较ptx和grep,使用 James Ellroy 的American Tabloid 的第一段:
$ cat text
America was never innocent. We popped our cherry on the boat over and looked back with no regrets. You can’t ascribe our fall from grace to any single event or set of circumstances. You can’t lose what you lacked at conception.
Run Code Online (Sandbox Code Playgroud)
这是grep(手动将颜色匹配更改为由 包围//):
$ grep -ni you text
1:America was never innocent. We popped our cherry on the boat over and looked back with no regrets. /You/ can’t ascribe our fall from grace to any single event or set of circumstances. /You/ can’t lose what /you/ lacked at conception.
Run Code Online (Sandbox Code Playgroud)
这是ptx:
$ ptx -Afo <(echo you) text
text:1: /back with no regrets. You can’t ascribe our fall/
text:1: /or set of circumstances. You can’t lose what you/
text:1: /. You can’t lose what you lacked at conception.
Run Code Online (Sandbox Code Playgroud)
因为grep是面向行的,而且这一段都是一行,所以grep输出不如ptx.