Tim*_*Tim 65 text-processing table
我有两个文本文件。第一个有内容:
Languages
Recursively enumerable
Regular
Run Code Online (Sandbox Code Playgroud)
而第二个有内容:
Minimal automaton
Turing machine
Finite
Run Code Online (Sandbox Code Playgroud)
我想将它们按列合并到一个文件中。所以我试过了paste 1 2,它的输出是:
Languages Minimal automaton
Recursively enumerable Turing machine
Regular Finite
Run Code Online (Sandbox Code Playgroud)
但是我想让列对齐良好,例如
Languages Minimal automaton
Recursively enumerable Turing machine
Regular Finite
Run Code Online (Sandbox Code Playgroud)
我想知道是否可以在不手动处理的情况下实现这一目标?
添加:
这是另一个例子,布鲁斯方法几乎可以解决它,除了一些轻微的错位,我想知道为什么?
$ cat 1
Chomsky hierarchy
Type-0
—
$ cat 2
Grammars
Unrestricted
$ paste 1 2 | pr -t -e20
Chomsky hierarchy Grammars
Type-0 Unrestricted
— (no common name)
Run Code Online (Sandbox Code Playgroud)
gle*_*man 82
您只需要该column命令,并告诉它使用制表符来分隔列
paste file1 file2 | column -s $'\t' -t
Run Code Online (Sandbox Code Playgroud)
为了解决“空单元格”争议,我们只需要-n选择column:
$ paste <(echo foo; echo; echo barbarbar) <(seq 3) | column -s $'\t' -t
foo 1
2
barbarbar 3
$ paste <(echo foo; echo; echo barbarbar) <(seq 3) | column -s $'\t' -tn
foo 1
2
barbarbar 3
Run Code Online (Sandbox Code Playgroud)
我的专栏手册页指出-n是“Debian GNU/Linux 扩展”。我的 Fedora 系统没有出现空单元问题:它似乎是从 BSD 派生的,手册页说“2.23 版将 -s 选项更改为非贪婪”
Bru*_*ger 16
您正在寻找方便的 dandypr命令:
paste file1 file2 | pr -t -e24
Run Code Online (Sandbox Code Playgroud)
“-e24”是“将制表位扩展到 24 个空格”。幸运的是,paste在列之间放置了一个制表符,因此pr可以扩展它。我通过计算“递归可枚举”中的字符并添加 2 来选择 24。
Pet*_*r.O 11
更新:这里有一个更简单的脚本(问题末尾的那个)用于列表输出。只需像您一样将文件名传递给它paste......它用于html制作框架,因此它是可调整的。它确实保留了多个空格,并且在遇到 unicode 字符时保留了列对齐方式。但是,编辑器或查看器渲染 unicode 的方式完全是另一回事......
?????????????????????????????????????????????????????????????????????????????????
? Languages ? Minimal ? Chomsky ? Unrestricted ?
?????????????????????????????????????????????????????????????????????????????????
? Recursive ? Turing machine ? Finite ? space indented ?
?????????????????????????????????????????????????????????????????????????????????
? Regular ? Grammars ? ? ? unicode may render oddly ?
?????????????????????????????????????????????????????????????????????????????????
? 1 2 3 4 spaces ? ? Symbol-& ? but the column count is ok ?
?????????????????????????????????????????????????????????????????????????????????
? ? ? ? Context ?
?????????????????????????????????????????????????????????????????????????????????
Run Code Online (Sandbox Code Playgroud)
#!/bin/bash
{ echo -e "<html>\n<table border=1 cellpadding=0 cellspacing=0>"
paste "$@" |sed -re 's#(.*)#\x09\1\x09#' -e 's#\x09# </pre></td>\n<td><pre> #g' -e 's#^ </pre></td>#<tr>#' -e 's#\n<td><pre> $#\n</tr>#'
echo -e "</table>\n</html>"
} |w3m -dump -T 'text/html'
Run Code Online (Sandbox Code Playgroud)
答案中提供的工具的概要(到目前为止)。
我非常仔细地观察过它们;这是我发现的:
paste# 这个工具对目前所有的答案都是通用的 # 它可以处理多个文件;因此多列......好!# 它用制表符分隔每一列...很好。# 它的输出没有列表。
下面的所有工具都删除了这个分隔符!...如果你需要一个分隔符,那就不好了。
column # 它删除了 Tab 分隔符,所以字段标识纯粹是由它似乎处理得很好的列决定的......我没有发现任何错误...... # 除了没有唯一的分隔符之外,它工作正常!
expand # 只有一个tab设置,所以超过2列是不可预知的
pr# 只有一个选项卡设置,所以超过 2 列是不可预测的。# 处理unicode时列的对齐方式不准确,去掉了Tab分隔符,所以字段标识纯粹靠列对齐方式
对我来说,column它显然是最好的单行解决方案。如果你想要分隔符或文件的 ASCII 艺术制表,请继续阅读,否则......columns非常好:)...
这是一个脚本,它接受任意数量的文件并创建一个 ASCII 艺术表格演示文稿..(请记住,unicode 可能不会呈现到预期的宽度,例如。? 是单个字符。这与列完全不同数字是错误的,就像上面提到的一些实用程序的情况一样。)......脚本的输出,如下所示,来自 4 个输入文件,命名为 F1 F2 F3 F4 ...
+------------------------+-------------------+-------------------+--------------+
| Languages | Minimal automaton | Chomsky hierarchy | Grammars |
| Recursively enumerable | Turing machine | Type-0 | Unrestricted |
| Regular | Finite | — | |
| Alphabet | | Symbol | |
| | | | Context |
+------------------------+-------------------+-------------------+--------------+
Run Code Online (Sandbox Code Playgroud)
#!/bin/bash
# Note: The next line is for testing purposes only!
set F1 F2 F3 F4 # Simulate commandline filename args $1 $2 etc...
p=' ' # The pad character
# Get line and column stats
cc=${#@}; lmax= # Count of columns (== input files)
for c in $(seq 1 $cc) ;do # Filenames from the commandline
F[$c]="${!c}"
wc=($(wc -l -L <${F[$c]})) # File length and width of longest line
l[$c]=${wc[0]} # File length (per file)
L[$c]=${wc[1]} # Longest line (per file)
((lmax<${l[$c]})) && lmax=${l[$c]} # Length of longest file
done
# Determine line-count deficits of shorter files
for c in $(seq 1 $cc) ;do
((${l[$c]}<lmax)) && D[$c]=$((lmax-${l[$c]})) || D[$c]=0
done
# Build '\n' strings to cater for short-file deficits
for c in $(seq 1 $cc) ;do
for n in $(seq 1 ${D[$c]}) ;do
N[$c]=${N[$c]}$'\n'
done
done
# Build the command to suit the number of input files
source=$(mktemp)
>"$source" echo 'paste \'
for c in $(seq 1 $cc) ;do
((${L[$c]}==0)) && e="x" || e=":a -e \"s/^.{0,$((${L[$c]}-1))}$/&$p/;ta\""
>>"$source" echo '<(sed -re '"$e"' <(cat "${F['$c']}"; echo -n "${N['$c']}")) \'
done
# include the ASCII-art Table framework
>>"$source" echo ' | sed -e "s/.*/| & |/" -e "s/\t/ | /g" \' # Add vertical frame lines
>>"$source" echo ' | sed -re "1 {h;s/[^|]/-/g;s/\|/+/g;p;g}" \' # Add top and botom frame lines
>>"$source" echo ' -e "$ {p;s/[^|]/-/g;s/\|/+/g}"'
>>"$source" echo
# Run the code
source "$source"
rm "$source"
exit
Run Code Online (Sandbox Code Playgroud)
这是我的原始答案(为了代替上面的脚本而进行了一些修剪)
使用wc得到的列宽,并sed与右侧垫可见的字符.(只是在这个例子中)...然后paste用加入了两列标签字符...
paste <(sed -re :a -e 's/^.{1,'"$(($(wc -L <F1)-1))"'}$/&./;ta' F1) F2
# output (No trailing whitespace)
Languages............. Minimal automaton
Recursively enumerable Turing machine
Regular............... Finite
Run Code Online (Sandbox Code Playgroud)
如果要填充右列:
paste <( sed -re :a -e 's/^.{1,'"$(($(wc -L <F1)-1))"'}$/&./;ta' F1 ) \
<( sed -re :a -e 's/^.{1,'"$(($(wc -L <F2)-1))"'}$/&./;ta' F2 )
# output (With trailing whitespace)
Languages............. Minimal automaton
Recursively enumerable Turing machine...
Regular............... Finite...........
Run Code Online (Sandbox Code Playgroud)
您快到了。paste在每列之间放置一个制表符,因此您需要做的就是展开制表符。(我假设您的文件不包含选项卡。)您确实需要确定左列的宽度。使用(最近足够的)GNU 实用程序,wc -L显示最长行的长度。在其他系统上,首先使用 awk。这+1是您想要的列之间的空白空间量。
paste left.txt right.txt | expand -t $(($(wc -L <left.txt) + 1))
paste left.txt right.txt | expand -t $(awk 'n<length {n=length} END {print n+1}')
Run Code Online (Sandbox Code Playgroud)
如果您有 BSD 列实用程序,您可以使用它来确定列宽并一次性展开选项卡。(?是一个文字制表符;在 bash/ksh/zsh 下你可以$'\t'改用,在任何 shell 中你都可以使用"$(printf '\t')".)
paste left.txt right.txt | column -s '?' -t
Run Code Online (Sandbox Code Playgroud)
我无法对 glenn jackman 的回答发表评论,因此我添加此内容以解决 Peter.O 指出的空单元格问题。在每个选项卡之前添加一个空字符可消除被视为单个中断的分隔符的运行并解决问题。(我最初使用空格,但使用空字符消除了列之间的额外空间。)
paste file1 file2 | sed 's/\t/\0\t/g' | column -s $'\t' -t
Run Code Online (Sandbox Code Playgroud)
如果空字符由于各种原因导致问题,请尝试:
paste file1 file2 | sed 's/\t/ \t/g' | column -s $'\t' -t
Run Code Online (Sandbox Code Playgroud)
或者
paste file1 file2 | sed $'s/\t/ \t/g' | column -s $'\t' -t
Run Code Online (Sandbox Code Playgroud)
双方sed并column展示在口味和Unix / Linux,BSD特别是(和Mac OS X)与GNU / Linux版本实现改变。