Mic*_*zek 18
这取决于您的 CSV 文件是否仅将逗号用作分隔符,或者您是否有以下疯狂行为:
第一场、“第二场”、第三场
这假设您使用的是一个简单的 CSV 文件:
您可以通过多种方式摆脱单列;我以第 2 列为例。最简单的方法可能是使用cut,它允许您指定分隔符-d以及要打印的字段-f;这告诉它在逗号和输出字段 1 和字段 3 上拆分到最后:
$ cut -d, -f1,3- /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
如果你真的需要使用sed,你可以写一个正则表达式匹配第一个n-1字段,n第一个字段和其余的字段,并跳过输出nth(这里n是2,所以第一组匹配1时间:)\{1\}:
$ sed 's/\(\([^,]\+,\)\{1\}\)[^,]\+,\(.*\)/\1\3/' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
在 中有很多方法可以做到这一点awk,但没有一种特别优雅。您可以使用for循环,但处理尾随逗号是一种痛苦;忽略它会是这样的:
$ awk -F, '{for(i=1; i<=NF; i++) if(i != 2) printf "%s,", $i; print NL}' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
我发现输出字段 1 更容易,然后用于substr在字段 2 之后完成所有内容:
$ awk -F, '{print $1 "," substr($0, length($1)+length($2)+3)}' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
不过,这对于进一步的列来说很烦人
在sed这本质上是一样的表达和以前一样,但你也捕捉目标列,包括该组中更换多次:
$ sed 's/\(\([^,]\+,\)\{1\}\)\([^,]\+,\)\(.*\)/\1\3\3\4/' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
在awkfor 循环方式中,它类似于(再次忽略尾随逗号):
$ awk -F, '{
for(i=1; i<=NF; i++) {
if(i == 2) printf "%s,", $i;
printf "%s,", $i
}
print NL
}' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
该substr方式:
$ awk -F, '{print $1 "," $2 "," substr($0, length($1)+2)}' /path/to/your/file
Run Code Online (Sandbox Code Playgroud)
(tcdyl 在他的回答中提出了一个更好的方法)
我认为sed解决方案很自然地遵循其他解决方案,但它开始变得荒谬可笑
Pan*_*her 15
awk是你最好的选择。awk按数字打印字段,所以...
awk 'BEGIN { FS=","; OFS=","; } {print $1,$2,$3}' file
Run Code Online (Sandbox Code Playgroud)
要删除列,而不是打印它:
awk 'BEGIN { FS=","; OFS=","; } {print $1,$3}' file
Run Code Online (Sandbox Code Playgroud)
要更改顺序:
awk 'BEGIN { FS=","; OFS=","; } {print $3,$1,$2}' file
Run Code Online (Sandbox Code Playgroud)
重定向到输出文件。
awk 'BEGIN { FS=","; OFS=","; } {print $3,$1,$2}' file > output.file
Run Code Online (Sandbox Code Playgroud)
awk 也可以格式化输出。
除了如何剪切和重新排列字段(在其他答案中介绍)之外,还有一个古怪的 CSV 字段问题。
如果您的数据就属于这一“古怪”的范畴,有点的前和后过滤可以照顾它。下面显示的过滤器要求字符\x01, \x02, \x03,\x04不会出现在数据中的任何位置。
以下是围绕简单awk字段转储的过滤器。
注意: 字段五具有无效/不完整的“引用字段”布局,但它在行尾是良性的(取决于 CSV 解析器)。但是,当然,如果将其从当前的行尾位置交换,则会导致有问题的非加速结果。
更新; user121196指出了一个错误,当逗号位于尾随引号之前。这是修复。
数据
cat <<'EOF' >file
field one,"fie,ld,two",field"three","field,\",four","field,five
"15111 N. Hayden Rd., Ste 160,",""
EOF
Run Code Online (Sandbox Code Playgroud)
编码
sed -r 's/^/,/; s/\\"/\x01/g; s/,"([^"]*)"/,\x02\1\x03/g; s/,"/,\x02/; :MC; s/\x02([^\x03]*),([^\x03]*)/\x02\1\x04\2/g; tMC; s/^,// ' file |
awk -F, '{ for(i=1; i<=NF; i++) printf "%s\n", $i; print NL}' |
sed -r 's/\x01/\\"/g; s/(\x02|\x03)/"/g; s/\x04/,/g'
Run Code Online (Sandbox Code Playgroud)
输出:
field one
"fie,ld,two"
field"three"
"field,\",four"
"field,five
"15111 N. Hayden Rd., Ste 160,"
""
Run Code Online (Sandbox Code Playgroud)
这是预过滤器,用注释进行了扩展。
在后置滤波器只是一个逆转\x01。\x02, \x03,\x04
sed -r '
s/^/,/ # add a leading comma delimiter
s/\\"/\x01/g # obfuscate escaped quotation-mark (\")
s/,"([^"]*)"/,\x02\1\x03/g # obfuscate quotation-marks
s/,"/,\x02/ # when no trailing quote on last field
:MC # obfuscate commas embedded in quotes
s/\x02([^\x03]*),([^\x03]*)/\x02\1\x04\2/g
tMC
s/^,// # remove spurious leading delimiter
'
Run Code Online (Sandbox Code Playgroud)
给定以下格式的空格分隔文件:
1 2 3 4 5
Run Code Online (Sandbox Code Playgroud)
您可以像这样使用 awk 删除字段 2:
awk '{ sub($2,""); print}' file
Run Code Online (Sandbox Code Playgroud)
返回
1 3 4 5
Run Code Online (Sandbox Code Playgroud)
在适当的情况下用第 n 列替换第 2 列。
要复制第 2 列,
awk '{ col = $2 " " $2; $2 = col; print }' file
Run Code Online (Sandbox Code Playgroud)
返回
1 2 2 3 4 5
Run Code Online (Sandbox Code Playgroud)
要切换第 2 列和第 3 列,
awk '{temp = $2; $2 = $3; $3 = temp; print}'
Run Code Online (Sandbox Code Playgroud)
返回
1 3 2 4 5
Run Code Online (Sandbox Code Playgroud)
awk 通常非常擅长处理字段的概念。如果您正在处理 CSV 而不是空格分隔的文件,您可以简单地使用
awk -F,
Run Code Online (Sandbox Code Playgroud)
将您的字段定义为逗号,而不是空格(这是默认值)。网上有很多不错的 awk 资源,我在下面列出了其中的一个资源。
#3 的来源