剥离字符串,但允许变音符号

Tim*_* B. 2 regex perl diacritics

#!/usr/bin/perl -T
use strict;
use warnings;
use utf8;
my $s = shift || die;
$s =~ s/[^A-Za-z ]//g;
print "$s\n";
exit;

> ./poc.pl "El Guapö"
El Guap
Run Code Online (Sandbox Code Playgroud)

有没有办法修改这个Perl代码,以便不删除各种变音符号和字符重音?谢谢!

zdi*_*dim 7

对于直接问题,您可能只需要\p{L}(字母)Unicode字符属性

但是,更重要的是,解码所有输入和编码输出.

use warnings;
use strict;
use feature 'say';

use utf8;   # allow non-ascii (UTF-8) characters in the source

use open ':std', ':encoding(UTF-8)';  # for standard streams

use Encode qw(decode_utf8);           # @ARGV escapes the above

my $string = 'El Guapö';
if (@ARGV) {
    $string = join ' ', map { decode_utf8($_) } @ARGV;
}
say "Input:     $string";

$string =~ s/[^\p{L} ]//g;

say "Processed: $string";
Run Code Online (Sandbox Code Playgroud)

当运行时   script.pl 123 El Guapö=_

Input:     123 El Guapö=_
Processed:  El Guapö

我使用了"毯子" \p{L}属性(Letter),因为缺乏具体的描述; 如果/需要调整.Unicode属性提供了很多,请参阅上面的链接和perluniprops上的完整列表.

123 El遗骸之间的空间,可能最终剥离前导(和尾随)空间.

请注意,也有\P{L}资本P表示否定的地方.


上述简单的意思\pL不适用于组合变音符号,因为标记也将被删除.感谢jm666指出这一点.

当重音的"逻辑"字符(显示为单个字符)是使用单独的字符为其基础和非间距标记(组合重音)编写时,会发生这种情况.通常,它的单个字符(扩展字形集群)及其代码点也存在.

例如:在niño所述ñIS U+OOF1,但它也可以被写为"n\x{303}".

要保持以这种方式编写的重音符号,请将\p{Mn}(\p{NonspacingMark})添加到字符类中

my $string = "El Guapö=_ ni\N{U+00F1}o.* nin\x{303}o+^";
say $string;

(my $nodiac = $string) =~ s/[^\pL ]//g;      #/ naive, accent chars get removed
say $nodiac;

(my $full = $string) =~ s/[^\pL\p{Mn} ]//g;  # add non-spacing mark
say $full;
Run Code Online (Sandbox Code Playgroud)

产量

El Guapö=_  niño.* niño+^
El Guapö niño nino
El Guapö niño niño

所以你想要s/[^\p{L}\p{Mn} ]//g保持组合的口音.