JS 字符串中的行结尾(也称为换行符)

joh*_* j. 5 javascript regex eol

众所周知,类 Unix 系统使用LF字符作为换行符,而 Windows 使用CR+LF.

然而,当我在 Windows PC 上从本地 HTML 文件测试此代码时,JS 似乎将所有换行符视为以LF. 这是正确的假设吗?

var string = `
    foo




    bar
`;

// There should be only one blank line between foo and bar.

// \n - Works
// string = string.replace(/^(\s*\n){2,}/gm, '\n');

// \r\n - Doesn't work
string = string.replace(/^(\s*\r\n){2,}/gm, '\r\n');

alert(string);

// That is, it seems that JS treat all newlines as separated with 
// `LF` instead of `CR+LF`?
Run Code Online (Sandbox Code Playgroud)

wp7*_*8de 4

我想我找到了一个解释。

您正在使用 ES6模板文字来构造多行字符串。

根据ECMAScript规范

.. 模板文字组件被解释为 Unicode 代码点序列。文字组件的模板值 (TV) 根据由模板文字组件的各个部分贡献的代码单元值 (SV, 11.8.4) 进行描述。作为此过程的一部分,模板组件中的一些 Unicode 代码点被解释为具有数学值(MV、11.8.3)。在确定 TV 时,转义序列被转义序列表示的 Unicode 代码点的 UTF-16 代码单元替换。模板原始值 (TRV) 与模板值类似,不同之处在于 TRV 中的转义序列按字面解释。

在此之下,定义为:

LineTerminatorSequence::<LF> 的 TRV 是代码单元 0x000A(换行)。
LineTerminatorSequence::<CR> 的 TRV 是代码单元 0x000A(换行)。

我的解释是,当您使用模板文字时,无论操作系统特定的换行定义如何,您总是只得到一个换行符。

最后,在JavaScript 的正则表达式中

\n 匹配换行符 (U+000A)。

它描述了观察到的行为。

但是,如果您定义字符串文字'\r\n'或从文件流中读取文本等,其中包含特定于操作系统的换行符,则必须处理它。

以下是一些演示模板文字行为的测试:

`a
b`.split('')
  .map(function (char) {
    console.log(char.charCodeAt(0));
  });

(String.raw`a
b`).split('')
  .map(function (char) {
    console.log(char.charCodeAt(0));
  });
  
 'a\r\nb'.split('')
  .map(function (char) {
    console.log(char.charCodeAt(0));
  });
  
"a\
b".split('')
  .map(function (char) {
    console.log(char.charCodeAt(0));
  });
Run Code Online (Sandbox Code Playgroud)

解释结果:
char(97) = a、char(98) = b
char(10) = \n、char(13) =\r