如何截断字符串最多包含N个字符?

Pet*_*nak 5 string unicode truncate rust

预期的方法String.truncate(usize)失败,因为它不考虑Unicode字符(考虑到Rust将字符串视为Unicode,这令人困惑).

let mut s = "??????".to_string();
s.truncate(4);
Run Code Online (Sandbox Code Playgroud)

线程''恐慌'断言失败:self.is_char_boundary(new_len)'

此外,truncate修改原始字符串,这并不总是需要.

我提出的最好的是转换为chars并收集到String.

fn truncate(s: String, max_width: usize) -> String {
    s.chars().take(max_width).collect()
}
Run Code Online (Sandbox Code Playgroud)

例如

fn main() {
    assert_eq!(truncate("??????".to_string(), 0), "");
    assert_eq!(truncate("??????".to_string(), 4), "????");
    assert_eq!(truncate("??????".to_string(), 100), "??????");
    assert_eq!(truncate("hello".to_string(), 4), "hell");
}
Run Code Online (Sandbox Code Playgroud)

然而,这感觉非常沉重.

She*_*ter 13

请务必阅读并理解德尔南的观点:

Unicode非常复杂.您确定要char(对应于代码点)作为单位而不是字形集群吗?

这个答案的其余部分假定您有充分的理由使用char而不是字形.

考虑到Rust将字符串视为Unicode,这令人费解

这不正确; Rust将字符串视为UTF-8.在UTF-8中,每个代码点都映射到可变数量的字节.没有O(1)算法将"6个字符"转换为"N个字节",因此标准库不会向您隐藏.

您可以使用char_indices逐个字符逐个字符串,并获取该字符的字节索引:

fn truncate(s: &str, max_chars: usize) -> &str {
    match s.char_indices().nth(max_chars) {
        None => s,
        Some((idx, _)) => &s[..idx],
    }
}

fn main() {
    assert_eq!(truncate("??????", 0), "");
    assert_eq!(truncate("??????", 4), "????");
    assert_eq!(truncate("??????", 100), "??????");
    assert_eq!(truncate("hello", 4), "hell");
}
Run Code Online (Sandbox Code Playgroud)

这也会返回一个切片,如果需要,可以选择移动到新的分配中,或者String在适当的位置进行变异:

// May not be as efficient as inlining the code...
fn truncate_in_place(s: &mut String, max_chars: usize) {
    let bytes = truncate(&s, max_chars).len();
    s.truncate(bytes);
}

fn main() {
    let mut s = "??????".to_string();
    truncate_in_place(&mut s, 0);
    assert_eq!(s, "");
}
Run Code Online (Sandbox Code Playgroud)

  • @Peter`chars`只返回字符.`char_indices`在概念上类似于`chars().enumerate()`,除了它返回字符在原始`str`中开始的`u8`的实际索引. (2认同)
  • @Veedrac 每个。单身的。时间。我永远不会记得它![Clippy 功能请求](https://github.com/Manishearth/rust-clippy/issues/1112)! (2认同)