从英语单词中删除字母重复的正确方法?

tal*_*a06 3 nlp linguistics sentiment-analysis

正如标题清楚地描述的那样,我想知道消除英语中人物重复的正确方法是什么,这种重复在社交媒体中常用来夸大这种感觉.由于我正在开发一种软​​件解决方案以纠正错误的单词,我需要一种可应用于大多数英语单词的全局算法.那么,我要求专家学习如何在不使用基于学习的方法的情况下消除英语单词中的其他字母的正确方法?

PS.(1)我使用WordNet 3.0数据库以编程方式检查单词是否有效.到目前为止,除了一些例子,例如在WordNet 3.0中veery定义的单词之外,这么好tawny brown North American trush noted for its song.当在WordNet中找到该单词时,我会中断字母消除过程.那么有没有其他知识库可以用来代替WordNet?

PS.(2)其实我在英语语言和使用社区问了这个问题.但是他们引导我在这里问问题.

一些例子:

haappyy --> happy
amaaazzinng --> amazing
veeerry --> very
Run Code Online (Sandbox Code Playgroud)

正如你在例子中看到的那样,字母重复的地方各不相同.

err*_*ist 5

将带有重复字母的非规范形式映射到其规范形式的一个主要问题是,在没有上下文的情况下它是模糊的:考虑例如meeet - 实际上是满足还是满足?

解决这个问题的一种方法可能是测量一个单词的"正规性"作为在n -gram模型中给出其上下文(即它前面的单词)的概率:被判断为所述单词的"别名"的单词(如下所述)并且最常见,因为手头的背景被视为"规范"形式:

正常化

能够想到的形式,如veeerry和非常作为相同的形式,即经修饰的"的字符有序袋"的变体,使得例如两者都为序列处理('v','e','r','y')尽管具有不同的权重:

  • veeerry:(('v', 1), ('e', 3), ('r', 2), ('y', 1))
  • 非常:(('v', 1), ('e', 1), ('r', 1), ('y', 1))

通过这种方式,可以像普通特征向量一样处理这些形式 - 这将用于加权下面的概率.

别名

正如你所描述的那样,我们需要能够将例如veeerry映射到与非常见的单词veery匹配.因此,简单地表示veeerry为"字符的命令包"没有任何的权重将意味着veeerry会就像是振振有词的一个变种veery -它显然不是.然而,由于veery最有可能有一个非常不同的词汇分配比的副词非常,将有可能得到一个标准化的意义('v','e','r','y')基础上给出它的上下文的概率-例如:

  1. 这辆车非常大.
  2. #这辆车很神圣.("这辆汽车是一种黄褐色的北美画眉,因其歌曲而闻名")

即使没有考虑这两个例子之间的句法词性区分(第一个是副词而第二个是名词),这两个不同词的原始统计分布是非常不同的.出于这个原因,给出了一个好的模型,P(veery_1 | "This car is") > P(veery_2 | "This car is a").

概率加权

为了例如关联veeerry以非常文本发现,同时仍保持当为了规范它veery被归到非常,我们就可以简单的使用特征向量表示字符的有序袋veeerry从两个计算的距离很和veery,然后使用该距离来加权每个在给定上下文的概率:

best_term(form, context) = argmax(arg=term, (P(term, context) * sim(ordered_bag_of_chars(form), ordered_bag_of_chars(term)))
Run Code Online (Sandbox Code Playgroud)

我写了一些Python代码,以便更好地解释这在现实生活中如何起作用:

#!/usr/bin/env python3

from scipy.spatial.distance import cosine

class Vocabulary(object):
    def __init__(self, forms):
        self.forms = forms
        self.form_variations = {}   
        for form in forms:
            unique_symbol_freqs = tuple(count_unique_symbol_freqs(form))
            unique_symbols = tuple(unique_symbol_freq[0] for unique_symbol_freq in unique_symbol_freqs)
            self.form_variations[unique_symbols] = WeightedSymbolSequence(unique_symbol_freqs)

class WeightedSymbolSequence(object):
    def __init__(self, unique_symbols):
        # TODO: Finish implementation
        pass

def count_unique_symbol_freqs(input_seq):
    if len(input_seq) > 0:
        # First process the head symbol
        previous_unique_symbol = (input_seq[0], 1)
        # Process tail symbols; Add extra iteration at the end in order to handle trailing single unique symbols
        tail_iter = iter(input_seq[1:])
        try:
            while True:
                input_symbol = next(tail_iter)
                if input_symbol == previous_unique_symbol[0]:
                    previous_unique_symbol = (previous_unique_symbol[0], previous_unique_symbol[1] + 1)
                else:
                    result_unique_symbol = previous_unique_symbol
                    previous_unique_symbol =  (input_symbol, 1)
                    yield result_unique_symbol
        except StopIteration:
            # The end of the sequence was encountered; Handle a potential last unique symbol
            yield previous_unique_symbol

if __name__ == '__main__':
    from sys import stderr

    tests = {"haappyy" : "happy", "amaaazzinng" : "amazing", "veeerry" : "very"}

    with open("/usr/share/dict/words", "r") as vocab_instream:
        # TODO: Alias uppercase versions of characters to their lowercase ones so that e.g. "dog" and "Dog" are highly similar but not identical
        vocab = Vocabulary(frozenset(line.strip().lower() for line in vocab_instream))

    for form1, form2 in tests.items():
        # First check if the token is an extant word 
        if form1 in vocab.forms:
            print("\"%s\" found in vocabulary list; No need to look for similar words." % form1)
        else:
            form1_unique_char_freqs = tuple(count_unique_symbol_freqs(form1))
            print(form1_unique_char_freqs)
            form2_unique_char_freqs = tuple(count_unique_symbol_freqs(form2))
            dist = cosine([form1_unique_char_freq[1] for form1_unique_char_freq in form1_unique_char_freqs], [form2_unique_char_freq[1] for form2_unique_char_freq in form2_unique_char_freqs])
            # TODO: Get probabilities of form variations using WeightedSymbolSequence objects in "vocab.form_variations" and then multiply the probability of each by the cosine similarity of bag of characters for form1 and form2

            try:
                # Get mapping to other variants, e.g. "haappyy" -> "happy"
                form1_unique_chars = tuple(form1_unique_char_freq[0] for form1_unique_char_freq in form1_unique_char_freqs)
                variations = vocab.form_variations[form1_unique_chars]
            except KeyError:
                # TODO: Use e.g. Levenshtein distance in order to find the next-most similar word: No other variations were found
                print("No variations of \"%s\" found." % form1, file=stderr)
Run Code Online (Sandbox Code Playgroud)