在句子列表中标记单词Python

Hyp*_*nja 11 text nltk python-2.7

我目前有一个文件,其中包含一个看起来像的列表

example = ['Mary had a little lamb' , 
           'Jack went up the hill' , 
           'Jill followed suit' ,    
           'i woke up suddenly' ,
           'it was a really bad dream...']
Run Code Online (Sandbox Code Playgroud)

"example"是这样的句子列表,我希望输出看起来像:

mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill' ....] 等等.我需要将每个单词标记为单独的句子,以便我可以将mod_example(一次使用for循环)的句子中的每个单词与参考句子进行比较.

我试过这个:

for sentence in example:
    text3 = sentence.split()
    print text3 
Run Code Online (Sandbox Code Playgroud)

得到了以下输出:

['it', 'was', 'a', 'really', 'bad', 'dream...']
Run Code Online (Sandbox Code Playgroud)

我如何为所有句子得到这个? 它会一直覆盖.是的,还要提一下我的方法是否正确?这应该是一个单词列表,标记为tokenized ..谢谢

alv*_*vas 24

你可以使用NLTK中的单词tokenizer(http://nltk.org/api/nltk.tokenize.html)和列表理解,参见http://docs.python.org/2/tutorial/datastructures.html#list -comprehensions

>>> from nltk.tokenize import word_tokenize
>>> example = ['Mary had a little lamb' , 
...            'Jack went up the hill' , 
...            'Jill followed suit' ,    
...            'i woke up suddenly' ,
...            'it was a really bad dream...']
>>> tokenized_sents = [word_tokenize(i) for i in example]
>>> for i in tokenized_sents:
...     print i
... 
['Mary', 'had', 'a', 'little', 'lamb']
['Jack', 'went', 'up', 'the', 'hill']
['Jill', 'followed', 'suit']
['i', 'woke', 'up', 'suddenly']
['it', 'was', 'a', 'really', 'bad', 'dream', '...']
Run Code Online (Sandbox Code Playgroud)

  • 我强烈建议不要使用 NLTK。尽管很受欢迎(因为它是第一个有据可查的 Python NLP 包),但它已经过时了。此外,`word_tokenize` 有转换输入的习惯。 (2认同)

Wah*_*ram 8

我制作这个脚本是为了让所有人都了解如何标记,这样他们就可以自己构建自然语言处理引擎。

import re
from contextlib import redirect_stdout
from io import StringIO

example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'

def token_to_sentence(str):
    f = StringIO()
    with redirect_stdout(f):
        regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
        regex_of_sentence = [x for x in regex_of_sentence if x is not '']
        for i in regex_of_sentence:
            print(i)
        first_step_to_sentence = (f.getvalue()).split('\n')
    g = StringIO()
    with redirect_stdout(g):
        for i in first_step_to_sentence:
            try:
                regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
                print(regex_to_clear_sentence.group(1))
            except:
                print(i)
        sentence = (g.getvalue()).split('\n')
    return sentence

def token_to_words(str):
    f = StringIO()
    with redirect_stdout(f):
        for i in str:
            regex_of_word = re.findall('([\w]{0,})', i)
            regex_of_word = [x for x in regex_of_word if x is not '']
            for word in regex_of_word:
                print(regex_of_word)
        words = (f.getvalue()).split('\n')
Run Code Online (Sandbox Code Playgroud)

我做了一个不同的过程,我从段落重新开始这个过程,让大家更了解文字处理。要处理的段落是:

example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
Run Code Online (Sandbox Code Playgroud)

将段落标记为句子:

sentence = token_to_sentence(example)
Run Code Online (Sandbox Code Playgroud)

将导致:

['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
Run Code Online (Sandbox Code Playgroud)

标记为单词:

words = token_to_words(sentence)
Run Code Online (Sandbox Code Playgroud)

将导致:

['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
Run Code Online (Sandbox Code Playgroud)

我将解释这是如何工作的。

首先,我使用正则表达式搜索所有分隔单词并停止直到找到标点符号的单词和空格,正则表达式是:

([\w\s]{0,})[^\w\s]{0,}
Run Code Online (Sandbox Code Playgroud)

所以计算将采用括号中的单词和空格:

'(Mary had a little lamb),( Jack went up the hill, Jill followed suit),( i woke up suddenly),( it was a really bad dream)...'
Run Code Online (Sandbox Code Playgroud)

结果仍然不清楚,包含一些“无”字符。所以我用这个脚本删除了“无”字符:

[x for x in regex_of_sentence if x is not '']
Run Code Online (Sandbox Code Playgroud)

所以段落将标记为句子,但不明确的句子结果是:

['Mary had a little lamb', ' Jack went up the hill', ' Jill followed suit', ' i woke up suddenly', ' it was a really bad dream']
Run Code Online (Sandbox Code Playgroud)

如您所见,结果显示一些句子以空格开头。所以为了在不开始空格的情况下制作一个清晰的段落,我制作了这个正则表达式:

\s([\w\s]{0,})
Run Code Online (Sandbox Code Playgroud)

它会做出一个清晰的句子,例如:

['Mary had a little lamb', 'Jack went up the hill', 'Jill followed suit', 'i woke up suddenly', 'it was a really bad dream']
Run Code Online (Sandbox Code Playgroud)

所以,我们必须做两个过程才能做出好的结果。

你的问题的答案是从这里开始......

为了将句子标记为单词,我进行了段落迭代并使用正则表达式来捕获单词,同时使用此正则表达式进行迭代:

([\w]{0,})
Run Code Online (Sandbox Code Playgroud)

并再次清除空字符:

[x for x in regex_of_word if x is not '']
Run Code Online (Sandbox Code Playgroud)

所以结果真的很清楚,只有单词列表:

['Mary', 'had', 'a', 'little', 'lamb', 'Jack', 'went, 'up', 'the', 'hill', 'Jill', 'followed', 'suit', 'i', 'woke', 'up', 'suddenly', 'it', 'was', 'a', 'really', 'bad', 'dream']
Run Code Online (Sandbox Code Playgroud)

以后要做一个好的NLP,你需要有自己的词组数据库,搜索词组是否在句子中,列出词组后,剩下的词就是一个词了。

使用这种方法,我可以用我的语言(印度尼西亚语)构建我自己的 NLP,这真的非常缺乏模块。

编辑:

我没有看到你想要比较的话的问题。所以你还有一句话要比较......我给你的​​奖金不仅是奖金,我还告诉你如何计算它。

mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']
Run Code Online (Sandbox Code Playgroud)

在这种情况下,您必须执行的步骤是: 1. 迭代 mod_example 2. 将第一个句子与 mod_example 中的单词进行比较。3.做一些计算

所以脚本将是:

import re
from contextlib import redirect_stdout
from io import StringIO

example = 'Mary had a little lamb, Jack went up the hill, Jill followed suit, i woke up suddenly, it was a really bad dream...'
mod_example = ["'Mary' 'had' 'a' 'little' 'lamb'" , 'Jack' 'went' 'up' 'the' 'hill']

def token_to_sentence(str):
    f = StringIO()
    with redirect_stdout(f):
        regex_of_sentence = re.findall('([\w\s]{0,})[^\w\s]', str)
        regex_of_sentence = [x for x in regex_of_sentence if x is not '']
        for i in regex_of_sentence:
            print(i)
        first_step_to_sentence = (f.getvalue()).split('\n')
    g = StringIO()
    with redirect_stdout(g):
        for i in first_step_to_sentence:
            try:
                regex_to_clear_sentence = re.search('\s([\w\s]{0,})', i)
                print(regex_to_clear_sentence.group(1))
            except:
                print(i)
        sentence = (g.getvalue()).split('\n')
    return sentence

def token_to_words(str):
    f = StringIO()
    with redirect_stdout(f):
        for i in str:
            regex_of_word = re.findall('([\w]{0,})', i)
            regex_of_word = [x for x in regex_of_word if x is not '']
            for word in regex_of_word:
                print(regex_of_word)
        words = (f.getvalue()).split('\n')

def convert_to_words(str):
    sentences = token_to_sentence(str)
    for i in sentences:
        word = token_to_words(i)
    return word

def compare_list_of_words__to_another_list_of_words(from_strA, to_strB):
        fromA = list(set(from_strA))
        for word_to_match in fromA:
            totalB = len(to_strB)
            number_of_match = (to_strB).count(word_to_match)
            data = str((((to_strB).count(word_to_match))/totalB)*100)
            print('words: -- ' + word_to_match + ' --' + '\n'
            '       number of match    : ' + number_of_match + ' from ' + str(totalB) + '\n'
            '       percent of match   : ' + data + ' percent')



#prepare already make, now we will use it. The process start with script below:

if __name__ == '__main__':
    #tokenize paragraph in example to sentence:
    getsentences = token_to_sentence(example)

    #tokenize sentence to words (sentences in getsentences)
    getwords = token_to_words(getsentences)

    #compare list of word in (getwords) with list of words in mod_example
    compare_list_of_words__to_another_list_of_words(getwords, mod_example)
Run Code Online (Sandbox Code Playgroud)