小编khr*_*a_s的帖子

为什么 sklearn tf-idf 矢量器给停用词最高分?

我使用 sklearn 为 nltk 库中 Brown 语料库的每个类别实现了 Tf-idf。有 15 个类别,每个类别的最高分都分配给一个停用词。

默认参数是use_idf=True,所以我使用 idf 。语料库足够大,可以计算出正确的分数。所以,我不明白 - 为什么停用词被赋予高值?

import nltk, sklearn, numpy
import pandas as pd
from nltk.corpus import brown, stopwords
from sklearn.feature_extraction.text import TfidfVectorizer

nltk.download('brown')
nltk.download('stopwords')

corpus = []
for c in brown.categories():
  doc = ' '.join(brown.words(categories=c))
  corpus.append(doc)

thisvectorizer = TfidfVectorizer()
X = thisvectorizer.fit_transform(corpus)
tfidf_matrix = X.toarray()
features = thisvectorizer.get_feature_names_out()

for array in tfidf_matrix:
  tfidf_per_doc = list(zip(features, array))
  tfidf_per_doc.sort(key=lambda x: x[1], reverse=True)
  print(tfidf_per_doc[:3])
Run Code Online (Sandbox Code Playgroud)

结果是:

[('the', 0.6893251240111703), ('and', 0.31175508121108203), ('he', 0.24393467757919754)] …
Run Code Online (Sandbox Code Playgroud)

python nltk tf-idf scikit-learn tfidfvectorizer

6
推荐指数
1
解决办法
1064
查看次数

标签 统计

nltk ×1

python ×1

scikit-learn ×1

tf-idf ×1

tfidfvectorizer ×1