小编Shu*_*jaj的帖子

tf-idf 模型如何处理测试数据期间看不见的单词?

我已经阅读了很多博客,但对答案并不满意,假设我在几个文档示例上训练 tf-idf 模型:

   " John like horror movie."
   " Ryan watches dramatic movies"
    ------------so on ----------
Run Code Online (Sandbox Code Playgroud)

我使用这个功能:

   from sklearn.feature_extraction.text import TfidfTransformer
   count_vect = CountVectorizer()
   X_train_counts = count_vect.fit_transform(twenty_train.data)
   X_train_tfidf = tfidf_transformer.fit_transform(X_train_counts)
   print((X_train_counts.todense()))
   # Gives count of words in each document

   But it doesn't tell which word? How to get words as headers in X_train_counts 
  outputs. Similarly in X_train_tfidf ?
Run Code Online (Sandbox Code Playgroud)

所以 X_train_tfidf 输出将是具有 tf-idf 分数的矩阵:

     Horror  watch  movie  drama
doc1  score1  --    -----------
doc2   ------------------------
Run Code Online (Sandbox Code Playgroud)

这样对吗?

做什么fit和做什么transformation?在 sklearn …

tf-idf python-3.x scikit-learn

3
推荐指数
1
解决办法
1778
查看次数

标签 统计

python-3.x ×1

scikit-learn ×1

tf-idf ×1