从Hit/Hits迁移到TopDocs/TopDocCollector

Pau*_*cas 13 java lucene

我有现有的代码,如:

final Term t = /* ... */;
final Iterator i = searcher.search( new TermQuery( t ) ).iterator();
while ( i.hasNext() ) {
    Hit hit = (Hit)i.next();
    // "FILE" is the field that recorded the original file indexed
    File f = new File( hit.get( "FILE" ) );
    // ...
}
Run Code Online (Sandbox Code Playgroud)

我不清楚如何使用TopDocs/ 重写代码TopDocCollector以及如何迭代所有结果.

its*_*dok 24

基本上,您必须决定您期望的结果数量的限制.然后迭代结果中的所有ScoreDocs TopDocs.

final MAX_RESULTS = 10000;
final Term t = /* ... */;
final TopDocs topDocs = searcher.search( new TermQuery( t ), MAX_RESULTS );
for ( ScoreDoc scoreDoc : topDocs.scoreDocs ) {
    Document doc = searcher.doc( scoreDoc.doc )
    // "FILE" is the field that recorded the original file indexed
    File f = new File( doc.get( "FILE" ) );
    // ...
}
Run Code Online (Sandbox Code Playgroud)

这基本上是Hits类所做的,只是它将限制设置为50个结果,如果你迭代过去,那么重复搜索,这通常是浪费的.这就是它被弃用的原因.

增加:如果没有限制你可以加上结果的数量,你应该使用HitCollector:

final Term t = /* ... */;
final ArrayList<Integer> docs = new ArrayList<Integer>();
searcher.search( new TermQuery( t ), new HitCollector() {
    public void collect(int doc, float score) {
        docs.add(doc);
    }
});

for(Integer docid : docs) {
    Document doc = searcher.doc(docid);
    // "FILE" is the field that recorded the original file indexed
    File f = new File( doc.get( "FILE" ) );
    // ...
}
Run Code Online (Sandbox Code Playgroud)