对于在Python脚本中运行Scrapy感到困惑

hh5*_*188 6 python scrapy web-scraping

在下面的文档中,我可以从Python脚本运行scrapy,但是我无法获得scrapy结果.

这是我的蜘蛛:

from scrapy.spider import BaseSpider
from scrapy.selector import HtmlXPathSelector
from items import DmozItem

class DmozSpider(BaseSpider):
    name = "douban" 
    allowed_domains = ["example.com"]
    start_urls = [
        "http://www.example.com/group/xxx/discussion"
    ]

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        rows = hxs.select("//table[@class='olt']/tr/td[@class='title']/a")
        items = []
        # print sites
        for row in rows:
            item = DmozItem()
            item["title"] = row.select('text()').extract()[0]
            item["link"] = row.select('@href').extract()[0]
            items.append(item)

        return items
Run Code Online (Sandbox Code Playgroud)

注意最后一行,我尝试使用返回的解析结果,如果我运行:

 scrapy crawl douban
Run Code Online (Sandbox Code Playgroud)

终端可以打印返回结果

但是我无法从Python脚本中获得返回结果.这是我的Python脚本:

from twisted.internet import reactor
from scrapy.crawler import Crawler
from scrapy.settings import Settings
from scrapy import log, signals
from spiders.dmoz_spider import DmozSpider
from scrapy.xlib.pydispatch import dispatcher

def stop_reactor():
    reactor.stop()
dispatcher.connect(stop_reactor, signal=signals.spider_closed)
spider = DmozSpider(domain='www.douban.com')
crawler = Crawler(Settings())
crawler.configure()
crawler.crawl(spider)
crawler.start()
log.start()
log.msg("------------>Running reactor")
result = reactor.run()
print result
log.msg("------------>Running stoped")
Run Code Online (Sandbox Code Playgroud)

我试图得到结果reactor.run(),但它什么也没有返回,

我怎样才能得到结果?

ale*_*cxe 8

终端打印结果,因为默认日志级别设置为DEBUG.

从脚本运行蜘蛛并调用时log.start(),默认日志级别设置为INFO.

只需更换:

log.start()
Run Code Online (Sandbox Code Playgroud)

log.start(loglevel=log.DEBUG)
Run Code Online (Sandbox Code Playgroud)

UPD:

要将结果作为字符串,您可以将所有内容记录到文件中,然后从中读取,例如:

log.start(logfile="results.log", loglevel=log.DEBUG, crawler=crawler, logstdout=False)

reactor.run()

with open("results.log", "r") as f:
    result = f.read()
print result
Run Code Online (Sandbox Code Playgroud)

希望有所帮助.

  • 该许可显示输出但不是从那里收集数据我认为正确的方法是写或使用管道. (2认同)

Ixi*_*xio 5

我在问自己同样的事情时发现了你的问题,即:“我怎样才能得到结果?”。由于这里没有回答,我努力自己找到答案,现在我可以分享它:

items = []
def add_item(item):
    items.append(item)
dispatcher.connect(add_item, signal=signals.item_passed)
Run Code Online (Sandbox Code Playgroud)

或者对于scrapy 0.22 ( http://doc.scrapy.org/en/latest/topics/practices.html#run-scrapy-from-a-script ) 将我的解决方案的最后一行替换为:

crawler.signals.connect(add_item, signals.item_passed)
Run Code Online (Sandbox Code Playgroud)

我的解决方案是从http://www.tryolabs.com/Blog/2011/09/27/calling-scrapy-python-script/自由改编的。