scrapy-splash如何处理无限滚动?

Bow*_*Liu 5 scrapy scrapy-splash splash-js-render

我想对通过在网页中向下滚动生成的内容进行逆向工程.问题出在网址上https://www.crowdfunder.com/user/following_page/80159?user_id=80159&limit=0&per_page=20&screwrand=933.screwrand似乎没有遵循任何模式,因此撤销网址不起作用.我正在考虑使用Splash进行自动渲染.如何使用Splash滚动浏览器?非常感谢!以下是两个请求的代码:

request1 = scrapy_splash.SplashRequest('https://www.crowdfunder.com/user/following/{}'.format(user_id),
                                                        self.parse_follow_relationship,
                                                        args={'wait':2},
                                                        meta={'user_id':user_id, 'action':'following'},
                                                        endpoint='http://192.168.99.100:8050/render.html')
yield request1

request2 = scrapy_splash.SplashRequest('https://www.crowdfunder.com/user/following_user/80159?user_id=80159&limit=0&per_page=20&screwrand=76',
                                                    self.parse_tmp,
                                                    meta={'user_id':user_id, 'action':'following'},
                                                    endpoint='http://192.168.99.100:8050/render.html')
yield request2
Run Code Online (Sandbox Code Playgroud)

浏览器控制台中显示的ajax请求

Mik*_*bov 13

要滚动页面,您可以编写自定义渲染脚本(请参阅http://splash.readthedocs.io/en/stable/scripting-tutorial.html),如下所示:

function main(splash)
    local num_scrolls = 10
    local scroll_delay = 1.0

    local scroll_to = splash:jsfunc("window.scrollTo")
    local get_body_height = splash:jsfunc(
        "function() {return document.body.scrollHeight;}"
    )
    assert(splash:go(splash.args.url))
    splash:wait(splash.args.wait)

    for _ = 1, num_scrolls do
        scroll_to(0, get_body_height())
        splash:wait(scroll_delay)
    end        
    return splash:html()
end
Run Code Online (Sandbox Code Playgroud)

要渲染此脚本,请使用'execute'端点而不是render.html端点:

script = """<Lua script> """
scrapy_splash.SplashRequest(url, self.parse,
                            endpoint='execute', 
                            args={'wait':2, 'lua_source': script}, ...)
Run Code Online (Sandbox Code Playgroud)