无法从网页中获取某些标题

SIM*_*SIM 10 php curl domdocument web-scraping

我在php中编写了一个脚本,从网页上刮下一个可以看作头发掉落标题.当我执行下面的脚本时,我收到以下错误:

注意:尝试在第16行的C:\ xampp\htdocs\runco​​de\testfile.php中获取非对象的属性"nodeValue".

链接到该网站

我试过的脚本:

<?php
    function get_content($url){
        $ch = curl_init();
        curl_setopt($ch, CURLOPT_URL, $url);
        curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0');
        curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
        curl_exec($ch);
        $htmlContent = curl_exec($ch);
        curl_close($ch);
        return $htmlContent;
    }
    $link = "https://www.purplle.com/search?q=hair%20fall%20shamboo"; 
    $xml = get_content($link);
    $dom = @DOMDocument::loadHTML($xml);
    $xpath = new DOMXPath($dom);
    $title = $xpath->query('//h1[@class="br-hdng"]/span')->item(0)->nodeValue;
    echo "{$title}";
?>
Run Code Online (Sandbox Code Playgroud)

我的预期输出是:

hair fall shamboo
Run Code Online (Sandbox Code Playgroud)

虽然xpath我在上面的脚本中使用的似乎是正确的,但我在这里粘贴了title可以找到的html元素的相关部分:

<h1 _ngcontent-c0="" class="br-hdng"><span _ngcontent-c0="" class="pr dib">hair fall shamboo<!----></span></h1>
Run Code Online (Sandbox Code Playgroud)

后记:title我要解析获取动态加载.由于我是php的新手,我不明白我尝试的方式是否准确.如果不是我应该做什么呢?

以下是我使用两种不同语言创建的脚本,发现它们像魔术一样工作.

我成功使用javascript:

const puppeteer = require('puppeteer');
function run () {
    return new Promise(async (resolve, reject) => {
        try {
            const browser = await puppeteer.launch();
            const page = await browser.newPage();
            await page.goto("https://www.purplle.com/search?q=hair%20fall%20shamboo");
            let urls = await page.evaluate(() => {
            let items = document.querySelector('h1.br-hdng span');
            return items.innerText;;
            })
            browser.close();
            return resolve(urls);
        } catch (e) {
            return reject(e);
        }
    })
}
run().then(console.log).catch(console.error);
Run Code Online (Sandbox Code Playgroud)

我再次成功使用python:

import requests_html

with requests_html.HTMLSession() as session:
    r = session.get('https://www.purplle.com/search?q=hair%20fall%20shamboo')
    r.html.render()
    item = r.html.find("h1.br-hdng span",first=True).text
    print(item)
Run Code Online (Sandbox Code Playgroud)

那怎么了php

Dec*_*ler 7

很可能你的代码存在的问题多于我在这个答案中所涉及的问题,但我看到的最突出的问题如下:

DOMDocument::loadHTML()不是静态方法,而是实例方法(返回布尔值).您应首先创建一个实例,DOMDocument然后调用loadHTML()该实例:

$dom = new DOMDocument;
$dom->loadHTML($xml);
Run Code Online (Sandbox Code Playgroud)

但是,由于您已@在该特定行上抑制了操作员的错误,因此您没有收到有关此操作的警告.虽然很常见的是错误抑制器运算符@用于抑制HTML验证错误,但是你应该考虑使用1,因为这不会抑制一般的PHP错误.libxml_use_internal_errors()

$dom = new DOMDocument;
$oldSetting = libxml_use_internal_errors(true);
$dom->loadHTML($xml);
libxml_use_internal_errors($oldSetting);
Run Code Online (Sandbox Code Playgroud)

最后要注意:如果您的PHP安装配置为允许通过配置设置加载URL,则可以
直接从URL加载DOM文档(无需cURL).请注意,出于安全原因,此设置通常会被禁用,因此如果您打算使用它,请小心使用它.DOMDocument::loadHTMLFile()allow_url_fopen


这是一个简单的测试用例,应该按预期工作:

<?php

$html = '
<html>
<head>
  <title>DOMDocument test-case</title>
</head>
<body>
  <div class="dummy-container">
    <h1 _ngcontent-c0="" class="br-hdng"><span _ngcontent-c0="" class="pr dib">hair fall shamboo<!----></span></h1>
  </div>
</body>';

$dom = new DOMDocument;

$oldSetting = libxml_use_internal_errors(true);
$dom->loadHTML( $html );
libxml_use_internal_errors($oldSetting);

$xpath = new DOMXPath( $dom );
$title = $xpath->query( '//h1[@class="br-hdng"]/span' )->item( 0 )->nodeValue;
echo $title;
Run Code Online (Sandbox Code Playgroud)

请参阅3v4l.org上在线解释的此示例

您应该用呼叫$html输出替换内容get_content().如果它不起作用,那么:

  1. 使用cURL(例如,var_dump( $html );在加载之前执行DOMDocument以查看您检索的内容之前执行操作)或者...

  2. 也许你正在一个命名空间内工作,在这种情况下你应该在之前DOMDocument和之前加一个反斜杠DOMXPath,即:new \DOMDocument;new \DOMXPath( $dom );.


1. LibXML是DOMDocument用于解析XML/HTML文档的XML库.


han*_*rik 5

那么php出了什么问题?

php不运行javascript.据推测,puppeteer从你的javascript代码,以及你的python代码中的requests_html,都运行javascript.

你的问题是这个页面br-hdng用javascript 加载标题和产品,它根本不是HTML的一部分.它实际上是从https://www.purplle.com/api/shop/itemsv3一堆GET参数加载而来的.你需要在这里进行JSON解析,而不是HTML解析:)但是在你可以访问那个api之前,你需要搜索页面给出的cookie,搜索字符串必须匹配api搜索字符串(否则api只返回错误),检查一下:

<?php
declare(strict_types = 0);
header ( "Content-Type: text/plain;charset=UTF-8" );
$ch = curl_init ();
curl_setopt_array ( $ch, array (
        CURLOPT_ENCODING => '',
        CURLOPT_COOKIEFILE => '', // enables cookie handling without saving them anywhere. this page requires cookie handling.
        CURLOPT_USERAGENT => 'Mozilla/5.0 (Windows NT 6.1; Win64; x64; rv:60.0) Gecko/20100101 Firefox/60.0', // 'libcurl/? PHP/' . PHP_VERSION, // many websites block requests without a useragent
        CURLOPT_RETURNTRANSFER => 1 
) );
// we don't care what's on this page, we just need to fetch it to create a cookie session.
$search_query = 'hair fall shamboo';
curl_setopt ( $ch, CURLOPT_URL, 'https://www.purplle.com/search?q=' . rawurlencode ( $search_query ) );
curL_exec ( $ch );
$url = 'https://www.purplle.com/api/shop/itemsv3?' . http_build_query ( array (
        'list_type' => 'search',
        'custom' => '',
        'list_type_value' => $search_query,
        'page' => '1',
        'sort_by' => 'rel',
        'elite' => '0' 
) );
// $url = 'https://www.purplle.com/api/shop/itemsv3?list_type=search&custom=&list_type_value=hair%20fall%20shamboo&page=1&sort_by=rel&elite=0';
// $out = tmpfile ();
// curl_setopt_array ( $ch, array (
// CURLOPT_HTTPHEADER => array (
// 'Accept: application/json, text/plain, */*',
// 'Accept-Language: en-US,en;q=0.5',
// 'Referer: https://www.purplle.com/search?q=hair%20fall%20shamboo',
// // Cookie: __cfduid=d3199415b5ce18cbff2779802b1f843331544901552; csrftoken=f8f18b5deae92972f63343e13c6a460b; purpllesession=hedxkc%2FkdGye%2BYi6ebmJktUN1LeqA5rdVXu96%2F0j0yqtP2xZ8LfwpK8daXqPSkeZulO9ZvqpMYXTmY8oMD03VcG9vdKGBm30R9fU%2FQygtXBFhZvfvsu0scyaL3FqHbePp2zG45MevWU961eg82KAkCuHk0qFM8URQBRyYV5gg8TeqnTPgI3tF87H5nJ%2BmfO4pn%2BRWmIuWXvgNXAO%2F8GEaH6lJVl17QZm9c5vwi10OYeLfmSdIMy6V2Pp0ZjLTzuFw2de5jpR0zsbHHKZ0C2e548PiDl3taHIE5wuZO4HYIeXUqTpE98%2Fo3kztoU1bTlXGZgu%2FxVQ3EWLRFWQ2t57UawA%2FuERlD8vvOyFGbYHGAWVxgFTR%2FObAhFLHns5kqoj; _autm30d=null; visitorppl=NZ5tqQpGlFYWg2MrDl1302113161544901552; session_initiated=Direct; _tmpsess=1; token=desktop_5c1553b07c61c_7955_16122018; __uzma=5c1553b085a480.63440826; __uzmb=1544901552; __uzmc=632121030774; __uzmd=1544901552
// 'Connection: keep-alive'
// ),
// // CURLOPT_CONNECT_TO=>array('www.purplle.com:443:dumpinput.ratma.net:80'),
// CURLOPT_STDERR => $out,
// CURLOPT_VERBOSE => 1
// ) );
// var_dump ( $url );
curl_setopt ( $ch, CURLOPT_URL, $url );
$json = curl_exec ( $ch );
$data = json_decode ( $json, true );

// var_dump ($json, $data );
$title = $data ['list_title'];
echo 'title: ' . $title . "\n";
foreach ( $data ['items'] as $item ) {
    echo "name: ", $item ['name'], "\n";
}
Run Code Online (Sandbox Code Playgroud)

输出:

title: hair fall shamboo
name: VLCC Hair fall Shampoo 350 ML (Buy1 Get1) & Ayurveda Hair Oil Combo (470 ml)
name: Biotique Bio Kelp Protein Shampoo For Falling Hair (190 ml)
name: Biotique Fresh Texture Shampoo - Bio Henna Leaf (120 ml)
name: Good Vibes Scalp Purifying Shampoo -Neem And Aloe Vera (200 ml)
name: Khadi Shikakai Sat Hair Cleanser Scalp Therapy (210 ml) By Swati Gramodyog
name: Good Vibes Apple Cider Vinegar Shampoo (120 ml)
name: Good Vibes Refreshing Shampoo - Green Apple (200 ml)
name: Good Vibes Hydrating Shampoo -Marigold (200 ml)
name: Alps Goodness Smoothening Shampoo - Keratin (50 ml)
name: Alps Goodness Softening Shampoo - Coconut & Almond (50 ml)
name: Alps Goodness Split End Control Shampoo - Coconut, Garlic & Shea Butter (50 ml)
name: Passion Indulge Papain Shampoo & Conditioner For Soft & Shiny Hair (200 ml + 100 ml)
name: Good Vibes Apple Cider Vinegar Shampoo (200 ml)
name: Alps Goodness Split End Control Shampoo - Coconut, Garlic & Shea Butter (200 ml)
name: Alps Goodness Nourishing Shampoo - Argan Oil & Olive (200 ml)
name: Alps Goodness Moisturizing Shampoo - Ginger & Egg (200 ml)
name: Alps Goodness Conditioning Shampoo - Pure Honey (200 ml)
name: Alps Goodness Hydrating Shampoo - Tea Tree (200 ml)
name: Alps Goodness Smoothening Shampoo - Keratin (200 ml)
name: Alps Goodness Softening Shampoo - Coconut & Almond (200 ml)
name: Good Vibes Scalp Purifying Shampoo -Neem And Aloe Vera (120 ml)
name: Good Vibes Hydrating Shampoo - Marigold (120 ml)
name: Alps Goodness Conditioning Shampoo - Pure Honey (50 ml)
name: Alps Goodness Moisturizing Shampoo - Ginger & Egg (50 ml)
Run Code Online (Sandbox Code Playgroud)