标签: domcrawler

Symfony 2 Dom Crawler:如何在Element中只获取text()

使用Dom Crawler只获取文本(没有标记).

$html = EOT<<<
  <div class="coucu">
    Get Description <span>Coucu</span>
  </div>
EOT;

$crawler = new Crawler($html);
$crawler = $crawler->filter('.coucu')->first()->text();
Run Code Online (Sandbox Code Playgroud)

输出:获取描述Coucu

我想输出(仅):获取描述

更新:

我找到了一个解决方案:(但这是非常糟糕的解决方案)

...
$html = $crawler->filter('.coucu')->html();
// use strip_tags_content in https://php.net/strip_tags
$html = strip_tags_content($html,'span');
Run Code Online (Sandbox Code Playgroud)

symfony domcrawler

8
推荐指数
2
解决办法
5782
查看次数

Goutte - Get inner values from $crawler-&gt;filter()

I am using PHP 7.1.33 and "fabpot/goutte": "^3.2". My composer file looks like the following:

{
    "name": "ubuntu/workspace",
    "require": {
        "fabpot/goutte": "^3.2"
    },
    "authors": [
        {
            "name": "admin",
            "email": "admin@admin.com"
        }
    ]
}

Run Code Online (Sandbox Code Playgroud)

I am trying to get details by a time range from a webpage but struggle how to pass the $crawler-values to my final result array $res1Array.

I tried the following:

<?php
require 'vendor/autoload.php';

use Goutte\Client;
use Symfony\Component\DomCrawler\Crawler;

/**
 * Crawls Detail Calender
 * Does …
Run Code Online (Sandbox Code Playgroud)

php goutte domcrawler

8
推荐指数
1
解决办法
1994
查看次数

如何使用symfony dom crawler将html表解析为数组

我有html表,我想从该表中创建数组

$html = '<table>
<tr>
    <td>satu</td>
    <td>dua</td>
</tr>
<tr>
    <td>tiga</td>
    <td>empat</td>
</tr>
</table>
Run Code Online (Sandbox Code Playgroud)

我的数组必须看起来像这样

array(
   array(
      "satu",
      "dua",
   ),
   array(
     "tiga",
     "empat",
   )
)
Run Code Online (Sandbox Code Playgroud)

我已经尝试了下面的代码,但无法获得我需要的数组

$crawler = new Crawler();
$crawler->addHTMLContent($html);
$row = array();
$tr_elements = $crawler->filterXPath('//table/tr');
foreach ($tr_elements as $tr) {
 // ???????
}
Run Code Online (Sandbox Code Playgroud)

php arrays symfony domcrawler

6
推荐指数
2
解决办法
6686
查看次数

如何使用Goutte Crawler提取数据?

这段代码,返回hrefs到内容,现在我想从这个hrefs中提取内容并将其发送到我的视图.我需要提取的名称div:

<div class="c_pad">
  <div class="c_label">
    <span class="std_header2">Contact:</span>
  </div>
<div class="c_name">
  <span class="std_text_b">Monkey</span>
</div>
<div class="clear"></div>
</div>
Run Code Online (Sandbox Code Playgroud)
<div class="c_pad">
    <div class="c_label">
      <span class="std_header2">Phone number:</span>
    </div>
    <div class="c_phone">
      <span class="std_text_b">001111111</span>
    </div>
    <div class="clear"></div>
</div>
Run Code Online (Sandbox Code Playgroud)
for($i=0; $i <= 1; $i++)
    {
      $p = new Client();
      $d = $p->request('GET', ''.$link.'&std=1&results='. $i);
      $n = $d->filter('a[class="o_title"]')->each(function ($node) 
        { 
         $pp = new Client();
         $dd = $pp->request('GET', $node->attr('href'));
         $kk = $dd->filter('div[id="adv_desc"]')->each(function ($tekst) {  echo $node->attr('href').'<br>'.$tekst->text(); 
                    });
         });
    }
Run Code Online (Sandbox Code Playgroud)

php goutte domcrawler

5
推荐指数
1
解决办法
4022
查看次数

如何从元素中获取文本,排除其中的一些其他元素

我domCrawler在symfony框架中使用.我使用它从html中抓取内容.现在我需要在带有ID的元素中获取文本.我可以使用下面的代码来完成文本:

$nodeValues = $crawler1->filter('#idOfTheElement')->each(function (Crawler $node, $i) {
            return $node->text();
        });
Run Code Online (Sandbox Code Playgroud)

element(#idOfTheElement)包含一些跨度,按钮等(也有一些类).我不想要那些内容.如何从元素中获取文本,排除其中的一些其他元素.

注意:我想要获取的文本除了元素#idOfTheElement之外没有任何其他包装器

Html如下所示:

<li id='#idOfTheElement'>Tel :<button data-pjtooltip="{dtanchor:'tooltipOpposeMkt'}" class="noMkt JS_PJ" type="button">text :</button><dl><dt><a name="tooltipOpposeMkt"></a></dt><dd><div class="wrapper"><p><strong>Signification des pictogrammes</strong></p><p>Devant un numéro, le picto <img width="11" height="9" alt="" src="something"> signale une opposition aux opérations de marketing direct.</p><span class="arrow">&nbsp;</span></div></dd></dl>12 23 45 88 99</li>
Run Code Online (Sandbox Code Playgroud)

symfony domcrawler

5
推荐指数
1
解决办法
846
查看次数

使用 DomCrawler 获取数据属性

是否可以使用 DomCrawler 获取数据?

$cralwer->attr('class')获取我节点的类属性,但->attr('data-something')或->attr('something')总是导致null.

编辑:标记 PHP 也是因为我在DomElement从 php操作对象时尝试过(使用->attributes->getNamedItem()),但它仍然无法正常工作。我想知道是否根本不可能返回数据属性?

php symfony domcrawler

5
推荐指数
1
解决办法
4155
查看次数

如何使用 Symfony 爬虫组件和 PHPUnit 测试表单提交的错误值?

当您通过浏览器使用该应用程序时,您发送了一个错误的值,系统会检查表单中的错误,如果出现问题(在本例中就是这样),它会重定向一条默认错误消息,该消息写在有罪的错误消息下方场地。

这是我试图用我的测试用例断言的行为,但我遇到了我没有预料到的 \InvalidArgumentException 。

我将 symfony/phpunit-bridge 与 phpunit/phpunit v8.5.23 和 symfony/dom-crawler v5.3.7 一起使用。这是它的示例:

public function testPayloadNotRespectingFieldLimits(): void
{
    $client = static::createClient();

    /** @var SomeRepository $repo */
    $repo = self::getContainer()->get(SomeRepository::class);
    $countEntries = $repo->count([]);
    
    $crawler = $client->request(
        'GET',
        '/route/to/form/add'
    );
    $this->assertResponseIsSuccessful(); // Goes ok.

    $form = $crawler->filter('[type=submit]')->form(); // It does retrieve my form node.
    
    // This is where it's not working.
    $form->setValues([
        'some[name]' => 'Someokvalue',
        'some[color]' => 'SomeNOTOKValue', // It is a ChoiceType with limited values, where 'SomeNOTOKValue' does not belong. This …
Run Code Online (Sandbox Code Playgroud)

php forms functional-testing symfony domcrawler

5
推荐指数
1
解决办法
913
查看次数

无法使用 Symfony Dom Crawler 获取元素的值

我正在使用 guzzle POST 方法获取 URL。它正在工作并返回我想要的页面。但问题是,当我想获取该页面中表单中的输入元素的值时,爬虫什么也不返回。我不知道为什么。

PHP:

<?php
use Symfony\Component\DomCrawler\Crawler;
use Guzzle\Http\Client;

$client = new Client();

$request = $client->get("https://example.com");
$response = $request->send();
$getRequest = $response->getBody();
$cookie = $response->getHeader("Set-Cookie");


$request = $client->post('https://example.com/page_example.php', array(
    'Content-Type' => 'application/x-www-form-urlencoded',
    'Cookie' => $cookie
    ), array(
        'param1' => 5,
        'param2' => 10,
        'param3' => 20
    ));

$response = $request->send();
$pageHTML = $response->getBody();

//fetch orderID
$crawler = new Crawler($pageHTML);
$orderID = $crawler->filter("input[name=orderId]")->attr('value');//there is only one element with this name

echo $orderID; //returns nothing
Run Code Online (Sandbox Code Playgroud)

我应该怎么办 ?

php dom symfony guzzle domcrawler

4
推荐指数
1
解决办法
2002
查看次数

在 symfony/goutte 中加入 URL

我有一个 Goutte/Client(goutte 使用 symfony 来处理请求),我想加入路径并获得最终 URL:

$client = new Goutte\Client();
$crawler = $client->request('GET', 'http://DOMAIN/some/path/')
// $crawler is instance of Symfony\Component\DomCrawler\Crawler

$new_path = '../new_page';
$final path = $crawler->someMagicFunction($new_path);
// final path == http://DOMAIN/some/new_page
Run Code Online (Sandbox Code Playgroud)

我正在寻找一种简单的方法,将$new_path变量与请求中的当前页面连接起来并获取新的 URL。

请注意,$new_page可以是以下任意一个:

new_page    ==> http://DOMAIN/some/path/new_page
../new_page ==> http://DOMAIN/some/new_page
/new_page   ==> http://DOMAIN/new_page
Run Code Online (Sandbox Code Playgroud)

symfony/goutte/guzzle 是否提供了任何简单的方法来做到这一点?

我找到了getUriForPathfrom Symfony\Component\HttpFoundation\Request,但我没有看到任何简单的方法将 the 转换Symfony\Component\BrowserKit\Request为HttpFoundation\Request

php symfony goutte guzzle domcrawler

2
推荐指数
1
解决办法
1389
查看次数

如何在 symfony 的 css 选择器组件中使用 :not 选择器

我想通过以下方式模拟我在 jQuery 中可以实现的目标 $('.someClass:not(.hidden)')

我试过下面的代码。

$crawler->filter('someClass:not(.hidden)')

但它似乎不起作用

php symfony domcrawler

2
推荐指数
1
解决办法
701
查看次数

标签 统计

domcrawler ×10

php ×8

symfony ×8

goutte ×3

guzzle ×2

arrays ×1

dom ×1

forms ×1

functional-testing ×1