使用Dom Crawler只获取文本(没有标记).
$html = EOT<<<
<div class="coucu">
Get Description <span>Coucu</span>
</div>
EOT;
$crawler = new Crawler($html);
$crawler = $crawler->filter('.coucu')->first()->text();
Run Code Online (Sandbox Code Playgroud)
输出:获取描述Coucu
我想输出(仅):获取描述
更新:
我找到了一个解决方案:(但这是非常糟糕的解决方案)
...
$html = $crawler->filter('.coucu')->html();
// use strip_tags_content in https://php.net/strip_tags
$html = strip_tags_content($html,'span');
Run Code Online (Sandbox Code Playgroud) I am using PHP 7.1.33 and "fabpot/goutte": "^3.2". My composer file looks like the following:
{
"name": "ubuntu/workspace",
"require": {
"fabpot/goutte": "^3.2"
},
"authors": [
{
"name": "admin",
"email": "admin@admin.com"
}
]
}
Run Code Online (Sandbox Code Playgroud)
I am trying to get details by a time range from a webpage but struggle how to pass the $crawler-values to my final result array $res1Array.
I tried the following:
<?php
require 'vendor/autoload.php';
use Goutte\Client;
use Symfony\Component\DomCrawler\Crawler;
/**
* Crawls Detail Calender
* Does …Run Code Online (Sandbox Code Playgroud) 我有html表,我想从该表中创建数组
$html = '<table>
<tr>
<td>satu</td>
<td>dua</td>
</tr>
<tr>
<td>tiga</td>
<td>empat</td>
</tr>
</table>
Run Code Online (Sandbox Code Playgroud)
我的数组必须看起来像这样
array(
array(
"satu",
"dua",
),
array(
"tiga",
"empat",
)
)
Run Code Online (Sandbox Code Playgroud)
我已经尝试了下面的代码,但无法获得我需要的数组
$crawler = new Crawler();
$crawler->addHTMLContent($html);
$row = array();
$tr_elements = $crawler->filterXPath('//table/tr');
foreach ($tr_elements as $tr) {
// ???????
}
Run Code Online (Sandbox Code Playgroud) 这段代码,返回hrefs到内容,现在我想从这个hrefs中提取内容并将其发送到我的视图.我需要提取的名称div:
<div class="c_pad">
<div class="c_label">
<span class="std_header2">Contact:</span>
</div>
<div class="c_name">
<span class="std_text_b">Monkey</span>
</div>
<div class="clear"></div>
</div>
Run Code Online (Sandbox Code Playgroud)
<div class="c_pad">
<div class="c_label">
<span class="std_header2">Phone number:</span>
</div>
<div class="c_phone">
<span class="std_text_b">001111111</span>
</div>
<div class="clear"></div>
</div>
Run Code Online (Sandbox Code Playgroud)
for($i=0; $i <= 1; $i++)
{
$p = new Client();
$d = $p->request('GET', ''.$link.'&std=1&results='. $i);
$n = $d->filter('a[class="o_title"]')->each(function ($node)
{
$pp = new Client();
$dd = $pp->request('GET', $node->attr('href'));
$kk = $dd->filter('div[id="adv_desc"]')->each(function ($tekst) { echo $node->attr('href').'<br>'.$tekst->text();
});
});
}
Run Code Online (Sandbox Code Playgroud) 我domCrawler在symfony框架中使用.我使用它从html中抓取内容.现在我需要在带有ID的元素中获取文本.我可以使用下面的代码来完成文本:
$nodeValues = $crawler1->filter('#idOfTheElement')->each(function (Crawler $node, $i) {
return $node->text();
});
Run Code Online (Sandbox Code Playgroud)
element(#idOfTheElement)包含一些跨度,按钮等(也有一些类).我不想要那些内容.如何从元素中获取文本,排除其中的一些其他元素.
注意:我想要获取的文本除了元素#idOfTheElement之外没有任何其他包装器
Html如下所示:
<li id='#idOfTheElement'>Tel :<button data-pjtooltip="{dtanchor:'tooltipOpposeMkt'}" class="noMkt JS_PJ" type="button">text :</button><dl><dt><a name="tooltipOpposeMkt"></a></dt><dd><div class="wrapper"><p><strong>Signification des pictogrammes</strong></p><p>Devant un numéro, le picto <img width="11" height="9" alt="" src="something"> signale une opposition aux opérations de marketing direct.</p><span class="arrow"> </span></div></dd></dl>12 23 45 88 99</li>
Run Code Online (Sandbox Code Playgroud) 是否可以使用 DomCrawler 获取数据?
$cralwer->attr('class')获取我节点的类属性,但->attr('data-something')或->attr('something')总是导致null.
编辑:标记 PHP 也是因为我在DomElement从 php操作对象时尝试过(使用->attributes->getNamedItem()),但它仍然无法正常工作。我想知道是否根本不可能返回数据属性?
当您通过浏览器使用该应用程序时,您发送了一个错误的值,系统会检查表单中的错误,如果出现问题(在本例中就是这样),它会重定向一条默认错误消息,该消息写在有罪的错误消息下方场地。
这是我试图用我的测试用例断言的行为,但我遇到了我没有预料到的 \InvalidArgumentException 。
我将 symfony/phpunit-bridge 与 phpunit/phpunit v8.5.23 和 symfony/dom-crawler v5.3.7 一起使用。这是它的示例:
public function testPayloadNotRespectingFieldLimits(): void
{
$client = static::createClient();
/** @var SomeRepository $repo */
$repo = self::getContainer()->get(SomeRepository::class);
$countEntries = $repo->count([]);
$crawler = $client->request(
'GET',
'/route/to/form/add'
);
$this->assertResponseIsSuccessful(); // Goes ok.
$form = $crawler->filter('[type=submit]')->form(); // It does retrieve my form node.
// This is where it's not working.
$form->setValues([
'some[name]' => 'Someokvalue',
'some[color]' => 'SomeNOTOKValue', // It is a ChoiceType with limited values, where 'SomeNOTOKValue' does not belong. This …Run Code Online (Sandbox Code Playgroud) 我正在使用 guzzle POST 方法获取 URL。它正在工作并返回我想要的页面。但问题是,当我想获取该页面中表单中的输入元素的值时,爬虫什么也不返回。我不知道为什么。
PHP:
<?php
use Symfony\Component\DomCrawler\Crawler;
use Guzzle\Http\Client;
$client = new Client();
$request = $client->get("https://example.com");
$response = $request->send();
$getRequest = $response->getBody();
$cookie = $response->getHeader("Set-Cookie");
$request = $client->post('https://example.com/page_example.php', array(
'Content-Type' => 'application/x-www-form-urlencoded',
'Cookie' => $cookie
), array(
'param1' => 5,
'param2' => 10,
'param3' => 20
));
$response = $request->send();
$pageHTML = $response->getBody();
//fetch orderID
$crawler = new Crawler($pageHTML);
$orderID = $crawler->filter("input[name=orderId]")->attr('value');//there is only one element with this name
echo $orderID; //returns nothing
Run Code Online (Sandbox Code Playgroud)
我应该怎么办 ?
我有一个 Goutte/Client(goutte 使用 symfony 来处理请求),我想加入路径并获得最终 URL:
$client = new Goutte\Client();
$crawler = $client->request('GET', 'http://DOMAIN/some/path/')
// $crawler is instance of Symfony\Component\DomCrawler\Crawler
$new_path = '../new_page';
$final path = $crawler->someMagicFunction($new_path);
// final path == http://DOMAIN/some/new_page
Run Code Online (Sandbox Code Playgroud)
我正在寻找一种简单的方法,将$new_path变量与请求中的当前页面连接起来并获取新的 URL。
请注意,$new_page可以是以下任意一个:
new_page ==> http://DOMAIN/some/path/new_page
../new_page ==> http://DOMAIN/some/new_page
/new_page ==> http://DOMAIN/new_page
Run Code Online (Sandbox Code Playgroud)
symfony/goutte/guzzle 是否提供了任何简单的方法来做到这一点?
我找到了getUriForPathfrom Symfony\Component\HttpFoundation\Request,但我没有看到任何简单的方法将 the 转换Symfony\Component\BrowserKit\Request为HttpFoundation\Request
我想通过以下方式模拟我在 jQuery 中可以实现的目标
$('.someClass:not(.hidden)')
我试过下面的代码。
$crawler->filter('someClass:not(.hidden)')
但它似乎不起作用