Hak*_*kim 0 html perl selenium
我一直在使用HTML :: SimpleLinkExtor从这个页面中提取链接:http://cpc.cs.qub.ac.uk/authorIndex/AUTHOR_index.html 虽然它适用于所有人,但当一个链接有'Ç'作为一个角色.它的作用是将它改为%C7.因此,当我在程序的其余部分中使用链接时,我收到代码404错误.这是我的代码:
#!/usr/bin/perl
use strict;
use warnings;
use HTML::SimpleLinkExtor;
use Time::HiRes qw(sleep);
use Test::WWW::Selenium;
use Test::More "no_plan"; #tests => 37; #
#use Test::Exception;
Test::More->builder->output ('result.txt');
Test::More->builder->failure_output ('errors.txt');
my $base = "http://cpc.cs.qub.ac.uk/authorIndex/AUTHOR_index.html";
my $sel = Test::WWW::Selenium->new( host => "localhost",
port => 4444,
browser => "*firefox",
browser_url => "http://cpc.cs.qub.ac.uk/" );
################################################
my $extor = HTML::SimpleLinkExtor->new($base);
$extor->parse_url($base);
my @all_links = $extor->a;
################################################
$sel->start();
$sel->open_ok($base);
$sel->open_ok($_) foreach (@all_links);
$sel->stop();
Run Code Online (Sandbox Code Playgroud)
同样,有什么想法我如何使用提取的链接实现click()函数.
谢谢
该网页以latin1编码提供,因此它编码为字节0xC7.尽管如此,HTML :: SimpleLinkExtor 应该足够聪明,可以将其转换为UTF-8作为链接,因为这几乎是标准的.但它没有这样做.在其来源中它说:
sub parse_url {
my $data = $_[0]->ua->get( $_[1] )->content;
return unless $data;
$_[0]->parse( $data );
}
Run Code Online (Sandbox Code Playgroud)
这里的错误是它应该使用- > decoding_content而不是- > content来正确进行编码转换.您可能想要为HTML :: SimpleLinkExtor提交错误报告.与此同时,你可以尝试编写一个自己的方法来替换这个破碎的方法.
编辑:这可能工作(未经测试):
# replace this:
$extor->parse_url($base);
# with this:
my $data = $extor->ua->get($base)->decoded_content;
if (defined $data) {
$extor->parse($data);
}
Run Code Online (Sandbox Code Playgroud)