use*_*981 6 screen-scraping r web-crawler web
我正在尝试编写将转到每个页面并从那里获取信息的代码.网址< - http://www.wikiart.org/en/claude-monet/mode/all-paintings-by-alphabet
我有代码输出所有hrefs.但它不起作用.
library(XML)
library(RCurl)
library(stringr)
tagrecode <- readHTMLTable ("http://www.wikiart.org/en/claude-monet/mode/all- paintings-by-alphabet")
tabla <- as.data.frame(tagrecode)
str(tabla)
names (tabla) <- c("name", "desc", "cat", "updated")
str(tabla)
res <- htmlParse ("http://www.wikiart.org/en/claude-monet/mode/all-paintings-by- alphabet")
enlaces <- getNodeSet (res, "//p[@class='pb5']/a/@href")
enlaces <- unlist(lapply(enlaces, as.character))
tabla$enlace <- paste("http://www.wikiart.org/en/claude-monet/mode/all-paintings-by- alphabet")
str(tabla)
lisurl <- tabla$enlace
fu1 <- function(url){
print(url)
pas1 <- htmlParse(url, useInternalNodes=T)
pas2 <- xpathSApply(pas1, "//p[@class='pb5']/a/@href")
}
urldef <- lapply(lisurl,fu1)
Run Code Online (Sandbox Code Playgroud)
在我有这个页面上所有图片的网址列表后,我想去第二个-...- 23页收集所有图片的网址.
下一步 - 删除有关每张图片的信息.我有一个工作代码,我需要在一个通用代码中构建它.
library(XML)
url = "http://www.wikiart.org/en/claude-monet/camille-and-jean-monet-in-the-garden-at-argenteuil"
doc = htmlTreeParse(url, useInternalNodes=T)
pictureName <- xpathSApply(doc,"//h1[@itemprop='name']", xmlValue)
date <- xpathSApply(doc, "//span[@itemprop='dateCreated']", xmlValue)
author <- xpathSApply(doc, "//a[@itemprop='author']", xmlValue)
style <- xpathSApply(doc, "//span[@itemprop='style']", xmlValue)
genre <- xpathSApply(doc, "//span[@itemprop='genre']", xmlValue)
pictureName
date
author
style
genre
Run Code Online (Sandbox Code Playgroud)
每个建议如何做到这一点将不胜感激!
这似乎有效.
library(XML)
library(httr)
url <- "http://www.wikiart.org/en/claude-monet/mode/all-paintings-by-alphabet/"
hrefs <- list()
for (i in 1:23) {
response <- GET(paste0(url,i))
doc <- content(response,type="text/html")
hrefs <- c(hrefs,doc["//p[@class='pb5']/a/@href"])
}
url <- "http://www.wikiart.org"
xPath <- c(pictureName = "//h1[@itemprop='name']",
date = "//span[@itemprop='dateCreated']",
author = "//a[@itemprop='author']",
style = "//span[@itemprop='style']",
genre = "//span[@itemprop='genre']")
get.picture <- function(href) {
response <- GET(paste0(url,href))
doc <- content(response,type="text/html")
info <- sapply(xPath,function(xp)ifelse(length(doc[xp])==0,NA,xmlValue(doc[xp][[1]])))
}
pictures <- do.call(rbind,lapply(hrefs,get.picture))
head(pictures)
# pictureName date author style genre
# [1,] "A Corner of the Garden at Montgeron" "1877" "Claude Monet" "Impressionism" "landscape"
# [2,] "A Corner of the Studio" "1861" "Claude Monet" "Realism" "self-portrait"
# [3,] "A Farmyard in Normandy" "c.1863" "Claude Monet" "Realism" "landscape"
# [4,] "A Windmill near Zaandam" NA "Claude Monet" "Impressionism" "landscape"
# [5,] "A Woman Reading" "1872" "Claude Monet" "Impressionism" "genre painting"
# [6,] "Adolphe Monet Reading in the Garden" "1866" "Claude Monet" "Impressionism" "genre painting"
Run Code Online (Sandbox Code Playgroud)
你实际上非常接近.你的xPath很好; 一个问题是,并非所有的图片都的信息(例如,对于一些网页你要访问的节点集是空的) - 注意,"A Windnill NEAD赞丹"的日期.所以代码必须处理这种可能性.
因此,在此示例中,第一个循环为每个页面(1:23)抓取锚标记的href属性的值,并将它们组合成长度为~1300的向量.
要处理这1300个页面中的每一个,并且由于我们必须处理缺失的标记,因此创建包含xPath字符串的向量并将该元素应用于每个页面更为直接.这就是功能的get.picture(...)作用.最后一个语句使用1300个href中的每一个调用此函数,并使用将行结果排在一起do.call(rbind,...).
另请注意,此代码对类HTMLInternalDocument的对象使用更紧凑的索引功能:doc[xpath]其中xpath是xPath字符串.这避免了使用xpathSApply(...),尽管后者会起作用.