ala*_*ncc 1 google-apps-script
我想创建一个 Google 脚本来检查给定的 URL 是否被 Google 索引,因此我编写了以下函数:
\nfunction CheckURLForGoogleIndex(url, activesheet) {// Delete the https:// and http:// prefix\n var cururl = url.replace("https://", ""); \n cururl = cururl.replace("http://", "");\n var googlesearchurl = "https://www.google.com/search?q=site:" + encodeURIComponent(cururl);\n var page = UrlFetchApp.fetch(googlesearchurl, {muteHttpExceptions: true}).getContentText();\n // Wait for 1 second before starting another fetch\n Utilities.sleep(1000);\n var number = page.match("did not match any documents");\n if (number) {\n activesheet.getSheetByName("Not Google Index").appendRow([url]);\n } else {\n activesheet.getSheetByName("Google Index").appendRow([url]);\n } \n} \nRun Code Online (Sandbox Code Playgroud)\n但是,在调试代码时,调用UrlFetchApp.fetch后,我只能看到变量页面的标题。
\n我尝试使用 Google 索引 URL 和非索引 URL 来测试该函数,但两者都会在 page.match 函数中返回 null,因此两者都放入“Google Index”表中。
\n我的功能有什么问题吗?
\n谢谢
\n笔记:
\n我在https://groups.google.com/g/google-apps-script-community/c/gs1qUuKwgn4上提出了这个问题,但没有人回答,所以我必须在这里问。
\n输入和输出示例
\n输入1:
\n网址 = https://www.datanumen.com/
\nactivesheet = 包含工作表“Google Index”和“Not Google Index”的 GoogleSheet
\n预期输出1:由于https://www.datanumen.com/已被 Google 索引,因此它将被添加到“Google Index”表中。
\npage = "<!doctype html><html lang="en"><head><meta charset="UTF-8"><meta content="/images/branding/googleg/1x/googleg_standard_color_128dp.png" itemprop="image"><title>site:www.datanumen.com/ - Google Search\xe2\x80\xa6"\nRun Code Online (Sandbox Code Playgroud)\n输入2:
\n网址 = https://www.datanumen.com/notindexedurl/
\nactivesheet = 包含工作表“Google Index”和“Not Google Index”的 GoogleSheet
\n预期输出2:由于https://www.datanumen.com/notindexedurl/未被 Google 索引,因此它将被添加到“NOT Google Index”表中。
\npage = "<!doctype html><html lang="en"><head><meta charset="UTF-8"><meta content="/images/branding/googleg/1x/googleg_standard_color_128dp.png" itemprop="image"><title>site:www.datanumen.com/notindexurl/ - G\xe2\x80\xa6"\nRun Code Online (Sandbox Code Playgroud)\n目前的问题是针对Input1和Input2,实际结果是:URL将始终被添加到“Google索引”表中,因为搜索结果将永远不会包含“不匹配任何文档”文本。
\n更新
\n我添加 console.log(page); 并再次调试。对于 Input1,我得到以下结果:
\n<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN">\n<html>\n<head><meta http-equiv="content-type" content="text/html; charset=utf-8"><meta name="viewport" content="initial-scale=1"><title>https://www.google.com/search?q=site:www.datanumen.com%2F</title></head>\n<body style="font-family: arial, sans-serif; background-color: #fff; color: #000; padding:20px; font-size:18px;" onload="e=document.getElementById(\'captcha\');if(e){e.focus();}">\n<div style="max-width:400px;">\n<hr noshade size="1" style="color:#ccc; background-color:#ccc;"><br>\n<form id="captcha-form" action="index" method="post">\n<script src="https://www.google.com/recaptcha/api.js" async defer></script>\n<script>var submitCallback = function(response) {document.getElementById(\'captcha-form\').submit();};</script>\n<div id="recaptcha" class="g-recaptcha" data-sitekey="6LfwuyUTAAAAAOAmoS0fdqijC2PbbdH4kjq62Y1b" data-callback="submitCallback" data-s="c5Hy4maqTFv3SzYRiWhpsqYF2isZmauUQnLVljOiED_PiaVWJWCsHMzRAZyh8HLCBHJ_mjET7yODJu8AlZ33_xGAQ8TcKuXAd7rQpsYakaGKPD8USiGSFhiII2ai-Cf_B26i1Ufpko-qYQ8V3rezhiSXxi5J2yHZ-_WwEj8ukzy5znxzVurTM_2cY243Q4ofwP7E7eWBaHIg6N3ofmPuFXd-uRIUU4z0cU_pas8"></div>\n<input type=\'hidden\' name=\'q\' value=\'EgRrsuB5GKmx6oUGIhBKAdWty9nssg-nAtyy9n7hMgFy\'><input type="hidden" name="continue" value="https://www.google.com/search?q=site:www.datanumen.com%2F">\n</form>\n<hr noshade size="1" style="color:#ccc; background-color:#ccc;">\n\n<div style="font-size:13px;">\n<b>About this page</b><br><br>\n\nOur systems have detected unusual traffic from your computer network. This page checks to see if it's really you sending the requests, and not a robot. <a href="#" onclick="document.getElementById(\'infoDiv\').style.display=\'block\';">Why did this happen?</a><br><br>\n\n<div id="infoDiv" style="display:none; background-color:#eee; padding:10px; margin:0 0 15px 0; line-height:1.4em;">\nThis page appears when Google automatically detects requests coming from your computer network which appear to be in violation of the <a href="//www.google.com/policies/terms/">Terms of Service</a>. The block will expire shortly after those requests stop. In the meantime, solving the above CAPTCHA will let you continue to use our services.<br><br>This traffic may have been sent by malicious software, a browser plug-in, or a script that sends automated requests. If you share your network connection, ask your administrator for help — a different computer using the same IP address may be responsible. <a href="//support.google.com/websearch/answer/86640">Learn more</a><br><br>Sometimes you may be asked to solve the CAPTCHA if you are using advanced terms that robots are known to use, or sending requests very quickly.\n</div>\n\nIP address: 107.178.224.121<br>Time: 2021-06-04T21:18:34Z<br>URL: https://www.google.com/search?q=site:www.datanumen.com%2F<br>\n</div>\n</div>\n</body>\n</html>\nRun Code Online (Sandbox Code Playgroud)\n
不幸的是,通过尝试使用 UrlFetchApp 来网络抓取搜索结果来直接执行此操作是行不通的。不过,您可以使用第三方工具来获取搜索结果的数量。
我使用指数退避方法对此进行了测试,该方法有时能够429在 . 调用获取请求时解决过去的错误UrlFetchApp。
当使用UrlFetchApp网络抓取或连接到 API 时,服务器可能会因too many requests- 或 而拒绝请求HTTP Error 429。
Google Apps 脚本在云端运行,通过 Google 拥有的池中的一组 IP 地址运行。实际上,您可以在这里看到所有 IP 范围。大多数网站(尤其是谷歌等大公司)都有适当的架构来防止机器人抓取其网站并减慢流量。
有时可以使用指数退避和随机时间间隔的混合来克服此错误,如 Binance API 所示(全面披露:此 GitHub 存储库是我编写的。)
我认为要么 Google 直接阻止了 Apps 脚本 IP 池,要么有太多人尝试同样的事情 - 因为使用相同的技术,我无法获得任何不涉及输入验证码的响应,正如我们在上面的注释可以在字符串的日志中看到page。
您可以使用许多第三方 API 来执行此操作,我建议您搜索一个能够满足您需求的第三方 API。
我测试了一个名为Authoritas 的工具,它返回不同关键字的搜索引擎索引。API 是异步的,因此可能需要长达一分钟才能获得响应,因此需要制定 Web 应用程序解决方案。
我使用的流程如下:
function makeApiCall(url, method, site) {
const public_key = ""
const private_key = ""
const salt = ""
let timestamp = Date.now()
const hash = Utilities.computeHmacSha256Signature(timestamp + public_key + salt, private_key)
const headers = {
"Authorization": "KeyAuth publicKey=" + public_key + " hash=" + toHexString(hash) + " ts=" + timestamp,
"Content-Type": "application/json"
}
const requestParameters = {
"search_engine": "google",
"region": "us",
"language": "en",
"max_results": 100,
"phrase": site,
"search_type": "web",
"user_agent": "pc",
"parameters": {
"priority": "standard"
},
"callback_type": "full",
"callback": "script-web-app-exec-url"
}
const options = {
"method": method,
"headers": headers,
"muteHttpExceptions": true,
"payload": JSON.stringify(requestParameters)
}
const response = UrlFetchApp.fetch(url, options)
return response
}
function toHexString(byteArray) {
const hexString = Array.from(byteArray, function(byte) {
return ('0' + (byte & 0xFF).toString(16)).slice(-2)
}).join('')
return hexString
}
Run Code Online (Sandbox Code Playgroud)
还有一个doPost(e)函数,以便当 API 返回数据时可以对其进行处理:
function doPost(e) {
const jsonData = JSON.parse(e.postData.contents)
const pages = jsonData.response.summary.pages
const ss = SpreadsheetApp.openById("1QBzDdGn1yaUxFJciLH_Ru-BbLHuBIZTUk2UnrUShGw0")
if (Object.keys(pages).length == 0) {
ss.getSheetByName("Not Google Index").appendRow([jsonData.request.phrase])
}
else {
ss.getSheetByName("Google Index").appendRow([jsonData.request.phrase])
}
}
Run Code Online (Sandbox Code Playgroud)
然后,我从这里发布了具有以下设置的 Web 应用程序:
Execute as: meWho has access: Anyone(不是 Anyone with a Google account)请记住在提供时复制 Web 应用程序 URL 并将其粘贴到"callback": "script-web-app-exec-url"有效负载的部分(通常可以使用,ScriptApp.getService().getUrl()但根据此问题,当从脚本编辑器运行代码时,此方法返回链接/dev而不是/exec不会返回的链接)工作)。
然后可以像这样简单地运行:
function run() {
const req = makeApiCall("v3.api.analyticsseo.com/serps/", "POST", "asdhfdhdfgdsfser.com")
console.log(req.getContentText())
}
Run Code Online (Sandbox Code Playgroud)
请求将运行,来自 API 的响应将被记录,其中包含请求对象,然后当请求准备就绪时,Authoritas API 将调用您在参数中提供的脚本 URL,callback该脚本将运行该doPost()方法。
这是一个复杂的解决方法,但不幸的是,如今网络抓取变得越来越困难。
| 归档时间: |
|
| 查看次数: |
1421 次 |
| 最近记录: |