如何查询某个URL是否被Google收录?

ala*_*ncc 1 google-apps-script

我想创建一个 Google 脚本来检查给定的 URL 是否被 Google 索引,因此我编写了以下函数:

\n
function CheckURLForGoogleIndex(url, activesheet) {// Delete the https:// and http:// prefix\n  var cururl = url.replace("https://", "");      \n  cururl = cururl.replace("http://", "");\n  var googlesearchurl = "https://www.google.com/search?q=site:" + encodeURIComponent(cururl);\n    var page = UrlFetchApp.fetch(googlesearchurl, {muteHttpExceptions: true}).getContentText();\n    // Wait for 1 second before starting another fetch\n    Utilities.sleep(1000);\n    var number = page.match("did not match any documents");\n    if (number) {\n      activesheet.getSheetByName("Not Google Index").appendRow([url]);\n    } else {\n      activesheet.getSheetByName("Google Index").appendRow([url]);\n    }  \n} \n
Run Code Online (Sandbox Code Playgroud)\n

但是,在调试代码时,调用UrlFetchApp.fetch后,我只能看到变量页面的标题。

\n

我尝试使用 Google 索引 URL 和非索引 URL 来测试该函数,但两者都会在 page.match 函数中返回 null,因此两者都放入“Google Index”表中。

\n

我的功能有什么问题吗?

\n

谢谢

\n

笔记:

\n

我在https://groups.google.com/g/google-apps-script-community/c/gs1qUuKwgn4上提出了这个问题,但没有人回答,所以我必须在这里问。

\n

输入和输出示例

\n

输入1:

\n

网址 = https://www.datanumen.com/

\n

activesheet = 包含工作表“Google Index”和“Not Google Index”的 GoogleSheet

\n

预期输出1:由于https://www.datanumen.com/已被 Google 索引,因此它将被添加到“Google Index”表中。

\n
page = "<!doctype html><html lang="en"><head><meta charset="UTF-8"><meta content="/images/branding/googleg/1x/googleg_standard_color_128dp.png" itemprop="image"><title>site:www.datanumen.com/ - Google Search\xe2\x80\xa6"\n
Run Code Online (Sandbox Code Playgroud)\n

输入2:

\n

网址 = https://www.datanumen.com/notindexedurl/

\n

activesheet = 包含工作表“Google Index”和“Not Google Index”的 GoogleSheet

\n

预期输出2:由于https://www.datanumen.com/notindexedurl/未被 Google 索引,因此它将被添加到“NOT Google Index”表中。

\n
page = "<!doctype html><html lang="en"><head><meta charset="UTF-8"><meta content="/images/branding/googleg/1x/googleg_standard_color_128dp.png" itemprop="image"><title>site:www.datanumen.com/notindexurl/ - G\xe2\x80\xa6"\n
Run Code Online (Sandbox Code Playgroud)\n

目前的问题是针对Input1和Input2,实际结果是:URL将始终被添加到“Google索引”表中,因为搜索结果将永远不会包含“不匹配任何文档”文本。

\n

更新

\n

我添加 console.log(page); 并再次调试。对于 Input1,我得到以下结果:

\n
<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN">\n<html>\n<head><meta http-equiv="content-type" content="text/html; charset=utf-8"><meta name="viewport" content="initial-scale=1"><title>https://www.google.com/search?q=site:www.datanumen.com%2F</title></head>\n<body style="font-family: arial, sans-serif; background-color: #fff; color: #000; padding:20px; font-size:18px;" onload="e=document.getElementById(\'captcha\');if(e){e.focus();}">\n<div style="max-width:400px;">\n<hr noshade size="1" style="color:#ccc; background-color:#ccc;"><br>\n<form id="captcha-form" action="index" method="post">\n<script src="https://www.google.com/recaptcha/api.js" async defer></script>\n<script>var submitCallback = function(response) {document.getElementById(\'captcha-form\').submit();};</script>\n<div id="recaptcha" class="g-recaptcha" data-sitekey="6LfwuyUTAAAAAOAmoS0fdqijC2PbbdH4kjq62Y1b" data-callback="submitCallback" data-s="c5Hy4maqTFv3SzYRiWhpsqYF2isZmauUQnLVljOiED_PiaVWJWCsHMzRAZyh8HLCBHJ_mjET7yODJu8AlZ33_xGAQ8TcKuXAd7rQpsYakaGKPD8USiGSFhiII2ai-Cf_B26i1Ufpko-qYQ8V3rezhiSXxi5J2yHZ-_WwEj8ukzy5znxzVurTM_2cY243Q4ofwP7E7eWBaHIg6N3ofmPuFXd-uRIUU4z0cU_pas8"></div>\n<input type=\'hidden\' name=\'q\' value=\'EgRrsuB5GKmx6oUGIhBKAdWty9nssg-nAtyy9n7hMgFy\'><input type="hidden" name="continue" value="https://www.google.com/search?q=site:www.datanumen.com%2F">\n</form>\n<hr noshade size="1" style="color:#ccc; background-color:#ccc;">\n\n<div style="font-size:13px;">\n<b>About this page</b><br><br>\n\nOur systems have detected unusual traffic from your computer network.  This page checks to see if it&#39;s really you sending the requests, and not a robot.  <a href="#" onclick="document.getElementById(\'infoDiv\').style.display=\'block\';">Why did this happen?</a><br><br>\n\n<div id="infoDiv" style="display:none; background-color:#eee; padding:10px; margin:0 0 15px 0; line-height:1.4em;">\nThis page appears when Google automatically detects requests coming from your computer network which appear to be in violation of the <a href="//www.google.com/policies/terms/">Terms of Service</a>. The block will expire shortly after those requests stop.  In the meantime, solving the above CAPTCHA will let you continue to use our services.<br><br>This traffic may have been sent by malicious software, a browser plug-in, or a script that sends automated requests.  If you share your network connection, ask your administrator for help &mdash; a different computer using the same IP address may be responsible.  <a href="//support.google.com/websearch/answer/86640">Learn more</a><br><br>Sometimes you may be asked to solve the CAPTCHA if you are using advanced terms that robots are known to use, or sending requests very quickly.\n</div>\n\nIP address: 107.178.224.121<br>Time: 2021-06-04T21:18:34Z<br>URL: https://www.google.com/search?q=site:www.datanumen.com%2F<br>\n</div>\n</div>\n</body>\n</html>\n
Run Code Online (Sandbox Code Playgroud)\n

Raf*_*rmo 5

回答:

不幸的是,通过尝试使用 UrlFetchApp 来网络抓取搜索结果来直接执行此操作是行不通的。不过,您可以使用第三方工具来获取搜索结果的数量。

更多信息:

我使用指数退避方法对此进行了测试,该方法有时能够429在 . 调用获取请求时解决过去的错误UrlFetchApp。

当使用UrlFetchApp网络抓取或连接到 API 时,服务器可能会因too many requests- 或 而拒绝请求HTTP Error 429。

Google Apps 脚本在云端运行,通过 Google 拥有的池中的一组 IP 地址运行。实际上,您可以在这里看到所有 IP 范围。大多数网站(尤其是谷歌等大公司)都有适当的架构来防止机器人抓取其网站并减慢流量。

有时可以使用指数退避和随机时间间隔的混合来克服此错误,如 Binance API 所示(全面披露:此 GitHub 存储库是我编写的。)

我认为要么 Google 直接阻止了 Apps 脚本 IP 池,要么有太多人尝试同样的事情 - 因为使用相同的技术,我无法获得任何不涉及输入验证码的响应,正如我们在上面的注释可以在字符串的日志中看到page。

可以做什么:

您可以使用许多第三方 API 来执行此操作,我建议您搜索一个能够满足您需求的第三方 API。

我测试了一个名为Authoritas 的工具,它返回不同关键字的搜索引擎索引。API 是异步的,因此可能需要长达一分钟才能获得响应,因此需要制定 Web 应用程序解决方案。

我使用的流程如下:

function makeApiCall(url, method, site) {
  const public_key = ""
  const private_key = ""
  const salt = "" 
  let timestamp = Date.now()

  const hash = Utilities.computeHmacSha256Signature(timestamp + public_key + salt, private_key)
  const headers = {
    "Authorization": "KeyAuth publicKey=" + public_key + " hash=" + toHexString(hash) + " ts=" + timestamp,
    "Content-Type": "application/json"
  }

  const requestParameters = {
    "search_engine": "google",
    "region": "us",
    "language": "en",
    "max_results": 100,
    "phrase": site,
    "search_type": "web",
    "user_agent": "pc",
    "parameters": {
      "priority": "standard"
    },
    "callback_type": "full",
    "callback": "script-web-app-exec-url"
  }

  const options = {
    "method": method,
    "headers": headers,
    "muteHttpExceptions": true,
    "payload": JSON.stringify(requestParameters)
  }

  const response = UrlFetchApp.fetch(url, options)
  return response
}

function toHexString(byteArray) {
  const hexString = Array.from(byteArray, function(byte) {
    return ('0' + (byte & 0xFF).toString(16)).slice(-2)
  }).join('')
  return hexString
}
Run Code Online (Sandbox Code Playgroud)

还有一个doPost(e)函数,以便当 API 返回数据时可以对其进行处理:

function doPost(e) {
  const jsonData = JSON.parse(e.postData.contents)
  const pages = jsonData.response.summary.pages
  const ss = SpreadsheetApp.openById("1QBzDdGn1yaUxFJciLH_Ru-BbLHuBIZTUk2UnrUShGw0") 
  
  if (Object.keys(pages).length == 0) {
    ss.getSheetByName("Not Google Index").appendRow([jsonData.request.phrase])
  }
  else {
    ss.getSheetByName("Google Index").appendRow([jsonData.request.phrase])
  }  
}
Run Code Online (Sandbox Code Playgroud)

然后,我从这里发布了具有以下设置的 Web 应用程序:

  • Execute as: me
  • Who has access: Anyone(不是 Anyone with a Google account)

请记住在提供时复制 Web 应用程序 URL 并将其粘贴到"callback": "script-web-app-exec-url"有效负载的部分(通常可以使用,ScriptApp.getService().getUrl()但根据此问题,当从脚本编辑器运行代码时,此方法返回链接/dev而不是/exec不会返回的链接)工作)。

然后可以像这样简单地运行:

function run() {
  const req = makeApiCall("v3.api.analyticsseo.com/serps/", "POST", "asdhfdhdfgdsfser.com")
  console.log(req.getContentText())
}
Run Code Online (Sandbox Code Playgroud)

请求将运行,来自 API 的响应将被记录,其中包含请求对象,然后当请求准备就绪时,Authoritas API 将调用您在参数中提供的脚本 URL,callback该脚本将运行该doPost()方法。

这是一个复杂的解决方法,但不幸的是,如今网络抓取变得越来越困难。

参考: