Leg*_*end 5 language-agnostic evaluation nlp information-retrieval
我编写了一个系统,该系统总结了包含数千个单词的长文档。是否有关于在用户调查的背景下如何评估这种系统的规范?
简而言之,是否存在评估我的工具拯救人类时间的度量标准?目前,我正在考虑使用(读取原始文档所花费的时间/阅读摘要所花费的时间)作为确定节省时间的一种方式,但是是否有更好的指标?
目前,我正在询问用户有关摘要准确性的主观问题。
蓝线
\n胭脂
\n胭脂措施召回
\n面向回忆的基础评估\nW(人类参考摘要) In w(机器生成摘要)
\n这就是机器生成摘要中出现的单词(和/或 n-gram)的数量。
\n系统和参考文献摘要之间 N 元语法的重叠。\n-Rouge N,此处 N 是 n 元语法
\nreference_text = """Artificial intelligence (AI, also machine intelligence, MI) is intelligence demonstrated by machines, in contrast to the natural intelligence (NI) displayed by humans and other animals. In computer science AI research is defined as the study of "intelligent agents": any device that perceives its environment and takes actions that maximize its chance of successfully achieving its goals. Colloquially, the term "artificial intelligence" is applied when a machine mimics "cognitive" functions that humans associate with other human minds, such as "learning" and "problem solving". See glossary of artificial intelligence. The scope of AI is disputed: as machines become increasingly capable, tasks considered as requiring "intelligence" are often removed from the definition, a phenomenon known as the AI effect, leading to the quip "AI is whatever hasn\'t been done yet." For instance, optical character recognition is frequently excluded from "artificial intelligence", having become a routine technology. Capabilities generally classified as AI as of 2017 include successfully understanding human speech, competing at a high level in strategic game systems (such as chess and Go), autonomous cars, intelligent routing in content delivery networks, military simulations, and interpreting complex data, including images and videos. Artificial intelligence was founded as an academic discipline in 1956, and in the years since has experienced several waves of optimism, followed by disappointment and the loss of funding (known as an "AI winter"), followed by new approaches, success and renewed funding. For most of its history, AI research has been divided into subfields that often fail to communicate with each other. These sub-fields are based on technical considerations, such as particular goals (e.g. "robotics" or "machine learning"), the use of particular tools ("logic" or "neural networks"), or deep philosophical differences. Subfields have also been based on social factors (particular institutions or the work of particular researchers). The traditional problems (or goals) of AI research include reasoning, knowledge, planning, learning, natural language processing, perception and the ability to move and manipulate objects. General intelligence is among the field\'s long-term goals. Approaches include statistical methods, computational intelligence, and traditional symbolic AI. Many tools are used in AI, including versions of search and mathematical optimization, neural networks and methods based on statistics, probability and economics. The AI field draws upon computer science, mathematics, psychology, linguistics, philosophy and many others. The field was founded on the claim that human intelligence "can be so precisely described that a machine can be made to simulate it". This raises philosophical arguments about the nature of the mind and the ethics of creating artificial beings endowed with human-like intelligence, issues which have been explored by myth, fiction and philosophy since antiquity. Some people also consider AI to be a danger to humanity if it progresses unabatedly. Others believe that AI, unlike previous technological revolutions, will create a risk of mass unemployment. In the twenty-first century, AI techniques have experienced a resurgence following concurrent advances in computer power, large amounts of data, and theoretical understanding; and AI techniques have become an essential part of the technology industry, helping to solve many challenging problems in computer science."""\nRun Code Online (Sandbox Code Playgroud)\n抽象概括
\n # Abstractive Summarize \n len(reference_text.split())\n from transformers import pipeline\n summarization = pipeline("summarization")\n abstractve_summarization = summarization(reference_text)[0]["summary_text"]\nRun Code Online (Sandbox Code Playgroud)\n抽象输出
\n In computer science AI research is defined as the study of "intelligent agents" Colloquially, the term "artificial intelligence" is applied when a machine mimics "cognitive" functions that humans associate with other human minds, such as "learning" and "problem solving" Capabilities generally classified as AI as of 2017 include successfully understanding human speech, competing at a high level in strategic game systems (such as chess and Go)\nRun Code Online (Sandbox Code Playgroud)\n精炼总结
\n # Extractive summarize\n from sumy.parsers.plaintext import PlaintextParser\n from sumy.nlp.tokenizers import Tokenizer\n from sumy.summarizers.lex_rank import LexRankSummarizer\n parser = PlaintextParser.from_string(reference_text, Tokenizer("english"))\n # parser.document.sentences\n summarizer = LexRankSummarizer()\n extractve_summarization = summarizer(parser.document,2)\n extractve_summarization) = \' \'.join([str(s) for s in list(extractve_summarization)])\nRun Code Online (Sandbox Code Playgroud)\n提取输出
\nColloquially, the term "artificial intelligence" is often used to describe machines that mimic "cognitive" functions that humans associate with the human mind, such as "learning" and "problem solving".As machines become increasingly capable, tasks considered to require "intelligence" are often removed from the definition of AI, a phenomenon known as the AI effect. Sub-fields have also been based on social factors (particular institutions or the work of particular researchers).The traditional problems (or goals) of AI research include reasoning, knowledge representation, planning, learning, natural language processing, perception and the ability to move and manipulate objects.\nRun Code Online (Sandbox Code Playgroud)\n使用 Rouge 评估抽象摘要
\n from rouge import Rouge\n r = Rouge()\n r.get_scores(abstractve_summarization, reference_text)\nRun Code Online (Sandbox Code Playgroud)\n使用 Rouge Abstractive 摘要输出
\n [{\'rouge-1\': {\'f\': 0.22299651364421083,\n \'p\': 0.9696969696969697,\n \'r\': 0.12598425196850394},\n \'rouge-2\': {\'f\': 0.21328671127225052,\n \'p\': 0.9384615384615385,\n \'r\': 0.1203155818540434},\n \'rouge-l\': {\'f\': 0.29041095634452996,\n \'p\': 0.9636363636363636,\n \'r\': 0.17096774193548386}}]\nRun Code Online (Sandbox Code Playgroud)\n使用 Rouge 评估抽象摘要
\n from rouge import Rouge\n r = Rouge()\n r.get_scores(extractve_summarization, reference_text)\nRun Code Online (Sandbox Code Playgroud)\n使用 Rouge Extractive 摘要输出
\n [{\'rouge-1\': {\'f\': 0.27860696251962963,\n \'p\': 0.8842105263157894,\n \'r\': 0.16535433070866143},\n \'rouge-2\': {\'f\': 0.22296172781038814,\n \'p\': 0.7127659574468085,\n \'r\': 0.13214990138067062},\n \'rouge-l\': {\'f\': 0.354755780824869,\n \'p\': 0.8734177215189873,\n \'r\': 0.22258064516129034}}]\nRun Code Online (Sandbox Code Playgroud)\n解释胭脂分数
\nROUGE 是一堆重叠的单词。ROUGE-N 指的是重叠的 n 元语法。具体来说:
\n\n与原始论文相比,我尝试简化符号。假设我们正在计算 ROUGE-2,又名二元匹配。分子 \xe2\x88\x91s 循环遍历单个参考摘要中的所有二元组,并计算在候选摘要中找到匹配二元组的次数(由摘要算法提出)。如果有多个参考摘要,\xe2\x88\x91r 确保我们对所有参考摘要重复该过程。
\n分母只是计算所有参考摘要中二元组的总数。这是一对文档摘要的过程。您对所有文档重复该过程,并对所有分数进行平均,这将为您提供 ROUGE-N 分数。因此,较高的分数意味着平均而言,您的摘要和参考文献之间的 n 元语法有很高的重叠度。
\n Example:\n\n S1. police killed the gunman\n \n S2. police kill the gunman\n \n S3. the gunman kill police\nRun Code Online (Sandbox Code Playgroud)\nS1 是参考,S2 和 S3 是候选。请注意,S2 和 S3 都有一个与参考重叠的二元组,因此它们具有相同的 ROUGE-2 分数,尽管 S2 应该更好。额外的 ROUGE-L 分数处理此问题,其中 L 代表最长公共子序列。在S2中,第一个单词和最后两个单词与参考匹配,因此得分为3/4,而S3仅与二元组匹配,因此得分为2/4。
\n从历史上看,通常通过与人工生成的参考摘要进行比较来评估摘要系统。在某些情况下,人类摘要者通过从原始文档中选择相关句子来构建摘要;在其他情况下,摘要是从头开始手写的。
这两种技术类似于自动摘要系统的两大类 - 提取与抽象(更多详细信息可在Wikipedia 上获得)。
一个标准工具是Rouge,它是一个脚本(或一组脚本;我记不清了),它计算自动摘要和参考摘要之间的 n-gram 重叠。Rough 可以选择性地计算允许在两个摘要之间插入或删除单词的重叠(例如,如果允许跳过 2 个单词,则“已安装的泵”将被视为与“已安装的有缺陷的防洪泵”匹配)。
我的理解是 Rouge 的 n-gram 重叠分数在一定程度上与人类对摘要的评估有很好的相关性,但随着摘要质量的提高,这种关系可能会破裂。即,超出某个质量阈值,被人类评估者判断为更好的摘要的评分可能与被判断为较差的摘要相似——或被超过。尽管如此,Rouge 分数在比较 2 个候选摘要系统时可能是有用的第一次削减,或者是一种在将系统传递给人类评估人员之前自动进行回归测试并剔除严重回归的方法。
如果您能够负担得起时间/金钱成本,那么您收集人工判断的方法可能是最好的评估方法。为了让该过程更加严谨,您可以查看最近总结任务中使用的评分标准(请参阅@John Lehmann 提到的各种会议)。这些评估者使用的评分表可能有助于指导您自己的评估。
一般来说:
Bleu衡量精度:机器生成的摘要中的单词(和/或n-gram)在人工参考摘要中出现了多少。
胭脂度量回想:人类参考摘要中的单词(和/或n-gram)在机器生成的摘要中出现了多少。
当然,这些结果是相辅相成的,就像精度与查全率一样。如果您从参考文献中看到的系统结果中有很多单词/词组,您的Bleu值就会很高;如果您在参考文献中出现的人类参考中有很多单词/词组的话,您的胭脂就会很高。
有一个称为“ 简短惩罚”的东西,这很重要,已经被添加到标准的Bleu实现中。它会惩罚比参考值的一般长度短的系统结果(在此处了解更多信息)。这补充了n-gram度量行为,实际上会惩罚比参考结果更长的时间,因为分母增长得越长,系统结果就越长。
您也可以对Rouge实施类似的操作,但是这次惩罚系统的结果要比一般参考长度更长,否则将使他们能够人为地获得更高的Rouge分数(因为结果越长,您击中某些球的机会就越高。在参考文献中出现的单词)。在Rouge中,我们将除以参考人的长度,因此,对于更长的系统结果(可能人为地提高其Rouge得分),我们将需要额外的罚款。
最后,您可以使用F1度量使度量标准协同工作:F1 = 2 *(Bleu * Rouge)/(Bleu + Rouge)