计算字典中的单词(Python)

use*_*517 1 python dictionary counting

我有这个代码,我想打开一个指定的文件,然后每次有一个while循环它会计算它,最后输出特定文件中的while循环总数.我决定将输入文件转换为字典,然后创建一个for循环,每次看到单词后跟一个空格时,它会在最后打印WHILE_之前向WHILE_添加+1计数.

然而,这似乎不起作用,我不知道为什么.任何帮助解决这个问题将非常感激.

这是我目前的代码:

WHILE_ = 0
INPUT_ = input("Enter file or directory: ")


OPEN_ = open(INPUT_)
READLINES_ = OPEN_.readlines()
STRING_ = (str(READLINES_))
STRIP_ = STRING_.strip()
input_str1 = STRIP_.lower()


dic = dict()
for w in input_str1.split():
    if w in dic.keys():
        dic[w] = dic[w]+1
    else:
        dic[w] = 1
DICT_ = (dic)


for LINE_ in DICT_:
    if  ("while\\n',") in LINE_:
        WHILE_ += 1
    elif ('while\\n",') in LINE_:
        WHILE_ += 1
    elif ('while ') in LINE_:
        WHILE_ += 1

print ("while_loops {0:>12}".format((WHILE_)))
Run Code Online (Sandbox Code Playgroud)

这是我正在使用的输入文件:

'''A trivial test of metrics
Author: Angus McGurkinshaw
Date: May 7 2013
'''

def silly_function(blah):
    '''A silly docstring for a silly function'''
    def nested():
        pass
    print('Hello world', blah + 36 * 14)
    tot = 0  # This isn't a for statement
    for i in range(10):
        tot = tot + i
        if_im_done = false  # Nor is this an if
    print(tot)

blah = 3
while blah > 0:
    silly_function(blah)
    blah -= 1
    while True:
        if blah < 1000:
            break
Run Code Online (Sandbox Code Playgroud)

输出应该是2,但我的代码目前打印0

aba*_*ert 6

这是一个非常奇怪的设计.您正在调用readlines获取字符串列表,然后调用str该列表,该列表将整个事物连接成一个大字符串,引号repr用逗号连接并用方括号括起,然后将结果拆分为空格.我不知道为什么你会做这样的事情.

你奇怪的变量名,额外无用的代码行DICT_ = (dic)等等只会使事情进一步混淆.

但我可以解释为什么它不起作用.DICT_在你做了所有那些愚蠢之后尝试打印,你会发现包含的唯一键while是while和'while.由于这些都不符合您要查找的任何模式,因此您的计数最终为0.

同样值得注意的是,WHILE_即使模式有多个实例,您也只会添加1 ,因此您的整个计数单是没用的.


如果您不混淆字符串,尝试恢复它们,然后尝试匹配错误恢复的版本,这将更容易.直接做吧.

虽然我正在使用它,但我还要解决其他一些问题,以便您的代码可读,更简单,并且不会泄漏文件,等等.以下是您试图手动破解的逻辑的完整实现:

import collections

filename = input("Enter file: ")
counts = collections.Counter()
with open(filename) as f:
    for line in f:
        counts.update(line.strip().lower().split())
print('while_loops {0:>12}'.format(counts['while']))
Run Code Online (Sandbox Code Playgroud)

当您在示例输入上运行此操作时,您就可以正确获取2.并将其扩展到处理if并且for是微不足道和明显的.


但请注意,您的逻辑中存在严重问题:任何看起来像关键字但位于注释或字符串中间的内容仍然会被拾取.如果不编写某种代码来删除注释和字符串,就无法解决这个问题.这意味着你将超额计算if并且for通过1.明显的剥离方式 - line.partition('#')[0]以及类似的报价 - 不会起作用.首先,在if关键字之前有一个字符串是完全有效的,如"foo" if x else "bar".其次,你不能用这种方式处理多行字符串.

这些问题和其他类似的问题就是你几乎肯定想要一个真正的解析器的原因.如果您只是尝试解析Python代码,那么标准库中的ast模块就是显而易见的方法.如果你想为各种不同的语言编写快速和脏的解析器,试试pyparsing,这是非常好的,并附带一些很好的例子.

这是一个简单的例子:

import ast

filename = input("Enter file: ")
with open(filename) as f:
    tree = ast.parse(f.read())
while_loops = sum(1 for node in ast.walk(tree) if isinstance(node, ast.While))
print('while_loops {0:>12}'.format(while_loops))
Run Code Online (Sandbox Code Playgroud)

或者,更灵活:

import ast
import collections

filename = input("Enter file: ")
with open(filename) as f:
    tree = ast.parse(f.read())
counts = collections.Counter(type(node).__name__ for node in ast.walk(tree))    
print('while_loops {0:>12}'.format(counts['While']))
print('for_loops {0:>14}'.format(counts['For']))
print('if_statements {0:>10}'.format(counts['If']))
Run Code Online (Sandbox Code Playgroud)