从字符串中删除格式

use*_*704 3 python formatting encoding parsing web

我正在尝试使用BeautifulSoup解析一些来自网络的数据。到目前为止,我已经使用以下代码从表中获取了我需要的数据:

def webParsing(canvas):
url='http://www.cmu.edu/dining/hours/index.html'
try:
    page= urllib.urlopen(url)
except:
    print 'Error while opening html file. Please ensure that you',
    print ' have a working internet connection.'
    return
sourceCode=page.read()
soup=BeautifulSoup(sourceCode)
#heading=soup.html.body.div
tableData=soup.table.tbody
parseTable(canvas,tableData)
def parseTable(canvas,tableData):
    canvas.data.hoursOfOperation=dict()
    rowTag='tr'
    colTag='td'
    for row in tableData.find_all(rowTag):
        row_text=[]
        for item in row.find_all(colTag):
            text=item.text.strip()
            row_text.append(text)
        (locations,hoursOpen)=(row_text[0],row_text[1])
        locations=locations.split(',')
        for location in locations:
            canvas.data.hoursOfOperation[location]=hoursOpen
    print canvas.data.hoursOfOperation
Run Code Online (Sandbox Code Playgroud)

如您所见,第一列中的“项目”通过字典映射到第二列中的“项目”。数据几乎完全是我在打印时想要的样子,但是在python中,这些字符串中有很多格式,例如'\ n'或'\ xe9'或'\ n \ xao'。有什么办法可以删除所有格式?换句话说,删除所有换行符,表示特定编码的任何内容,表示带重音符号的任何内容,并仅获取字符串文字?我不需要最有效或最安全的方法,我是一名初学者,所以最好能采用最简单的方法!谢谢!

aIK*_*Kid 6

这是个窍门:您可以将其编码为ascii,然后删除其余所有内容:

>>> 'abc\xe9'.encode('ascii', errors='ignore')
b'abc'
Run Code Online (Sandbox Code Playgroud)

编辑:

啊,我忘记了您也不需要标准的特殊字符。使用此代替:

''.join(s for s in string if ord(s)>31 and ord(s)<126)
Run Code Online (Sandbox Code Playgroud)

希望这可以帮助!