python可以编码为utf-8但不能解码

Question

python可以编码为utf-8但不能解码

gho*_*bzi 4 python decode utf-8 python-3.x

下面的代码可以将字符串编码为 utf-8 ：

#!/usr/bin/python
# -*- coding: utf-8 -*-

str = '????'
print(str.encode('utf-8'))

Run Code Online (Sandbox Code Playgroud)

打印：

b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'

Run Code Online (Sandbox Code Playgroud)

但我无法使用此代码解码此字符串：

#!/usr/bin/python
# -*- coding: utf-8 -*-

str = b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'
print(str.decode('utf-8'))

Run Code Online (Sandbox Code Playgroud)

错误是：

Traceback (most recent call last):
  File "C:\test.py", line 5, in <module>
    print(str.decode('utf-8'))
AttributeError: 'str' object has no attribute 'decode'

Run Code Online (Sandbox Code Playgroud)

请帮我 ...

编辑

从答案切换到字节字符串：

#!/usr/bin/python
# -*- coding: utf-8 -*-

str = b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'
print(str.decode('utf-8'))

Run Code Online (Sandbox Code Playgroud)

现在错误是：

Traceback (most recent call last):
  File "C:\test.py", line 5, in <module>
    print(str.decode('utf-8'))
  File "C:\Python34\lib\encodings\cp437.py", line 19, in encode
    return codecs.charmap_encode(input,self.errors,encoding_map)[0]
UnicodeEncodeError: 'charmap' codec can't encode characters in position 0-3: character maps to <undefined>

Run Code Online (Sandbox Code Playgroud)

Answer 1

Mar*_*nen 5

看起来您使用的是 Python 3.X。你.encode()Unicode 字符串（u'xxx'或'xxx'）。你.decode()字节串b'xxxx'。

#!/usr/bin/python
# -*- coding: utf-8 -*-

s = b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'
#   ^
#   Need a 'b'
#
print(s.decode('utf-8'))

Run Code Online (Sandbox Code Playgroud)

请注意，您的终端可能无法显示 Unicode 字符串。我的 Windows 控制台没有：

Python 3.3.5 (v3.3.5:62cf4e77f785, Mar  9 2014, 10:35:05) [MSC v.1600 64 bit (AMD64)] on win32
Type "help", "copyright", "credits" or "license" for more information.
>>> s = b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'
>>> #   ^
... #   Need a 'b'
... #
... print(s.decode('utf-8'))
Traceback (most recent call last):
  File "<stdin>", line 4, in <module>
  File "D:\dev\Python33x64\lib\encodings\cp437.py", line 19, in encode
    return codecs.charmap_encode(input,self.errors,encoding_map)[0]
UnicodeEncodeError: 'charmap' codec can't encode characters in position 0-3: character maps to <undefined>

Run Code Online (Sandbox Code Playgroud)

但它确实进行了解码。 '\uxxxx'表示一个 Unicode 代码点。

>>> s.decode('utf-8')
'\u0648\u0631\u0648\u062f'

Run Code Online (Sandbox Code Playgroud)

我的 PythonWin IDE 支持 UTF-8 并且可以显示字符：

>>> s = b'\xd9\x88\xd8\xb1\xd9\x88\xd8\xaf'
>>> print(s.decode('utf-8'))
????

Run Code Online (Sandbox Code Playgroud)

您还可以将数据写入文件并在支持 UTF-8 的编辑器（如记事本）中显示。由于您的原始字符串已经是 UTF-8，因此只需将其直接作为字节写入文件即可。 'wb'以二进制模式打开文件，字节按原样写入：

>>> with open('out.txt','wb') as f:
...     f.write(s)

Run Code Online (Sandbox Code Playgroud)

如果你有一个 Unicode 字符串，你可以将它写成 UTF-8：

>>> with open('out.txt','w',encoding='utf8') as f:
...     f.write(u)  # assuming "u" is already a decoded Unicode string.

Run Code Online (Sandbox Code Playgroud)

PSstr是内置类型。不要将它用于变量名。

Python 2.x 的工作方式不同。 'xxxx'是一个字节串并且u'xxxx'是一个Unicode串，但你仍然.encode()是Unicode串和.decode()字节串。

归档时间：	11 年，3 月前
查看次数：	5804 次
最近记录：	11 年，3 月前