解析非标准分号"JSON"

mor*_*ree 1 python parsing json

我有一个非标准的"JSON"文件来解析.每个项目以分号分隔,而不是以逗号分隔.我不能简单地更换;与,因为可能有一些含有价值;,恩."你好,世界".我如何将其解析为JSON通常会解析的结构?

{
  "client" : "someone";
  "server" : ["s1"; "s2"];
  "timestamp" : 1000000;
  "content" : "hello; world";
  ...
}
Run Code Online (Sandbox Code Playgroud)

Mar*_*ers 6

使用Python tokenize模块将文本流转换为逗号而不是分号.Python tokenizer也很乐意处理JSON输入,甚至包括分号.标记生成器将字符串作为整个标记呈现,并且"原始"分号在流中作为单个token.OP标记供您替换:

import tokenize
import json

corrected = []

with open('semi.json', 'r') as semi:
    for token in tokenize.generate_tokens(semi.readline):
        if token[0] == tokenize.OP and token[1] == ';':
            corrected.append(',')
        else:
            corrected.append(token[1])

data = json.loads(''.join(corrected))
Run Code Online (Sandbox Code Playgroud)

这假设一旦用逗号替换分号,格式就变为有效的JSON; 例如,在关闭]或}允许之前没有尾随逗号,尽管您甚至可以跟踪添加的最后一个逗号,如果下一个非换行标记是右括号,则再次删除它.

演示:

>>> import tokenize
>>> import json
>>> open('semi.json', 'w').write('''\
... {
...   "client" : "someone";
...   "server" : ["s1"; "s2"];
...   "timestamp" : 1000000;
...   "content" : "hello; world"
... }
... ''')
>>> corrected = []
>>> with open('semi.json', 'r') as semi:
...     for token in tokenize.generate_tokens(semi.readline):
...         if token[0] == tokenize.OP and token[1] == ';':
...             corrected.append(',')
...         else:
...             corrected.append(token[1])
...
>>> print ''.join(corrected)
{
"client":"someone",
"server":["s1","s2"],
"timestamp":1000000,
"content":"hello; world"
}
>>> json.loads(''.join(corrected))
{u'content': u'hello; world', u'timestamp': 1000000, u'client': u'someone', u'server': [u's1', u's2']}
Run Code Online (Sandbox Code Playgroud)

国米令牌空白被放弃了,而是开始重视可以重新设置tokenize.NL令牌和(lineno, start)和(lineno, end)是每个标记的部分位置的元组.由于令牌周围的空白对JSON解析器无关紧要,我对此并不感到困扰.