Pav*_*put 6 python excel json etl
我正在创建一个 ML 模型,它将使用 JSON 文件来理解模式和响应格式。由于我的数据为 excel 格式,因此我在 python 中将其转换为 JSON。
这是代码:
import xlrd
from collections import OrderedDict
import simplejson as json
# Open the workbook and select the first worksheet
wb = xlrd.open_workbook('D:\\android\\testdata2.xlsx')
sh = wb.sheet_by_index(0)
# List to hold dictionaries
data_list = []
# Iterate through each row in worksheet and fetch values into dict
for rownum in range(1, sh.nrows):
data = OrderedDict()
row_values = sh.row_values(rownum)
data['pattern'] = row_values[0]
data['response'] = row_values[1]
data_list.append(data)
# Serialize the list of dicts to JSON
j = json.dumps(data_list)
# Write to file
with open('data1.json', 'w') as f:
f.write(j)
Run Code Online (Sandbox Code Playgroud)
我得到的输出为:
[{
"pattern": "WALLSTENT NON COUVERTE ",
"response": "ENDOPROTHESE STENT VASCULAIRE "
}, {
"pattern": "PRIMEADVANCED SURSCAN MRI ",
"response": "NEUROSTIMULATEUR NERF VAGUE GAUCHE "
}, {
"pattern": "AVASTIN FLACON DE",
"response": "BEVACIZUMAB"
}, {
"pattern": "PERJETA SOLUTION A DILUER POUR PERFUSION",
"response": "BRENTUXIMAB VEDOTIN"
}]
Run Code Online (Sandbox Code Playgroud)
我正在寻找的所需输出是这样的:
{
"intents": [{
"pattern": ["WALLSTENT, NON, COUVERTE "],
"response": ["ENDOPROTHESE STENT VASCULAIRE] "
}, {
"pattern": ["PRIMEADVANCED ,SURSCAN ,MRI"] ,
"response": ["NEUROSTIMULATEUR NERF VAGUE GAUCHE "]
}, {
"pattern": ["AVASTIN , FLACON ,DE"],
"response": ["BEVACIZUMAB"]
}, {
"pattern": ["PERJETA, SOLUTION, A, DILUER, POUR ,PERFUSION"],
"response": ["BRENTUXIMAB VEDOTIN"]
}]
}
Run Code Online (Sandbox Code Playgroud)
我可以在我的函数中做哪些修改来获得我正在寻找的输出。
应该这样做:
import xlrd
from collections import OrderedDict
import simplejson as json
# Open the workbook and select the first worksheet
wb = xlrd.open_workbook('D:\\android\\testdata2.xlsx')
sh = wb.sheet_by_index(0)
# List to hold dictionaries
data_list = []
# Iterate through each row in worksheet and fetch values into dict
for rownum in range(1, sh.nrows):
data = OrderedDict()
row_values = sh.row_values(rownum)
data['pattern'] = row_values[0]
data['response'] = row_values[1]
data_list.append(data)
data_list = {'intents': data_list} # Added line
# Serialize the list of dicts to JSON
j = json.dumps(data_list)
# Write to file
with open('data1.json', 'w') as f:
f.write(j)
Run Code Online (Sandbox Code Playgroud)
注意添加的data_list = {'intents': data_list}.
| 归档时间: |
|
| 查看次数: |
13445 次 |
| 最近记录: |