如何将具有 numpy 数组值的 pandas 系列转换为数据帧

Olg*_*run 4 python numpy pandas

  1. 我有一个 pandas 系列,其值是 numpy 数组。为了简单起见,说
    系列 = pd.Series([np.array([1,2,3,4]), np.array([5,6,7,8]), np.array([9,10,11,12] )], 索引=['文件1', '文件2', '文件3'])
file1       [1, 2, 3, 4]
file2       [5, 6, 7, 8]
file3    [9, 10, 11, 12]
Run Code Online (Sandbox Code Playgroud)

如何将其扩展为以下形式的数据框df_concatenated:

       0   1   2   3
file1  1   2   3   4
file2  5   6   7   8
file3  9  10  11  12
Run Code Online (Sandbox Code Playgroud)
  1. 同一问题的更广泛版本。实际上,它series是从以下形式的不同数据帧获得的:

数据框:

              0   1
file  slide        
file1 1       1   2
      2       3   4
file2 1       5   6
      2       7   8
file3 1       9  10
      2      11  12
Run Code Online (Sandbox Code Playgroud)

通过对“文件”索引进行分组并串联列。

   def concat_sublevel(data):
        return np.concatenate(data.values)

   series = data.groupby(level=[0]).apply(concat_sublevel)
Run Code Online (Sandbox Code Playgroud)

可能有人看到了从 dataframedata到 df_concatenated.

警告。slide对于不同的值,子索引可以有不同数量的值file。在这种情况下,我需要重复其中一行才能在所有结果行中获得相同的尺寸

Nag*_*ran 5

您可以尝试使用记录中的 pandas Dataframe

pd.DataFrame.from_records(series.values,index=series.index)
Run Code Online (Sandbox Code Playgroud)

出去:

    0   1   2   3
file1   1   2   3   4
file2   5   6   7   8
file3   9   10  11  12
Run Code Online (Sandbox Code Playgroud)

  • 等效替代方案:`pd.DataFrame(series.values.tolist(), index=series.index)` (2认同)