如何修改此函数以返回 4d 数组而不是 3d 数组?

ari*_*wan 5 python numpy dataframe pandas numpy-ndarray

我创建了这个函数,它接受 adataframe以返回ndarrays输入和标签。

def transform_to_array(dataframe, chunk_size=100):
    
    grouped = dataframe.groupby('id')

    # initialize accumulators
    X, y = np.zeros([0, 1, chunk_size, 4]), np.zeros([0,]) # original inpt shape: [0, 1, chunk_size, 4]

    # loop over each group (df[df.id==1] and df[df.id==2])
    for _, group in grouped:

        inputs = group.loc[:, 'A':'D'].values 
        label = group.loc[:, 'label'].values[0]

        # calculate number of splits
        N = (len(inputs)-1) // chunk_size

        if N > 0:
            inputs = np.array_split(
                 inputs, [chunk_size + (chunk_size*i) for i in range(N)])
        else:
            inputs = [inputs]

        # loop over splits
        for inpt in inputs:
            inpt = np.pad(
                inpt, [(0, chunk_size-len(inpt)),(0, 0)], 
                mode='constant')
            # add each inputs split to accumulators
            X = np.concatenate([X, inpt[np.newaxis, np.newaxis]], axis=0)
            y = np.concatenate([y, label[np.newaxis]], axis=0) 

    return X, y
Run Code Online (Sandbox Code Playgroud)

该函数返回Xshape(n_samples, 1, chunk_size, 4)和yshape (n_samples, )。

举些例子:

N = 10_000
id = np.arange(N)
labels = np.random.randint(5, size=N)
df = pd.DataFrame(data = np.random.randn(N, 4),  columns=list('ABCD'))

df['label'] = labels
df.insert(0, 'id', id)
df = df.loc[df.id.repeat(157)]

df.head()
    id      A            B          C            D    label
0   0   -0.571676   -0.337737   -0.019276   -1.377253   1
0   0   -0.571676   -0.337737   -0.019276   -1.377253   1
0   0   -0.571676   -0.337737   -0.019276   -1.377253   1
0   0   -0.571676   -0.337737   -0.019276   -1.377253   1
0   0   -0.571676   -0.337737   -0.019276   -1.377253   1
Run Code Online (Sandbox Code Playgroud)

生成以下内容:

X, y = transform_to_array(df)

X.shape   # shape of input
(20000, 1, 100, 4)
y.shape   # shape of label
(20000,)
Run Code Online (Sandbox Code Playgroud)

该函数按预期工作正常,但是需要很长时间才能完成执行:

start_time = time.time()
X, y = transform_to_array(df)
end_time = time.time()
print(f'Time taken: {end_time - start_time} seconds.')
Time taken: 227.83956217765808 seconds.
Run Code Online (Sandbox Code Playgroud)

为了提高函数的性能(最小化执行时间),我创建了以下修改后的函数:

def modified_transform_to_array(dataframe, chunk_size=100):
    # group data by 'id'
    grouped = dataframe.groupby('id')
    # initialize lists to store transformed data
    X, y = [], []

    # loop over each group (df[df.id==1] and df[df.id==2])
    for _, group in grouped:
        # get input and label data for group
        inputs = group.loc[:, 'A':'D'].values 
        label = group.loc[:, 'label'].values[0]

        # calculate number of splits
        N = (len(inputs)-1) // chunk_size

        if N > 0:
            # split input data into chunks
            inputs = np.array_split(
             inputs, [chunk_size + (chunk_size*i) for i in range(N)])
        else:
            inputs = [inputs]

        # loop over splits
        for inpt in inputs:
            # pad input data to have a chunk size of chunk_size
            inpt = np.pad(
            inpt, [(0, chunk_size-len(inpt)),(0, 0)], 
                mode='constant')
            # add each input split and corresponding label to lists
            X.append(inpt)
            y.append(label)

    # convert lists to numpy arrays
    X = np.array(X)
    y = np.array(y)

    return X, y
Run Code Online (Sandbox Code Playgroud)

起初,我似乎成功地减少了所花费的时间:

start_time = time.time()
X2, y2 = modified_transform_to_array(df)
end_time = time.time()
print(f'Time taken: {end_time - start_time} seconds.')
Time taken: 5.842168092727661 seconds.
Run Code Online (Sandbox Code Playgroud)

然而,结果是它改变了预期返回值的形状。

X2.shape  # this should be (20000, 1, 100, 4)
(20000, 100, 4)

y.shape  # this is fine
(20000, )
Run Code Online (Sandbox Code Playgroud)

问题

由于速度更快,如何修改modified_transform_to_array()以返回预期的数组形状?(n_samples, 1, chunk_size, 4)

nor*_*ok2 4

您可以简单地reshape在X末尾返回它之前modified_transform_to_array(),例如:

def modified_transform_to_array( ... ):

    ...

    # convert lists to numpy arrays
    X = np.array(X)
    y = np.array(y)
    X = X.reshape((X.shape[0], 1, *X.shape[1:]))  # <-- THIS LINE
    return X, y
Run Code Online (Sandbox Code Playgroud)

或者,等效地:

X = X.reshape((X.shape[0], 1, X.shape[1], X.shape[2]))
Run Code Online (Sandbox Code Playgroud)

正如@MSS 的回答中所指出的,您也可以通过切片实现相同的重塑结果,方法是从选择整个数组(即)的 aa 切片开始,然后在您想要的位置X[:, :, :]插入 a None(或其更明确的别名)np.newaxis增加维数:

X = X[:, None, :, :]
X = X[:, np.newaxis, :, :]
Run Code Online (Sandbox Code Playgroud)

最后两个切片可以用省略号代替,省略号...本质上会产生足够的全轴切片(即:或slice(None))来填充整个数组维度。

X = X[:, None, ...]
X = X[:, np.newaxis, ...]
Run Code Online (Sandbox Code Playgroud)

您可能需要阅读NumPy 用户指南的相关部分,以获取有关NumPy 切片的使用和切片的进一步说明。NoneEllipsis

  • * 运算符将列表或元组解包为其组成部分。在这里,它将形状切片从 1 扩展到末尾。在我的回答中,我使用 np.newaxis 和 ellipses... 来执行相同的操作。 (2认同)
  • @MSS我编译了很多(也许这些天有更多的训练)https://xkcd.com/303/:-) (2认同)