如何使用python(Pandas)堆积条形集群

jrj*_*rjc 46 python plot matplotlib pandas seaborn

所以这是我的数据集的样子:

In [1]: df1=pd.DataFrame(np.random.rand(4,2),index=["A","B","C","D"],columns=["I","J"])

In [2]: df2=pd.DataFrame(np.random.rand(4,2),index=["A","B","C","D"],columns=["I","J"])

In [3]: df1
Out[3]: 
          I         J
A  0.675616  0.177597
B  0.675693  0.598682
C  0.631376  0.598966
D  0.229858  0.378817

In [4]: df2
Out[4]: 
          I         J
A  0.939620  0.984616
B  0.314818  0.456252
C  0.630907  0.656341
D  0.020994  0.538303
Run Code Online (Sandbox Code Playgroud)

我希望每个数据帧都有堆积条形图,但由于它们具有相同的索引,我希望每个索引有2个堆叠条形.

我试图在同一轴上绘制两个:

In [5]: ax = df1.plot(kind="bar", stacked=True)

In [5]: ax2 = df2.plot(kind="bar", stacked=True, ax = ax)
Run Code Online (Sandbox Code Playgroud)

但它重叠.

然后我尝试先连接两个数据集:

pd.concat(dict(df1 = df1, df2 = df2),axis = 1).plot(kind="bar", stacked=True)
Run Code Online (Sandbox Code Playgroud)

但这里一切都堆积如山

我最好的尝试是:

 pd.concat(dict(df1 = df1, df2 = df2),axis = 0).plot(kind="bar", stacked=True)
Run Code Online (Sandbox Code Playgroud)

这使 :

在此输入图像描述

这基本上是我想要的,除了我想要的酒吧订购

(df1,A)(df2,A)(df1,B)(df2,B)等......

我想有一个技巧,但我找不到它!


在@ bgschiller的回答之后我得到了这个:

在此输入图像描述

这几乎是我想要的.我希望通过索引对条形图进行聚类,以便在视觉上清晰.

额外:x标签不是多余的,如:

df1 df2    df1 df2
_______    _______ ...
   A          B
Run Code Online (Sandbox Code Playgroud)

谢谢你的帮助.

jrj*_*rjc 61

所以,我最终找到了一个技巧(编辑:请参阅下面的使用seaborn和longform数据帧):

解决方案与熊猫和matplotlib

这是一个更完整的例子:

import pandas as pd
import matplotlib.cm as cm
import numpy as np
import matplotlib.pyplot as plt

def plot_clustered_stacked(dfall, labels=None, title="multiple stacked bar plot",  H="/", **kwargs):
    """Given a list of dataframes, with identical columns and index, create a clustered stacked bar plot. 
labels is a list of the names of the dataframe, used for the legend
title is a string for the title of the plot
H is the hatch used for identification of the different dataframe"""

    n_df = len(dfall)
    n_col = len(dfall[0].columns) 
    n_ind = len(dfall[0].index)
    axe = plt.subplot(111)

    for df in dfall : # for each data frame
        axe = df.plot(kind="bar",
                      linewidth=0,
                      stacked=True,
                      ax=axe,
                      legend=False,
                      grid=False,
                      **kwargs)  # make bar plots

    h,l = axe.get_legend_handles_labels() # get the handles we want to modify
    for i in range(0, n_df * n_col, n_col): # len(h) = n_col * n_df
        for j, pa in enumerate(h[i:i+n_col]):
            for rect in pa.patches: # for each index
                rect.set_x(rect.get_x() + 1 / float(n_df + 1) * i / float(n_col))
                rect.set_hatch(H * int(i / n_col)) #edited part     
                rect.set_width(1 / float(n_df + 1))

    axe.set_xticks((np.arange(0, 2 * n_ind, 2) + 1 / float(n_df + 1)) / 2.)
    axe.set_xticklabels(df.index, rotation = 0)
    axe.set_title(title)

    # Add invisible data to add another legend
    n=[]        
    for i in range(n_df):
        n.append(axe.bar(0, 0, color="gray", hatch=H * i))

    l1 = axe.legend(h[:n_col], l[:n_col], loc=[1.01, 0.5])
    if labels is not None:
        l2 = plt.legend(n, labels, loc=[1.01, 0.1]) 
    axe.add_artist(l1)
    return axe

# create fake dataframes
df1 = pd.DataFrame(np.random.rand(4, 5),
                   index=["A", "B", "C", "D"],
                   columns=["I", "J", "K", "L", "M"])
df2 = pd.DataFrame(np.random.rand(4, 5),
                   index=["A", "B", "C", "D"],
                   columns=["I", "J", "K", "L", "M"])
df3 = pd.DataFrame(np.random.rand(4, 5),
                   index=["A", "B", "C", "D"], 
                   columns=["I", "J", "K", "L", "M"])

# Then, just call :
plot_clustered_stacked([df1, df2, df3],["df1", "df2", "df3"])
Run Code Online (Sandbox Code Playgroud)

它给出了:

多个堆积条形图

您可以通过传递cmap参数来更改栏的颜色:

plot_clustered_stacked([df1, df2, df3],
                       ["df1", "df2", "df3"],
                       cmap=plt.cm.viridis)
Run Code Online (Sandbox Code Playgroud)

用seaborn解决方案:

给定相同的df1,df2,df3,我将它们转换为长形式:

df1["Name"] = "df1"
df2["Name"] = "df2"
df3["Name"] = "df3"
dfall = pd.concat([pd.melt(i.reset_index(),
                           id_vars=["Name", "index"]) # transform in tidy format each df
                   for i in [df1, df2, df3]],
                   ignore_index=True)
Run Code Online (Sandbox Code Playgroud)

seaborn的问题在于它本身不会堆叠条形图,因此诀窍是将每个条形图的累积总和绘制在彼此之上:

dfall.set_index(["Name", "index", "variable"], inplace=1)
dfall["vcs"] = dfall.groupby(level=["Name", "index"]).cumsum()
dfall.reset_index(inplace=True) 

>>> dfall.head(6)
  Name index variable     value       vcs
0  df1     A        I  0.717286  0.717286
1  df1     B        I  0.236867  0.236867
2  df1     C        I  0.952557  0.952557
3  df1     D        I  0.487995  0.487995
4  df1     A        J  0.174489  0.891775
5  df1     B        J  0.332001  0.568868
Run Code Online (Sandbox Code Playgroud)

然后循环遍历每组variable并绘制累积总和:

c = ["blue", "purple", "red", "green", "pink"]
for i, g in enumerate(dfall.groupby("variable")):
    ax = sns.barplot(data=g[1],
                     x="index",
                     y="vcs",
                     hue="Name",
                     color=c[i],
                     zorder=-i, # so first bars stay on top
                     edgecolor="k")
ax.legend_.remove() # remove the redundant legends 
Run Code Online (Sandbox Code Playgroud)

多堆栈条形图seaborn

它缺乏我认为可以轻松添加的传奇.问题是,为了区分数据框而不是阴影(可以很容易地添加),我们有一个亮度梯度,而且它对于第一个有点太亮了,我真的不知道如何改变它而不改变每个矩形逐个(如第一个解决方案).

如果你不理解代码中的某些内容,请告诉我.

随意重复使用CC0下的代码.

  • 你能给我一个巨大的帮助,把这个snipplet放在BSD/MIT/CC-0下吗?谢谢 :) (2认同)

bil*_*oie 6

@jrjc 对 use of 的回答seaborn很聪明,但是有几个问题,作者指出:

  1. 当只需要两三个类别时,“浅色”阴影太苍白了。它使颜色系列(淡蓝色、蓝色、深蓝色等)难以区分。
  2. 生成图例不是为了区分阴影的含义(“苍白”是什么意思?)

然而,更重要的是,我发现,因为groupby代码中的语句:

  1. 此解决方案仅适用于按字母顺序排列列的情况。如果我["I", "J", "K", "L", "M"]用反字母 ( ["zI", "yJ", "xK", "wL", "vM"])重命名列,我会得到这个图

如果列不按字母顺序排列,则堆叠条形图构建失败


我努力通过这个开源 python 模块中plot_grouped_stackedbars()函数来解决这些问题。

  1. 将阴影保持在合理范围内
  2. 它会自动生成解释阴影的图例
  3. 它不依赖 groupby

带有图例和窄阴影范围的适当分组堆积条形图

它还允许

  1. 各种标准化选项(见下文标准化为最大值的 100%)
  2. 误差线的添加

带有归一化和误差条的示例

在此处查看完整演示。我希望这证明是有用的,并且可以回答最初的问题。


小智 6

这是Cord Kaldemeyer答案的更简洁的实现。这个想法是为绘图保留尽可能多的宽度。然后每个簇得到所需长度的子图。

# Data and imports

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
from matplotlib.ticker import MaxNLocator
import matplotlib.gridspec as gridspec
import matplotlib

matplotlib.style.use('ggplot')

np.random.seed(0)

df = pd.DataFrame(np.asarray(1+5*np.random.random((10,4)), dtype=int),columns=["Cluster", "Bar", "Bar_part", "Count"])
df = df.groupby(["Cluster", "Bar", "Bar_part"])["Count"].sum().unstack(fill_value=0)
display(df)

# plotting

clusters = df.index.levels[0]
inter_graph = 0
maxi = np.max(np.sum(df, axis=1))
total_width = len(df)+inter_graph*(len(clusters)-1)

fig = plt.figure(figsize=(total_width,10))
gridspec.GridSpec(1, total_width)
axes=[]

ax_position = 0
for cluster in clusters:
    subset = df.loc[cluster]
    ax = subset.plot(kind="bar", stacked=True, width=0.8, ax=plt.subplot2grid((1,total_width), (0,ax_position), colspan=len(subset.index)))
    axes.append(ax)
    ax.set_title(cluster)
    ax.set_xlabel("")
    ax.set_ylim(0,maxi+1)
    ax.yaxis.set_major_locator(MaxNLocator(integer=True))
    ax_position += len(subset.index)+inter_graph

for i in range(1,len(clusters)):
    axes[i].set_yticklabels("")
    axes[i-1].legend().set_visible(False)
axes[0].set_ylabel("y_label")

fig.suptitle('Big Title', fontsize="x-large")
legend = axes[-1].legend(loc='upper right', fontsize=16, framealpha=1).get_frame()
legend.set_linewidth(3)
legend.set_edgecolor("black")

plt.show()
Run Code Online (Sandbox Code Playgroud)

结果如下:

(还无法直接在网站上发布图像)


Cor*_*yer 5

我已经设法通过基本命令使用pandas和matplotlib子图来做同样的事情。

这是一个例子:

fig, axes = plt.subplots(nrows=1, ncols=3)

ax_position = 0
for concept in df.index.get_level_values('concept').unique():
    idx = pd.IndexSlice
    subset = df.loc[idx[[concept], :],
                    ['cmp_tr_neg_p_wrk', 'exp_tr_pos_p_wrk',
                     'cmp_p_spot', 'exp_p_spot']]     
    print(subset.info())
    subset = subset.groupby(
        subset.index.get_level_values('datetime').year).sum()
    subset = subset / 4  # quarter hours
    subset = subset / 100  # installed capacity
    ax = subset.plot(kind="bar", stacked=True, colormap="Blues",
                     ax=axes[ax_position])
    ax.set_title("Concept \"" + concept + "\"", fontsize=30, alpha=1.0)
    ax.set_ylabel("Hours", fontsize=30),
    ax.set_xlabel("Concept \"" + concept + "\"", fontsize=30, alpha=0.0),
    ax.set_ylim(0, 9000)
    ax.set_yticks(range(0, 9000, 1000))
    ax.set_yticklabels(labels=range(0, 9000, 1000), rotation=0,
                       minor=False, fontsize=28)
    ax.set_xticklabels(labels=['2012', '2013', '2014'], rotation=0,
                       minor=False, fontsize=28)
    handles, labels = ax.get_legend_handles_labels()
    ax.legend(['Market A', 'Market B',
               'Market C', 'Market D'],
              loc='upper right', fontsize=28)
    ax_position += 1

# look "three subplots"
#plt.tight_layout(pad=0.0, w_pad=-8.0, h_pad=0.0)

# look "one plot"
plt.tight_layout(pad=0., w_pad=-16.5, h_pad=0.0)
axes[1].set_ylabel("")
axes[2].set_ylabel("")
axes[1].set_yticklabels("")
axes[2].set_yticklabels("")
axes[0].legend().set_visible(False)
axes[1].legend().set_visible(False)
axes[2].legend(['Market A', 'Market B',
                'Market C', 'Market D'],
               loc='upper right', fontsize=28)
Run Code Online (Sandbox Code Playgroud)

分组之前“子集”的数据帧结构如下所示:

<class 'pandas.core.frame.DataFrame'>
MultiIndex: 105216 entries, (D_REC, 2012-01-01 00:00:00) to (D_REC, 2014-12-31 23:45:00)
Data columns (total 4 columns):
cmp_tr_neg_p_wrk    105216 non-null float64
exp_tr_pos_p_wrk    105216 non-null float64
cmp_p_spot          105216 non-null float64
exp_p_spot          105216 non-null float64
dtypes: float64(4)
memory usage: 4.0+ MB
Run Code Online (Sandbox Code Playgroud)

和这样的情节:

在此处输入图片说明

它的格式为“ ggplot”,带有以下标头:

import pandas as pd
import matplotlib.pyplot as plt
import matplotlib
matplotlib.style.use('ggplot')
Run Code Online (Sandbox Code Playgroud)

  • 您能否添加示例数据,以便可以重现。 (3认同)
  • 很好的答案,但如果没有要复制的数据,就很难遵循。是否可以在某处下载数据? (2认同)

Gra*_*eth 5

这是一个很好的开始,但我认为可以对颜色进行一些修改以使其更加清晰。另外,在导入Altair中的每个参数时也要小心,因为这可能导致与命名空间中的现有对象发生冲突。这是一些重新配置的代码,用于在堆叠值时显示正确的颜色显示:

Altair集群柱形图

导入包

import pandas as pd
import numpy as np
import altair as alt
Run Code Online (Sandbox Code Playgroud)

产生一些随机数据

df1=pd.DataFrame(10*np.random.rand(4,3),index=["A","B","C","D"],columns=["I","J","K"])
df2=pd.DataFrame(10*np.random.rand(4,3),index=["A","B","C","D"],columns=["I","J","K"])
df3=pd.DataFrame(10*np.random.rand(4,3),index=["A","B","C","D"],columns=["I","J","K"])

def prep_df(df, name):
    df = df.stack().reset_index()
    df.columns = ['c1', 'c2', 'values']
    df['DF'] = name
    return df

df1 = prep_df(df1, 'DF1')
df2 = prep_df(df2, 'DF2')
df3 = prep_df(df3, 'DF3')

df = pd.concat([df1, df2, df3])
Run Code Online (Sandbox Code Playgroud)

用Altair绘制数据

alt.Chart(df).mark_bar().encode(

    # tell Altair which field to group columns on
    x=alt.X('c2:N', title=None),

    # tell Altair which field to use as Y values and how to calculate
    y=alt.Y('sum(values):Q',
        axis=alt.Axis(
            grid=False,
            title=None)),

    # tell Altair which field to use to use as the set of columns to be  represented in each group
    column=alt.Column('c1:N', title=None),

    # tell Altair which field to use for color segmentation 
    color=alt.Color('DF:N',
            scale=alt.Scale(
                # make it look pretty with an enjoyable color pallet
                range=['#96ceb4', '#ffcc5c','#ff6f69'],
            ),
        ))\
    .configure_view(
        # remove grid lines around column clusters
        strokeOpacity=0    
    )
Run Code Online (Sandbox Code Playgroud)