将两列设置为 Pandas 数据框中的索引以进行时间序列分析

yos*_*rry 3 python indexing time-series pandas

在天气或股票市场数据的情况下,温度和股票价格都是在任何给定日期的多个站点或股票行情中测量的。

因此,设置包含两个字段的索引的最有效方法是什么?

对于天气:weather_station 然后是 Date

对于股票数据:stock_code 然后是日期

以这种方式设置索引将允许过滤,例如:

  • stock_df["code"]["start_date":"end_date"]
  • weather_df["station"]["start_date":"end_date"]

Ale*_*der 7

该功能目前已存在。请参阅文档以获取更多示例。

stock_df = pd.DataFrame({'symbol': ['AAPL', 'AAPL', 'F', 'F', 'F'], 
                         'date': ['2016-1-1', '2016-1-2', '2016-1-1', '2016-1-2', '2016-1-3'], 
                         'price': [100., 101, 50, 47.5, 49]}).set_index(['symbol', 'date'])

>>> stock_df
                 price
symbol date           
AAPL   2016-1-1  100.0
       2016-1-2  101.0
F      2016-1-1   50.0
       2016-1-2   47.5
       2016-1-3   49.0

>>> stock_df.loc['AAPL']
          price
date           
2016-1-1    100
2016-1-2    101

>>> stock_df.loc['AAPL', '2016-1-2']
price    101
Name: (AAPL, 2016-1-2), dtype: float64
Run Code Online (Sandbox Code Playgroud)


小智 7

正如安东所提到的,您需要按如下方式使用 MultiIndex:

stock_df.index = pd.MultiIndex.from_arrays(stock_df[['code', 'date']].values.T, names=['idx1', 'idx2'])

weather_df.index = pd.MultiIndex.from_arrays(weather_df[['station', 'date']].values.T, names=['idx1', 'idx2'])
Run Code Online (Sandbox Code Playgroud)