我有一个 pyspark 数据框:
Location Month Brand Sector TrueValue PickoutValue
USA 1/1/2021 brand1 cars1 7418 30000
USA 2/1/2021 brand1 cars1 1940 2000
USA 3/1/2021 brand1 cars1 4692 2900
USA 4/1/2021 brand1 cars1
USA 1/1/2021 brand2 cars2 16383104.2 16666667
USA 2/1/2021 brand2 cars2 26812874.2 16666667
USA 3/1/2021 brand2 cars2
USA 1/1/2021 brand3 cars3 75.6% 70.0%
USA 3/1/2021 brand3 cars3 73.1% 70.0%
USA 2/1/2021 brand3 cars3 77.1% 70.0%
Run Code Online (Sandbox Code Playgroud)
我有每个品牌从 1/1/2021 到 12/1/2021 的月份值。我需要创建另一列,其中包含基于品牌和部门并按月排序的 TrueValue 列的累积总和。具有%值的行应该是累积总和除以月数。
我的预期数据框是:
Location Month Brand Sector TrueValue PickoutValue TotalSumValue …Run Code Online (Sandbox Code Playgroud)