根据 Pandas DataFrame 中其他列的条件创建新列

ebu*_*168 3 python mapping dataframe pandas

我有这个数据框:

+------+--------------+------------+
| ID   | Education    |      Score | 
+------+--------------+------------+
|    1 |  High School |      7.884 |     
|    2 |  Bachelors   |      6.952 |     
|    3 |  High School |      8.185 |   
|    4 |  High School |      6.556 | 
|    5 |  Bachelors   |      6.347 | 
|    6 |  Master      |      6.794 |   
+------+--------------+------------+
Run Code Online (Sandbox Code Playgroud)

我想创建一个对分数列进行分类的新列。我想给它贴上标签:“差”、“好”、“非常好”。

这可能看起来像这样:

+------+--------------+------------+------------+
| ID   | Education    |      Score | Labels     |
+------+--------------+------------+------------+
|    1 |  High School |      7.884 | Good       |
|    2 |  Bachelors   |      6.952 | Bad        |
|    3 |  High School |      8.185 | Very good  |   
|    4 |  High School |      6.556 | Bad        |
|    5 |  Bachelors   |      6.347 | Bad        |
|    6 |  Master      |      6.794 | Bad        |
+------+--------------+------------+------------+
Run Code Online (Sandbox Code Playgroud)

我怎样才能做到这一点?

提前致谢

Asg*_*eer 7

我想这是您想要映射到标签的分数。您可以定义一个映射函数,将分数作为输入,然后返回标签:

def map_score(score):
  if score >= 8:
    return "Very good"
  elif score >= 7:
    return "Good"
  else:
    return "Bad"

df["Labels"] = df["Score"].apply(map_score)
Run Code Online (Sandbox Code Playgroud)


Pra*_*ini 5

import pandas as pd 

# initialize list of lists 
data = [[1,'High School',7.884], [2,'Bachelors',6.952], [3,'High School',8.185], [4,'High School',6.556],[5,'Bachelors',6.347],[6,'Master',6.794]] 

# Create the pandas DataFrame 
df = pd.DataFrame(data, columns = ['ID', 'Education', 'Score']) 

df['Labels'] = ['Bad' if x<7.000 else 'Good' if 7.000<=x<8.000 else 'Very Good' for x in df['Score']]
df

    ID  Education    Score    Labels
0   1   High School  7.884    Good
1   2   Bachelors    6.952    Bad
2   3   High School  8.185    Very Good
3   4   High School  6.556    Bad
4   5   Bachelors    6.347    Bad
5   6   Master       6.794    Bad
Run Code Online (Sandbox Code Playgroud)

  • 只是一个提示: `df['labels']=np.select([df['Score']&lt;7,df['Score']. Between(7,8)],['Bad','Good'] ,'非常好')`,`np.select` 将以矢量化方式工作,速度更快:) (2认同)