小编Ivo*_*Ivo的帖子

熊猫使用基于第二列的另一列方法替换NaN

我有一个带有两列的熊猫数据框,city和country。双方city并country包含缺失值。考虑以下数据帧:

temp = pd.DataFrame({"country": ["country A", "country A", "country A", "country A", "country B","country B","country B","country B", "country C", "country C", "country C", "country C"],
                     "city": ["city 1", "city 2", np.nan, "city 2", "city 3", "city 3", np.nan, "city 4", "city 5", np.nan, np.nan, "city 6"]})
Run Code Online (Sandbox Code Playgroud)

我现在想在填写NaN在S city与列模式该国在剩余的数据帧的城市,例如对于A国:城市1被提及一次; 提到城市2两次;因此,用etc 填充city索引处的列。2city 2

我已经做好了

cities = [city for city in temp["country"].value_counts().index]
modes = temp.groupby(["country"]).agg(pd.Series.mode)
dict_locations …
Run Code Online (Sandbox Code Playgroud)

python pandas

3
推荐指数
1
解决办法
65
查看次数

检查命令行输入是否为数字

我是C的新手,想要编写一个简单的程序,用两个命令行参数调用(第一个是程序的名称,第二个是字符串).我已经验证了参数的数量,现在想验证输入只包含数字.

#include <stdio.h>
#include <string.h>
#include <ctype.h>
#include <cs50.h>

int main(int argc, string argv[])
{
    if (argc == 2)
        {
            for (int i = 0, n = strlen(argv[1]); i < n; i++)
            {
                if isdigit(argv[1][i])
                {
                    printf("Success!\n");
                }
                else
                    {
                    printf("Error. Second argument can be numbers only\n");
                    return 1;
                    }    
            }
    else
        {
            printf("Error. Second argument can be numbers only\n");
            return 1;
        }
}
Run Code Online (Sandbox Code Playgroud)

虽然上面的代码没有抛出错误,但它没有正确验证输入,我无法弄清楚原因.有人能指出我正确的方向吗?

非常感谢提前和所有最好的

c cs50

2
推荐指数
1
解决办法
182
查看次数

从列中提取国家/地区名称(或其他实体)

我data.frame在列中包含国家和城市location,我想通过与world.cities$country.etc来自library(maps)(或任何其他国家/地区名称集合)的数据框匹配来提取前者。

考虑这个例子:

df <- data.frame(location = c("Aarup, Denmark",
                              "Switzerland",
                              "Estonia: Aaspere"),
                 other_col = c(2,3,4))
Run Code Online (Sandbox Code Playgroud)

我尝试使用这段代码

df %>% extract(location,
               into = c("country", "rest_location"),
               remove = FALSE,
               function(x) x[which x %in% world.cities$country.etc])
Run Code Online (Sandbox Code Playgroud)

但我没有成功;我期望这样的事情:

          location other_col     country rest_location
1   Aarup, Denmark         2     Denmark       Aarup, 
2      Switzerland         3 Switzerland              
3 Estonia: Aaspere         4     Estonia     : Aaspere
Run Code Online (Sandbox Code Playgroud)

r dataframe

2
推荐指数
1
解决办法
2074
查看次数

根据标记数据创建带有标签的列

我有一个带有标签数据的数据集,并且想创建一个仅包含标签作为字符的新列。

考虑以下示例:

value_labels <- tibble(value = 1:6, labels = paste0("value", 1:6))
df_data <- tibble(id = 1:10, var = floor(runif(10, 1, 6)))
df_data <- df_data %>% mutate(var = haven::labelled(var, labels = deframe(value_labels[2:1])))
Run Code Online (Sandbox Code Playgroud)

这产生:

# A tibble: 10 x 2
      id        var
   <int>  <dbl+lbl>
 1     1 2 [value2]
 2     2 2 [value2]
 3     3 4 [value4]
 4     4 2 [value2]
 5     5 4 [value4]
 6     6 3 [value3]
 7     7 5 [value5]
 8     8 4 [value4]
 9     9 3 [value3]
10    10 …
Run Code Online (Sandbox Code Playgroud)

r r-haven tidyverse r-labelled

2
推荐指数
1
解决办法
1088
查看次数

在 sns.lmplot() 中格式化 x 轴(日期)

我需要用 来绘制每日数据sns.lmplot()。

数据具有以下结构:

df = pd.DataFrame(columns=['date', 'origin', 'group', 'value'],
                  data = [['2001-01-01', "Peter", "A", 1.0],
                          ['2011-01-01', "Peter", "A", 1.1],
                          ['2011-01-02', "Peter", "B", 1.2],
                          ['2012-01-03', "Peter", "A", 1.3],
                          ['2012-01-01', "Peter", "B", 1.4],
                          ['2013-01-02', "Peter", "A", 1.5],
                          ['2013-01-03', "Peter", "B", 1.6],
                          ['2021-01-01', "Peter", "A", 1.7]])
Run Code Online (Sandbox Code Playgroud)

我现在想用每月平均值绘制数据sns.lmplot()(我的原始数据比玩具数据更细粒度)并使用hueforgroup列。为此,我按月汇总:

df['date'] = pd.to_datetime(df['date']).dt.strftime('%Y%M').astype(int)
df = df.groupby(['date', 'origin', 'group']).agg(['mean'])
df.columns = ["_".join(pair) for pair in df.columns]  # reset col multi-index
df = df.reset_index()  # reset index
Run Code Online (Sandbox Code Playgroud)

然后我绘制数据:

sns.lmplot(data=df, x="date", …
Run Code Online (Sandbox Code Playgroud)

python datetime matplotlib seaborn lmplot

1
推荐指数
1
解决办法
1355
查看次数

标签 统计

python ×2

r ×2

c ×1

cs50 ×1

dataframe ×1

datetime ×1

lmplot ×1

matplotlib ×1

pandas ×1

r-haven ×1

r-labelled ×1

seaborn ×1

tidyverse ×1