我有一个带有两列的熊猫数据框,city和country。双方city并country包含缺失值。考虑以下数据帧:
temp = pd.DataFrame({"country": ["country A", "country A", "country A", "country A", "country B","country B","country B","country B", "country C", "country C", "country C", "country C"],
"city": ["city 1", "city 2", np.nan, "city 2", "city 3", "city 3", np.nan, "city 4", "city 5", np.nan, np.nan, "city 6"]})
Run Code Online (Sandbox Code Playgroud)
我现在想在填写NaN在S city与列模式该国在剩余的数据帧的城市,例如对于A国:城市1被提及一次; 提到城市2两次;因此,用etc 填充city索引处的列。2city 2
我已经做好了
cities = [city for city in temp["country"].value_counts().index]
modes = temp.groupby(["country"]).agg(pd.Series.mode)
dict_locations …Run Code Online (Sandbox Code Playgroud) 我是C的新手,想要编写一个简单的程序,用两个命令行参数调用(第一个是程序的名称,第二个是字符串).我已经验证了参数的数量,现在想验证输入只包含数字.
#include <stdio.h>
#include <string.h>
#include <ctype.h>
#include <cs50.h>
int main(int argc, string argv[])
{
if (argc == 2)
{
for (int i = 0, n = strlen(argv[1]); i < n; i++)
{
if isdigit(argv[1][i])
{
printf("Success!\n");
}
else
{
printf("Error. Second argument can be numbers only\n");
return 1;
}
}
else
{
printf("Error. Second argument can be numbers only\n");
return 1;
}
}
Run Code Online (Sandbox Code Playgroud)
虽然上面的代码没有抛出错误,但它没有正确验证输入,我无法弄清楚原因.有人能指出我正确的方向吗?
非常感谢提前和所有最好的
我data.frame在列中包含国家和城市location,我想通过与world.cities$country.etc来自library(maps)(或任何其他国家/地区名称集合)的数据框匹配来提取前者。
考虑这个例子:
df <- data.frame(location = c("Aarup, Denmark",
"Switzerland",
"Estonia: Aaspere"),
other_col = c(2,3,4))
Run Code Online (Sandbox Code Playgroud)
我尝试使用这段代码
df %>% extract(location,
into = c("country", "rest_location"),
remove = FALSE,
function(x) x[which x %in% world.cities$country.etc])
Run Code Online (Sandbox Code Playgroud)
但我没有成功;我期望这样的事情:
location other_col country rest_location
1 Aarup, Denmark 2 Denmark Aarup,
2 Switzerland 3 Switzerland
3 Estonia: Aaspere 4 Estonia : Aaspere
Run Code Online (Sandbox Code Playgroud) 我有一个带有标签数据的数据集,并且想创建一个仅包含标签作为字符的新列。
考虑以下示例:
value_labels <- tibble(value = 1:6, labels = paste0("value", 1:6))
df_data <- tibble(id = 1:10, var = floor(runif(10, 1, 6)))
df_data <- df_data %>% mutate(var = haven::labelled(var, labels = deframe(value_labels[2:1])))
Run Code Online (Sandbox Code Playgroud)
这产生:
# A tibble: 10 x 2
id var
<int> <dbl+lbl>
1 1 2 [value2]
2 2 2 [value2]
3 3 4 [value4]
4 4 2 [value2]
5 5 4 [value4]
6 6 3 [value3]
7 7 5 [value5]
8 8 4 [value4]
9 9 3 [value3]
10 10 …Run Code Online (Sandbox Code Playgroud) 我需要用 来绘制每日数据sns.lmplot()。
数据具有以下结构:
df = pd.DataFrame(columns=['date', 'origin', 'group', 'value'],
data = [['2001-01-01', "Peter", "A", 1.0],
['2011-01-01', "Peter", "A", 1.1],
['2011-01-02', "Peter", "B", 1.2],
['2012-01-03', "Peter", "A", 1.3],
['2012-01-01', "Peter", "B", 1.4],
['2013-01-02', "Peter", "A", 1.5],
['2013-01-03', "Peter", "B", 1.6],
['2021-01-01', "Peter", "A", 1.7]])
Run Code Online (Sandbox Code Playgroud)
我现在想用每月平均值绘制数据sns.lmplot()(我的原始数据比玩具数据更细粒度)并使用hueforgroup列。为此,我按月汇总:
df['date'] = pd.to_datetime(df['date']).dt.strftime('%Y%M').astype(int)
df = df.groupby(['date', 'origin', 'group']).agg(['mean'])
df.columns = ["_".join(pair) for pair in df.columns] # reset col multi-index
df = df.reset_index() # reset index
Run Code Online (Sandbox Code Playgroud)
然后我绘制数据:
sns.lmplot(data=df, x="date", …Run Code Online (Sandbox Code Playgroud)