小编Hac*_*ore的帖子

熊猫describe()未显示

我正在学习有关机器学习的Google课程,并试图使其在Atom上运行,而不是简单地使用colab版本。模型训练和其他工作进展顺利,但使用describe()函数时遇到问题。我已经查阅了文档,但是仍然无法显示摘要。它仅在我在命令行上尝试交互式python时有效。我的代码的相关部分如下。谢谢您的帮助。

import math

from IPython import display
from matplotlib import cm
from matplotlib import gridspec
from matplotlib import pyplot as plt
import numpy as np
import pandas as pd
from sklearn import metrics
import tensorflow as tf
from tensorflow.python.data import Dataset

tf.logging.set_verbosity(tf.logging.ERROR)
pd.options.display.max_rows = 10
pd.options.display.float_format = '{:.lf}'.format

# Load data set
california_housing_dataframe = pd.read_csv("https://storage.googleapis.com/mledu-datasets/california_housing_train.csv", sep=",")

......

# Split the data set into training sets of the first 12000/17000 examples,
training_examples = preprocess_features(california_housing_dataframe.head(12000))
training_targets = preprocess_targets(california_housing_dataframe.head(12000))

# and validation sets …
Run Code Online (Sandbox Code Playgroud)

python machine-learning pandas

1
推荐指数
1
解决办法
1502
查看次数

如何从具有不同长度的列表列表中创建 Pandas DataFrame?

我的数据格式如下

data = [["a", "b", "c"],
        ["b", "c"],
        ["d", "e", "f", "c"]]
Run Code Online (Sandbox Code Playgroud)

并且我想要一个 DataFrame 将所有唯一的字符串作为列和出现的二进制值

    a  b  c  d  e  f
0   1  1  1  0  0  0
1   0  1  1  0  0  0
2   0  0  1  1  1  1
Run Code Online (Sandbox Code Playgroud)

我有一个使用列表推导式的工作代码,但对于大数据来说速度很慢。

# vocab_list contains all the unique keys, which is obtained when reading in data from file
df = pd.DataFrame([[1 if word in entry else 0 for word in vocab_list] for entry in data])
Run Code Online (Sandbox Code Playgroud)

有没有办法优化这个任务?谢谢。

编辑(实际数据的小样本):

[['a', 'about', …

python dataframe pandas

1
推荐指数
1
解决办法
300
查看次数

标签 统计

pandas ×2

python ×2

dataframe ×1

machine-learning ×1