Dan*_*ner 5 python scikit-learn tensorflow
这个学期我开始研究机器学习。我们只使用了诸如 Microsoft 的 Azure 和亚马逊的 AWS 之类的 API,但我们还没有深入了解这些服务的工作原理。我的好朋友是数学专业的大四学生,他让我帮助他根据.csv他提供的文件使用 TensorFlow 创建一个股票预测器。
我有几个问题。第一个是他的.csv档案。该文件只有日期和结束值,它们没有分开,因此我不得不手动分开日期和值。我已经设法做到了,现在我在 MinMaxScaler() 上遇到了麻烦。有人告诉我,我几乎可以忽略日期,只测试收盘值,将它们标准化,然后根据它们进行预测。
我不断收到此错误:
ValueError: 找到具有 0 个样本的数组 (shape=(0, 1)) 而 MinMaxScaler() 需要最小值为 1
老实说,我以前从未使用过SKLearning和 TensorFlow,这是我第一次从事这样的项目。我在该主题上看到的所有指南都使用了 Pandas,但就我而言,该.csv文件一团糟,我不相信我可以使用 Pandas。
我正在遵循本指南:
但不幸的是,由于我缺乏经验,有些事情对我来说并不真正有效,我希望能更清楚地了解我应该如何处理我的情况。
下面附上我的(凌乱的)代码:
import pandas as pd
import numpy as np
import tensorflow as tf
import sklearn
from sklearn.model_selection import KFold
from sklearn.preprocessing import scale
from sklearn.preprocessing import MinMaxScaler
import matplotlib
import matplotlib.pyplot as plt
from dateutil.parser import parse
from datetime import datetime, timedelta
from collections import deque
stock_data = []
stock_date = []
stock_value = []
f = open("s&p500closing.csv","r")
data = f.read()
rows = data.split("\n")
rows_noheader = rows[1:len(rows)]
#Separating values from messy `.csv`, putting each value to it's list and also a combined list of both
for row in rows_noheader:
[date, value] = row[1:len(row)-1].split('\t')
stock_date.append(date)
stock_value.append((value))
stock_data.append((date, value))
#Numpy array of all closing values converted to floats and normalized against the maximum
stock_value = np.array(stock_value, dtype=np.float32)
normvalue = [i/max(stock_value) for i in stock_value]
#Number of closing values and days. Since there is one closing value for each, they both match and there are 4528 of them (each)
nclose_and_days = 0
for i in range(len(stock_data)):
nclose_and_days+=1
train_data = stock_value[:2264]
test_data = stock_value[2264:]
scaler = MinMaxScaler()
train_data = train_data.reshape(-1,1)
test_data = test_data.reshape(-1,1)
# Train the Scaler with training data and smooth data
smoothing_window_size = 1100
for di in range(0,4400,smoothing_window_size):
#error occurs here
scaler.fit(train_data[di:di+smoothing_window_size,:])
train_data[di:di+smoothing_window_size,:] = scaler.transform(train_data[di:di+smoothing_window_size,:])
# You normalize the last bit of remaining data
scaler.fit(train_data[di+smoothing_window_size:,:])
train_data[di+smoothing_window_size:,:] = scaler.transform(train_data[di+smoothing_window_size:,:])
# Reshape both train and test data
train_data = train_data.reshape(-1)
# Normalize test data
test_data = scaler.transform(test_data).reshape(-1)
# Now perform exponential moving average smoothing
# So the data will have a smoother curve than the original ragged data
EMA = 0.0
gamma = 0.1
for ti in range(1100):
EMA = gamma*train_data[ti] + (1-gamma)*EMA
train_data[ti] = EMA
# Used for visualization and test purposes
all_mid_data = np.concatenate([train_data,test_data],axis=0)
window_size = 100
N = train_data.size
std_avg_predictions = []
std_avg_x = []
mse_errors = []
for pred_idx in range(window_size,N):
std_avg_predictions.append(np.mean(train_data[pred_idx-window_size:pred_idx]))
mse_errors.append((std_avg_predictions[-1]-train_data[pred_idx])**2)
std_avg_x.append(date)
print('MSE error for standard averaging: %.5f'%(0.5*np.mean(mse_errors)))
Run Code Online (Sandbox Code Playgroud)
我知道这篇文章很旧,但是当我在这里偶然发现时,其他人会..在遇到同样的问题并进行大量谷歌搜索后,我发现了一篇文章 https://github.com/llSourcell/Make_Money_with_Tensorflow_2.0/issues/7
因此,如果您下载的数据集太小,它似乎会抛出该错误。下载 1962 年的 .csv,它就足够大了;)。
现在,我只需要为我的数据集找到正确的参数..因为我正在将其适应另一种类型的预测..希望它有帮助
小智 1
您的问题不是 CSV 或 pandas。实际上,您可以将带有 pandas 的 CSV 直接读取到数据框中,这是我建议您这样做的。df = pd.read_csv(path)
我使用相同的代码遇到同样的问题。发生的事情是Scaler = MinMaxScaler, then 在 for di in range 部分中,您正在将数据拟合到训练集中,然后对其进行转换并将其重新分配回自身。
问题是,它试图在训练集中找到更多数据以适合缩放器,但数据耗尽了。这很奇怪,因为您所遵循的教程呈现它的方式。