我想用NaN替换数据帧列中的错误值.
mydata = {'x' : [10, 50, 18, 32, 47, 20], 'y' : ['12', '11', 'N/A', '13', '15', 'N/A']}
df = pd.DataFrame(mydata)
df[df.y == 'N/A']['y'] = np.nan
Run Code Online (Sandbox Code Playgroud)
虽然,最后一行失败并抛出警告,因为它正在处理df的副本.那么,处理这个问题的正确方法是什么?我已经看过许多使用iloc或ix的解决方案但是在这里,我需要使用布尔条件.
我想将给定自定距离的点聚类,奇怪的是,似乎scipy和sklearn聚类方法都不允许指定距离函数.
例如,sklearn.cluster.AgglomerativeClustering我唯一能做的就是输入一个亲和力矩阵(这将是一个非常大的内存).为了构建这个矩阵,建议使用sklearn.neighbors.kneighbors_graph,但我不明白如何在两点之间指定距离函数.有人可以开导我吗?
我正在编写一个存储a的模板类,std::function以便稍后调用它.这是简化的代码:
template <typename T>
struct Test
{
void call(T type)
{
function(type);
}
std::function<void(T)> function;
};
Run Code Online (Sandbox Code Playgroud)
问题是该模板不能为该void类型编译,因为
void call(void type)
Run Code Online (Sandbox Code Playgroud)
变得不确定.
专门针对该void类型并不能解决问题,因为
template <>
void Test<void>::call(void)
{
function();
}
Run Code Online (Sandbox Code Playgroud)
仍然与宣言不符call(T Type).
因此,使用C++ 11的新功能,我试过std::enable_if:
typename std::enable_if_t<std::is_void_v<T>, void> call()
{
function();
}
typename std::enable_if_t<!std::is_void_v<T>, void> call(T type)
{
function(type);
}
Run Code Online (Sandbox Code Playgroud)
但它不能用Visual Studio编译:
错误C2039:'type':不是'std :: enable_if'的成员
你会如何解决这个问题?
我通过运行“pip install theano”在 Windows 7 64 位上安装了带有 Spyder 2.3.8 的 theano。它运作良好。但是,当我尝试运行“import theano”时,出现以下错误:
回溯(最近一次调用最后一次):
文件“”,第 1 行,在
进口theano
文件“D:\Anaconda3\lib\site-packages\theano\__init__.py”,第 55 行,在
从 theano.compile 导入 \
文件“D:\Anaconda3\lib\site-packages\theano\compile\__init__.py”,第 9 行,在
从 theano.compile.function_module 导入 *
文件“D:\Anaconda3\lib\site-packages\theano\compile\function_module.py”,第 18 行,在
导入 theano.compile.mode
文件“D:\Anaconda3\lib\site-packages\theano\compile\mode.py”,第 11 行,在
导入 theano.gof.vm
文件“D:\Anaconda3\lib\site-packages\theano\gof\vm.py”,第 25 行,在
in_c_key=假)
AddConfigVar 中的文件“D:\Anaconda3\lib\site-packages\theano\configparser.py”,第 231 行
configparam.fullname)
AttributeError: ('这个名字已经被占用', 'profile')
这意味着什么?
我正在用 C 编写 Python 扩展。我的函数将 Numpy 数组(包含整数)列表作为参数。我遍历列表,获取每个数组,增加引用,获取指向数组的 C 指针,然后减少引用。
if (!PyArg_ParseTuple(args, "O", &list))
return NULL;
long nb_arrays = PyList_Size(list);
arrays = (int **) malloc(nb_arrays * sizeof(int *));
for (i = 0; i < nb_arrays; i++)
{
PyArrayObject *array = (PyArrayObject *) PyList_GetItem(list, i);
Py_INCREF(array);
arrays[i] = (int *) PyArray_DATA(array);
Py_DECREF(array);
}
Run Code Online (Sandbox Code Playgroud)
在这个循环之后,我使用指针进行计算。是正确的还是我必须等待函数结束才能减少引用计数?
我正在尝试使用Numpy genfromtxt导入一个简单的CSV文件,但无法将第一列的数据转换为日期。
这是我的代码:
import numpy as np
from datetime import datetime
str2date = lambda x: datetime.strptime(x, '%Y-%m-%d %H:%M:%S')
data = np.genfromtxt('C:\\\\data.csv',dtype=None,names=True, delimiter=',', converters = {0: str2date})
Run Code Online (Sandbox Code Playgroud)
我在str2date中收到以下错误:
TypeError:必须为str,而不是字节
问题是有很多列,所以我宁愿避免指定所有列类型(基本上是数字)。
我用C语言编写了一个简单的Python扩展函数,它仅读取一个Numpy数组,然后崩溃。
static PyObject *test(PyObject *self, PyObject *args)
{
PyArrayObject *array = NULL;
if (!PyArg_ParseTuple(args, "O!", &PyArray_Type, &array)) // Crash
return NULL;
return Py_BuildValue("d", 0);
}
Run Code Online (Sandbox Code Playgroud)
这就是它的称呼:
l = np.array([1,2,3,1,2,2,1,3])
print("%d" % extension.test(l))
Run Code Online (Sandbox Code Playgroud)
我的代码有什么问题?
python ×6
numpy ×3
c ×2
arrays ×1
c++ ×1
c++11 ×1
genfromtxt ×1
installation ×1
nan ×1
pandas ×1
scikit-learn ×1
scipy ×1
sfinae ×1
templates ×1
theano ×1