小编LN3*_*LN3的帖子

如何将包含 S4 对象的大型列表编写为 CSV 文件?

我有运行并输出一个大列表的代码。我一直坚持将输出写入文件,因为我不断收到不同的错误,因此我无法以通常用于数据帧的任何方式写入文件。

\n

我使用的代码和数据是这样的:

\n
library(GeneOverlap)\nlibrary(dplyr)\nlibrary(stringr)\n\ndataset1 <- structure(list(Gene = c("Gene1", "Gene1", "Gene2", "Gene3", "Gene3.", \n"Gene3"), Gene_count = c(5L, 5L, 3L, 16L, 16L, 16L), Phenotype = c("Phenotype1", \n"Phenotype2", "Phenotype1", "Phenotype6", "Phenotype2", "Phenotype1"\n)), row.names = c(NA, -6L), class = c("data.table", "data.frame"\n))\n\n\ndataset2 <- structure(list(Gene = c("Gene1", "Gene1", "Gene4", "Gene2", "Gene6", \n"Gene7"), Gene_count = c(10L, 10L, 4L, 17L, 3L, 2L), Phenotype = c("Phenotype1", \n"Phenotype2", "Phenotype1", "Phenotype6", "Phenotype2", "Phenotype1"\n)), row.names = c(NA, -6L), class = c("data.table", "data.frame"\n))\n\nd1_split <- split(dataset1, dataset1$Phenotype)\nd2_split <- split(dataset2, dataset2$Phenotype)\n\n# this should be …
Run Code Online (Sandbox Code Playgroud)

r bioinformatics r-s4

2
推荐指数
1
解决办法
679
查看次数

AttributeError:“numpy.int64”对象没有属性“_get_object_id”

我在 pyspark 中有一个数据集,我为其创建了 row_num 列,因此我的数据如下所示:

#data:
+-----------------+-----------------+-----+------------------+-------+
|F1_imputed       |F2_imputed       |label|          features|row_num|
+-----------------+-----------------+-----+------------------+-------+
|        -0.002353|           0.9762|    0|[-0.002353,0.9762]|      1|
|           0.1265|           0.1176|    0|   [0.1265,0.1176]|      2|
|         -0.08637|          0.06524|    0|[-0.08637,0.06524]|      3|
|          -0.1428|           0.4705|    0|  [-0.1428,0.4705]|      4|
|          -0.1015|           0.6811|    0|  [-0.1015,0.6811]|      5|
|         -0.01146|           0.8273|    0| [-0.01146,0.8273]|      6|
|           0.0853|           0.2525|    0|   [0.0853,0.2525]|      7|
|           0.2186|           0.2725|    0|   [0.2186,0.2725]|      8|
|           -0.145|           0.3592|    0|   [-0.145,0.3592]|      9|
|          -0.1176|           0.4225|    0|  [-0.1176,0.4225]|     10|
+-----------------+-----------------+-----+------------------+-------+
Run Code Online (Sandbox Code Playgroud)

我正在尝试使用以下方法过滤掉随机选择的行:

count = data.count()
sample = [np.random.choice(np.arange(count), replace=True, size=50)] …
Run Code Online (Sandbox Code Playgroud)

python numpy apache-spark pyspark

1
推荐指数
1
解决办法
1万
查看次数

标签 统计

apache-spark ×1

bioinformatics ×1

numpy ×1

pyspark ×1

python ×1

r ×1

r-s4 ×1