我有运行并输出一个大列表的代码。我一直坚持将输出写入文件,因为我不断收到不同的错误,因此我无法以通常用于数据帧的任何方式写入文件。
\n我使用的代码和数据是这样的:
\nlibrary(GeneOverlap)\nlibrary(dplyr)\nlibrary(stringr)\n\ndataset1 <- structure(list(Gene = c("Gene1", "Gene1", "Gene2", "Gene3", "Gene3.", \n"Gene3"), Gene_count = c(5L, 5L, 3L, 16L, 16L, 16L), Phenotype = c("Phenotype1", \n"Phenotype2", "Phenotype1", "Phenotype6", "Phenotype2", "Phenotype1"\n)), row.names = c(NA, -6L), class = c("data.table", "data.frame"\n))\n\n\ndataset2 <- structure(list(Gene = c("Gene1", "Gene1", "Gene4", "Gene2", "Gene6", \n"Gene7"), Gene_count = c(10L, 10L, 4L, 17L, 3L, 2L), Phenotype = c("Phenotype1", \n"Phenotype2", "Phenotype1", "Phenotype6", "Phenotype2", "Phenotype1"\n)), row.names = c(NA, -6L), class = c("data.table", "data.frame"\n))\n\nd1_split <- split(dataset1, dataset1$Phenotype)\nd2_split <- split(dataset2, dataset2$Phenotype)\n\n# this should be …Run Code Online (Sandbox Code Playgroud) 我在 pyspark 中有一个数据集,我为其创建了 row_num 列,因此我的数据如下所示:
#data:
+-----------------+-----------------+-----+------------------+-------+
|F1_imputed |F2_imputed |label| features|row_num|
+-----------------+-----------------+-----+------------------+-------+
| -0.002353| 0.9762| 0|[-0.002353,0.9762]| 1|
| 0.1265| 0.1176| 0| [0.1265,0.1176]| 2|
| -0.08637| 0.06524| 0|[-0.08637,0.06524]| 3|
| -0.1428| 0.4705| 0| [-0.1428,0.4705]| 4|
| -0.1015| 0.6811| 0| [-0.1015,0.6811]| 5|
| -0.01146| 0.8273| 0| [-0.01146,0.8273]| 6|
| 0.0853| 0.2525| 0| [0.0853,0.2525]| 7|
| 0.2186| 0.2725| 0| [0.2186,0.2725]| 8|
| -0.145| 0.3592| 0| [-0.145,0.3592]| 9|
| -0.1176| 0.4225| 0| [-0.1176,0.4225]| 10|
+-----------------+-----------------+-----+------------------+-------+
Run Code Online (Sandbox Code Playgroud)
我正在尝试使用以下方法过滤掉随机选择的行:
count = data.count()
sample = [np.random.choice(np.arange(count), replace=True, size=50)] …Run Code Online (Sandbox Code Playgroud)