我想置换一个向量,这样一个元素在排列之后就不能像在原始元素中那样在同一个地方.假设我有一个像这样的元素列表:AABBCCADEF
有效的洗牌是:BBAADEFCCA
但这些都是无效的:B A ACFEDCAB或BCA B FEDCAB
我能找到的最接近的答案是:python shuffle这样的位置永远不会重复.但这不是我想要的,因为在这个例子中没有重复的元素.
我想要一个快速算法,在重复的情况下推广该答案.
MWE:
library(microbenchmark)
set.seed(1)
x <- sample(letters, size=295, replace=T)
terrible_implementation <- function(x) {
xnew <- sample(x)
while(any(x == xnew)) {
xnew <- sample(x)
}
return(xnew)
}
microbenchmark(terrible_implementation(x), times=10)
Unit: milliseconds
expr min lq mean median uq max neval
terrible_implementation(x) 479.5338 2346.002 4738.49 2993.29 4858.254 17005.05 10
Run Code Online (Sandbox Code Playgroud)
另外,如何确定是否可以以这种方式置换序列?
编辑:为了清楚地说明我想要的东西,新的载体应满足以下条件:
1)all(table(newx) == table(x))
2)all(x != newx)
例如:
newx <- terrible_implementation(x)
all(table(newx) == table(x))
[1] TRUE
all(x != newx)
[1] …Run Code Online (Sandbox Code Playgroud) 我受到了fst包的启发,试图编写一个C++函数来快速序列化我在R到磁盘中的一些数据结构.
但即使在非常简单的对象上,我也无法达到相同的写入速度.下面的代码是一个将大的1 GB向量写入磁盘的简单示例.
使用自定义C++代码,我实现了135 MB/s的写入速度,这是我根据CrystalBench的磁盘限制.
在相同的数据上,write_fst实现了223 MB/s的写入速度,这似乎是不可能的,因为我的磁盘无法快速写入.(注意,我正在使用fst::threads_fst(1)和compress=0设置,并且文件具有相同的数据大小.)
我错过了什么?
如何让C++函数更快地写入磁盘?
C++代码:
#include <Rcpp.h>
#include <fstream>
#include <cstring>
#include <iostream>
// [[Rcpp::plugins(cpp11)]]
using namespace Rcpp;
// [[Rcpp::export]]
void test(SEXP x) {
char* d = reinterpret_cast<char*>(REAL(x));
long dl = Rf_xlength(x) * 8;
std::ofstream OutFile;
OutFile.open("/tmp/test.raw", std::ios::out | std::ios::binary);
OutFile.write(d, dl);
OutFile.close();
}
Run Code Online (Sandbox Code Playgroud)
R代码:
library(microbenchmark)
library(Rcpp)
library(dplyr)
library(fst)
fst::threads_fst(1)
sourceCpp("test.cpp")
x <- runif(134217728) # 1 gigabyte
df <- data.frame(x)
microbenchmark(test(x), write_fst(df, "/tmp/test.fst", compress=0), times=3)
Unit: seconds …Run Code Online (Sandbox Code Playgroud) 我有一个很大的原始向量,例如:
x <- rep(as.raw(1:10), 4e8) # this vector is about 4 GB
Run Code Online (Sandbox Code Playgroud)
我只想删除第一个元素,但是无论我做什么,它都会占用大量内存。
> x <- tail(x, length(x)-1)
Error: cannot allocate vector of size 29.8 Gb
> x <- x[-1L]
Error: cannot allocate vector of size 29.8 Gb
> x <- x[seq(2, length(x)-1)]
Error: cannot allocate vector of size 29.8 Gb
Run Code Online (Sandbox Code Playgroud)
这是怎么回事?我真的必须依靠C来进行如此简单的操作吗?(我知道使用Rcpp很简单,但这不是重点)。
SessionInfo:
R version 3.6.1 (2019-07-05)
Platform: x86_64-pc-linux-gnu (64-bit)
Running under: Ubuntu 16.04.6 LTS
Matrix products: default
BLAS: /usr/lib/libblas/libblas.so.3.6.0
LAPACK: /usr/lib/lapack/liblapack.so.3.6.0
locale:
[1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
[3] LC_TIME=en_US.UTF-8 LC_COLLATE=en_US.UTF-8
[5] …Run Code Online (Sandbox Code Playgroud) 不起作用:
mydat <- data.frame(`Col 1`=1:5, `Col 2`=1:5, check.names=F)
xcol <- "Col 1"
ycol <- "Col 2"
ggplot(data=mydat, aes_string(x=xcol, y=ycol)) + geom_point()
Run Code Online (Sandbox Code Playgroud)
作品:
mydat <- data.frame(`A`=1:5, `B`=1:5)
xcol <- "A"
ycol <- "B"
ggplot(data=mydat, aes_string(x=xcol, y=ycol)) + geom_point()
Run Code Online (Sandbox Code Playgroud)
作品。
mydat <- data.frame(`Col 1`=1:5, `Col 2`=1:5, check.names=F)
ggplot(data=mydat, aes(x=`Col 1`, y=`Col 2`)) + geom_point()
Run Code Online (Sandbox Code Playgroud)
有什么问题?
我继续读到这样做是不好的,但我不认为这些答案完全回答了我的具体问题.
在某些情况下,它似乎真的很有用.我想做类似以下的事情:
class Example {
private:
int val;
public:
void my_function() {
#if defined (__AVX2__)
#include <function_internal_code_using_avx2.h>
#else
#include <function_internal_code_without_avx2.h>
#endif
}
};
Run Code Online (Sandbox Code Playgroud)
如果#include在这个例子中使用代码中间是不好的,那么实现我想要做的事情的好方法是什么?也就是说,我试图在avx2可用和不可编译的情况下区分成员函数实现.
我正在尝试在包中使用外部指针,但遇到一个问题,即好像没有调用终结器并且内存泄漏。
下面是一个非常人为的问题示例:
#include <Rcpp.h>
using namespace Rcpp;
void finalize(SEXP xp){
delete static_cast< std::vector<double> *>(R_ExternalPtrAddr(xp));
}
// [[Rcpp::export]]
SEXP ext_ref_ex() {
std::vector<double> * x = new std::vector<double>(1000000);
SEXP xp = PROTECT(R_MakeExternalPtr(x, R_NilValue, R_NilValue));
R_RegisterCFinalizer(xp, finalize);
UNPROTECT(1);
return xp;
}
Run Code Online (Sandbox Code Playgroud)
R面:
library(Rcpp)
sourceCpp("tests.cpp")
# breaks and/or crashes
for(i in 1:10000) {
z <- ext_ref_ex()
}
# no issue
for(i in 1:10000) {
z <- ext_ref_ex()
rm(z)
gc()
}
Run Code Online (Sandbox Code Playgroud)
运行第一个循环,R会在进行足够的迭代(问题1)后最终发生段错误,而预期的行为是应清理数据且不应出现段错误。
问题2是,如果您中断进程并调用gc(),有时内存将被清除,但通常不会清除。根据该htop报告,即使在rm(list=ls())和之后,R也会使用60-70%的内存gc()。
第二个循环没有明显的内存问题。
我在C端做错了吗?我遇到错误了吗?
(Windows上的R版本3.5.2 ubuntu …
我正在尝试编写一个函数来计算字符串向量中唯一项的数量(我的问题稍微复杂一点,但这是可重现的.我根据我在C++中找到的答案做了这个.这是我的代码:
C++
int unique_sort(vector<string> x) {
sort(x.begin(), x.end());
return unique(x.begin(), x.end()) - x.begin();
}
int unique_set(vector<string> x) {
unordered_set<string> tab(x.begin(), x.end());
return tab.size();
}
Run Code Online (Sandbox Code Playgroud)
R:
x <- paste0("x", sample(1:1e5, 1e7, replace=T))
microbenchmark(length(unique(x)),unique_sort(x), unique_set(x), times=3)
Run Code Online (Sandbox Code Playgroud)
结果:
Unit: milliseconds
expr min lq mean median uq
length(unique(x)) 365.0213 373.4018 406.0209 381.7823 426.5206
unique_sort(x) 10732.1918 10847.0532 10907.6882 10961.9146 10995.4363
unique_set(x) 1948.6517 2230.3383 2334.4040 2512.0249 2527.2802
Run Code Online (Sandbox Code Playgroud)
查看该unique函数的R源代码(它有点难以理解),它似乎在数组上使用循环向哈希添加唯一元素,并检查该哈希是否已存在元素.
因此,我认为它应该等同于unordered_set方法.我不明白为什么unordered_set方法慢了5倍.
TLDR:为什么我的C++代码变慢?
在最新版本的 boost 中,定义了 4 个字节序宏:
* `BOOST_ENDIAN_BIG_BYTE`, byte-swapped big-endian.
* `BOOST_ENDIAN_BIG_WORD`, word-swapped big-endian.
* `BOOST_ENDIAN_LITTLE_BYTE`, byte-swapped little-endian.
* `BOOST_ENDIAN_LITTLE_WORD`, word-swapped little-endian.
Run Code Online (Sandbox Code Playgroud)
https://www.boost.org/doc/libs/1_69_0/boost/predef/other/endian.h
我不清楚_BYTE和_WORD宏之间的区别。