我们需要比较两个CSV文件.假设文件一行有几行,第二个文件可以有相同的行数或更多行.大多数行可以在两个文件上保持相同.寻找在这两个文件之间进行差异的最佳方法,并只读取第二个文件与第一个文件有差异的那些行.处理文件的应用程序是Java.
有什么最好的方法?
注意:如果我们知道在第二个文件中更新,插入或删除了一行,那将会很棒.
要求:-
这样做的一种方法是使用java的Set接口; 将每一行作为字符串读取,将其添加到集合中,然后removeAll()在第一个集合上执行第二个集合,从而保留不同的行.当然,这假定文件中没有重复的行.
// using FileUtils to read in the files.
HashSet<String> f1 = new HashSet<String>(FileUtils.readLines("file1.csv"));
HashSet<String> f2 = new HashSet<String>(FileUtils.readLines("file2.csv"));
f1.removeAll(f2); // f1 now contains only the lines which are not in f2
Run Code Online (Sandbox Code Playgroud)
更新
好的,所以你有一个PK领域.我只是假设你知道如何从你的字符串中得到它; 使用openCSV或正则表达式或任何你想要的.制作一个实际HashMap而不是HashSet如上所述,使用PK作为键,行作为值.
HashMap<String, String> f1 = new HashMap<String, String>();
HashMap<String, String> f2 = new HashMap<String, String>();
// read f1, f2; use PK field as the key
List<String> deleted = new ArrayList<String>();
List<String> updated = new ArrayList<String>();
for(Map.Entry<String, String> entry : f1.keySet()) {
if(!f2.containsKey(entry.getKey()) {
deleted.add(entry.getValue());
} else {
if(!f2.get(entry.getKey().equals(f1.getValue())) {
updated.add(f1.getValue());
}
}
}
for(String key : f1.keySet()) {
f2.remove(key);
}
// f2 now contains only "new" rows
Run Code Online (Sandbox Code Playgroud)
读取整个第一个文件,并将其放入List. 然后一次读取第二个文件,并将每一行与第一个文件的所有行进行比较,看它是否重复。如果它不是重复的,那么它就是新信息。如果您在阅读时遇到问题,请查看http://opencsv.sourceforge.net/,这是一个非常好的 Java 读取 CSV 文件库。