Hadoop:提供目录作为MapReduce作业的输入

sgo*_*les 8 java hadoop mapreduce input cloudera

我正在使用Cloudera Hadoop.我能够运行简单的mapreduce程序,我提供了一个文件作为MapReduce程序的输入.

此文件包含mapper函数要处理的所有其他文件.

但是,我陷入了困境.

/folder1
  - file1.txt
  - file2.txt
  - file3.txt
Run Code Online (Sandbox Code Playgroud)

如何指定MapReduce程序的输入路径"/folder1",以便它可以开始处理该目录中的每个文件?

有任何想法吗 ?

编辑:

1)Intiailly,我提供了inputFile.txt作为mapreduce程序的输入.它工作得很好.

>inputFile.txt
file1.txt
file2.txt
file3.txt
Run Code Online (Sandbox Code Playgroud)

2)但是现在,我想在命令行上提供一个输入目录作为arg [0],而不是给出一个输入文件.

hadoop jar ABC.jar /folder1 /output
Run Code Online (Sandbox Code Playgroud)

sha*_*ovo 13

问题是FileInputFormat不会在输入路径dir中递归读取文件.

解决方案: 使用以下代码

FileInputFormat.setInputDirRecursive(job, true); 在Map Reduce Code下面的行之前

FileInputFormat.addInputPath(job, new Path(args[0]));

您可以在此处查看修复的版本.


zhu*_*ala 2

您可以使用FileSystem.listStatus从给定目录获取文件列表,代码如下:

//get the FileSystem, you will need to initialize it properly
FileSystem fs= FileSystem.get(conf); 
//get the FileStatus list from given dir
FileStatus[] status_list = fs.listStatus(new Path(args[0]));
if(status_list != null){
    for(FileStatus status : status_list){
        //add each file to the list of inputs for the map-reduce job
        FileInputFormat.addInputPath(conf, status.getPath());
    }
}
Run Code Online (Sandbox Code Playgroud)