You are viewing a plain text version of this content. The canonical link for it is here.
Posted to reviews@spark.apache.org by vanzin <gi...@git.apache.org> on 2018/04/25 22:42:55 UTC

[GitHub] spark pull request #21066: [SPARK-23977][CLOUD][WIP] Add commit protocol bin...

Github user vanzin commented on a diff in the pull request:

    https://github.com/apache/spark/pull/21066#discussion_r184223981
  
    --- Diff: hadoop-cloud/src/hadoop-3/main/scala/org/apache/spark/internal/io/cloud/PathOutputCommitProtocol.scala ---
    @@ -0,0 +1,260 @@
    +/*
    + * Licensed to the Apache Software Foundation (ASF) under one or more
    + * contributor license agreements.  See the NOTICE file distributed with
    + * this work for additional information regarding copyright ownership.
    + * The ASF licenses this file to You under the Apache License, Version 2.0
    + * (the "License"); you may not use this file except in compliance with
    + * the License.  You may obtain a copy of the License at
    + *
    + *    http://www.apache.org/licenses/LICENSE-2.0
    + *
    + * Unless required by applicable law or agreed to in writing, software
    + * distributed under the License is distributed on an "AS IS" BASIS,
    + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
    + * See the License for the specific language governing permissions and
    + * limitations under the License.
    + */
    +
    +package org.apache.spark.internal.io.cloud
    +
    +import java.io.IOException
    +
    +import org.apache.hadoop.fs.Path
    +import org.apache.hadoop.mapreduce.{JobContext, TaskAttemptContext}
    +import org.apache.hadoop.mapreduce.lib.output.{FileOutputCommitter, PathOutputCommitter, PathOutputCommitterFactory}
    +
    +import org.apache.spark.internal.io.{FileCommitProtocol, HadoopMapReduceCommitProtocol}
    +import org.apache.spark.internal.io.FileCommitProtocol.TaskCommitMessage
    +
    +/**
    + * Spark Commit protocol for Path Output Committers.
    + * This committer will work with the `FileOutputCommitter` and subclasses.
    + * All implementations *must* be serializable.
    + *
    + * Rather than ask the `FileOutputFormat` for a committer, it uses the
    + * `org.apache.hadoop.mapreduce.lib.output.PathOutputCommitterFactory` factory
    + * API to create the committer.
    + * This is what [[org.apache.hadoop.mapreduce.lib.output.FileOutputFormat]] does,
    + * but as [[HadoopMapReduceCommitProtocol]] still uses the original
    + * `org.apache.hadoop.mapred.FileOutputFormat` binding
    + * subclasses do not do this, overrides those subclasses to using the
    + * factory mechanism now supported in the base class.
    + *
    + * In `setupCommitter` the factory is bonded to and the committer for
    + * the destination path chosen.
    + *
    + * @constructor Instantiate. dynamic partition overwrite is not supported,
    + *              so that committers for stores which do not support rename
    + *              will not get confused.
    + * @param jobId                     job
    + * @param destination               destination
    + * @param dynamicPartitionOverwrite does the caller want support for dynamic
    + *                                  partition overwrite. If so, it will be
    + *                                  refused.
    + * @throws IOException when an unsupported dynamicPartitionOverwrite option is supplied.
    + */
    +class PathOutputCommitProtocol(
    +  jobId: String,
    +  destination: String,
    +  dynamicPartitionOverwrite: Boolean = false)
    +  extends HadoopMapReduceCommitProtocol(
    +    jobId,
    +    destination,
    +    false) with Serializable {
    +
    +  @transient var committer: PathOutputCommitter = _
    --- End diff --
    
    `private`?


---

---------------------------------------------------------------------
To unsubscribe, e-mail: reviews-unsubscribe@spark.apache.org
For additional commands, e-mail: reviews-help@spark.apache.org