You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@hbase.apache.org by "Liyin Tang (JIRA)" <ji...@apache.org> on 2016/03/18 05:19:33 UTC
[jira] [Created] (HBASE-15482) Provide an option to skip
calculating block locations for SnapshotInputFormat
Liyin Tang created HBASE-15482:
----------------------------------
Summary: Provide an option to skip calculating block locations for SnapshotInputFormat
Key: HBASE-15482
URL: https://issues.apache.org/jira/browse/HBASE-15482
Project: HBase
Issue Type: Improvement
Components: mapreduce
Reporter: Liyin Tang
Priority: Minor
When a MR job is reading from SnapshotInputFormat, it needs to calculate the splits based on the block locations in order to get best locality. However, this process may take a long time for large snapshots.
In some setup, the computing layer, Spark, Hive or Presto could run out side of HBase cluster. In these scenarios, the block locality doesn't matter. Therefore, it will be great to have an option to skip calculating the block locations for every job. That will super useful for the Hive/Presto/Spark connectors.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)