You are viewing a plain text version of this content. The canonical link for it is here.
Posted to mapreduce-issues@hadoop.apache.org by "Bikas Saha (JIRA)" <ji...@apache.org> on 2012/12/19 20:59:12 UTC
[jira] [Created] (MAPREDUCE-4892) CombineFileInputFormat node input
split can be skewed on small clusters
Bikas Saha created MAPREDUCE-4892:
-------------------------------------
Summary: CombineFileInputFormat node input split can be skewed on small clusters
Key: MAPREDUCE-4892
URL: https://issues.apache.org/jira/browse/MAPREDUCE-4892
Project: Hadoop Map/Reduce
Issue Type: Bug
Reporter: Bikas Saha
Assignee: Bikas Saha
Fix For: 3.0.0
The CombineFileInputFormat split generation logic tries to group blocks by node in order to create splits. It iterates through the nodes and creates splits on them until there aren't enough blocks left on a node that can be grouped into a valid split. If the first few nodes have a lot of blocks on them then they can end up getting a disproportionately large share of the total number of splits created. This can result in poor locality of maps. This problem is likely to happen on small clusters where its easier to create a skew in the distribution of blocks on nodes.
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira