You are viewing a plain text version of this content. The canonical link for it is here.
Posted to issues@beam.apache.org by "Alan Myrvold (Jira)" <ji...@apache.org> on 2019/11/26 21:13:00 UTC

[jira] [Commented] (BEAM-8832) Slow GCS uploads for large Java staging jars

    [ https://issues.apache.org/jira/browse/BEAM-8832?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16982903#comment-16982903 ] 

Alan Myrvold commented on BEAM-8832:
------------------------------------

Ran test upload of ./gradlew :runners:google-cloud-dataflow-java:GCSUpload for a 100M file, 10 times for various sizes.

(using current 1M default and max)
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 16 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 14 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 16 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 21 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 15 seconds

(using 8M chunk size)
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 7 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds

(using 16M chunk size)
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 7 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 5 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds

(using 64m chunk size)
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 3 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 3 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 3 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 3 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 3 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds
INFO: Staging files complete: 0 files cached, 1 files newly uploaded in 4 seconds

> Slow GCS uploads for large Java staging jars
> --------------------------------------------
>
>                 Key: BEAM-8832
>                 URL: https://issues.apache.org/jira/browse/BEAM-8832
>             Project: Beam
>          Issue Type: Improvement
>          Components: runner-dataflow
>    Affects Versions: 2.16.0
>            Reporter: Alan Myrvold
>            Assignee: Alan Myrvold
>            Priority: Minor
>          Time Spent: 10m
>  Remaining Estimate: 0h
>
> The default and max upload chunk size is 1M, in 
> [https://github.com/apache/beam/blob/master/runners/google-cloud-dataflow-java/src/main/java/org/apache/beam/runners/dataflow/util/GcsStager.java#L96]
> This should be at least 8M to improve performance.
> 8m is the default python write chunk size
> [https://github.com/apache/beam/blob/master/sdks/python/apache_beam/io/gcp/gcsio.py#L86]
>  



--
This message was sent by Atlassian Jira
(v8.3.4#803005)