You are viewing a plain text version of this content. The canonical link for it is here.
Posted to gitbox@hive.apache.org by "simhadri-g (via GitHub)" <gi...@apache.org> on 2023/03/23 12:23:45 UTC

[GitHub] [hive] simhadri-g commented on a diff in pull request #4131: HIVE-27158: Store hive columns stats in puffin files for iceberg tables

simhadri-g commented on code in PR #4131:
URL: https://github.com/apache/hive/pull/4131#discussion_r1146105589


##########
iceberg/iceberg-handler/src/main/java/org/apache/iceberg/mr/hive/HiveIcebergStorageHandler.java:
##########
@@ -349,6 +364,98 @@ public Map<String, String> getBasicStatistics(Partish partish) {
     return stats;
   }
 
+
+  @Override
+  public boolean canSetColStatistics() {
+    String statsSource = HiveConf.getVar(conf, HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+    return statsSource.equals(PUFFIN) ? true : false;
+  }
+
+  @Override
+  public boolean canProvideColStatistics(org.apache.hadoop.hive.ql.metadata.Table tbl) {
+
+    org.apache.hadoop.hive.ql.metadata.Table hmsTable = tbl;
+    TableDesc tableDesc = Utilities.getTableDesc(hmsTable);
+    Table table = Catalogs.loadTable(conf, tableDesc.getProperties());
+    if (table.currentSnapshot() != null) {
+      String statsSource = HiveConf.getVar(conf, HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+      String statsPath = table.location() + "/stats/" + table.name() + table.currentSnapshot().snapshotId();
+      if (statsSource.equals(PUFFIN)) {
+        try (FileSystem fs = new Path(table.location()).getFileSystem(conf)) {
+          if (fs.exists(new Path(statsPath))) {
+            return true;
+          }
+        } catch (IOException e) {
+          LOG.warn(e.getMessage());
+        }
+      }
+    }
+    return false;
+  }
+
+  @Override
+  public List<ColumnStatisticsObj> getColStatistics(org.apache.hadoop.hive.ql.metadata.Table tbl) {
+
+    org.apache.hadoop.hive.ql.metadata.Table hmsTable = tbl;
+    TableDesc tableDesc = Utilities.getTableDesc(hmsTable);
+    Table table = Catalogs.loadTable(conf, tableDesc.getProperties());
+    String statsSource = HiveConf.getVar(conf, HiveConf.ConfVars.HIVE_USE_STATS_FROM).toLowerCase();
+    switch (statsSource) {
+      case ICEBERG:
+        // Place holder for iceberg stats
+        break;
+      case PUFFIN:
+        String snapshotId = table.name() + table.currentSnapshot().snapshotId();

Review Comment:
   Hi Rajesh, 
   Thanks for the review! :) 
   
   Currently, the scope of this patch is for column stats. We can include basic stats in this puffin file as well if needed.
   
   But currently, basics stats for iceberg tables are obtained from `table.currentSnapshot().summary() `  here : https://github.com/apache/hive/blob/master/iceberg/iceberg-handler/src/main/java/org/apache/iceberg/mr/hive/HiveIcebergStorageHandler.java#L316 .
   I do not think we are querying the metastore for basic stats. Please correct me if i am missing something.
   
   
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: gitbox-unsubscribe@hive.apache.org

For queries about this service, please contact Infrastructure at:
users@infra.apache.org


---------------------------------------------------------------------
To unsubscribe, e-mail: gitbox-unsubscribe@hive.apache.org
For additional commands, e-mail: gitbox-help@hive.apache.org