You are viewing a plain text version of this content. The canonical link for it is here.
Posted to issues@hive.apache.org by "mahesh kumar behera (Jira)" <ji...@apache.org> on 2022/01/20 04:57:00 UTC

[jira] [Resolved] (HIVE-25638) Select returns deleted records in Hive ACID table

     [ https://issues.apache.org/jira/browse/HIVE-25638?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]

mahesh kumar behera resolved HIVE-25638.
----------------------------------------
    Resolution: Fixed

> Select returns deleted records in Hive ACID table
> -------------------------------------------------
>
>                 Key: HIVE-25638
>                 URL: https://issues.apache.org/jira/browse/HIVE-25638
>             Project: Hive
>          Issue Type: Bug
>          Components: HiveServer2
>            Reporter: mahesh kumar behera
>            Assignee: mahesh kumar behera
>            Priority: Major
>              Labels: pull-request-available
>          Time Spent: 20m
>  Remaining Estimate: 0h
>
> Hive stores the stripe stats in the ORC files. During select, these stats are used to create the SARG. The SARG is used to reduce the records read from the delete-delta files. Currently, in case where the number of stripes are more than 1, the SARG generated is not proper as it uses the first stripe index for both min and max key interval. The max key interval should be obtained from last stripe index. This cause some valid deleted records to be skipped. And those records are return to the user. We need the last stripe here instead of the first one, is the fact the keys are ordered in the file.



--
This message was sent by Atlassian Jira
(v8.20.1#820001)