You are viewing a plain text version of this content. The canonical link for it is here.
Posted to github@arrow.apache.org by GitBox <gi...@apache.org> on 2022/06/19 00:59:26 UTC
[GitHub] [arrow-rs] saethlin opened a new pull request, #1906: Fix misaligned reference and logic error in crc32
saethlin opened a new pull request, #1906:
URL: https://github.com/apache/arrow-rs/pull/1906
Previously, this code tried to turn a &[u8] into a &[u32] without
checking alignment. This means it could and did create misaligned
references, which is UB. This can be detected by running the tests with
-Zbuild-std --target=x86_64-unknown-linux-gnu (or whatever your host
is). This change adopts the approach from the murmurhash implementation.
The previous implementation also ignored the tail bytes. The loop at the
end treats num_bytes as if it is the full length of the slice, but it
isn't, num_bytes number of bytes after the last 4-byte group. This can
be observed for example by changing "hello" to just "hell" in the tests.
Under the old implementation, the test will still pass. Now, the value
that comes out changes, and "hello" and "hell" hash to different values.
# Which issue does this PR close?
<!---
We generally require a GitHub issue to be filed for all bug fixes and enhancements and this helps us generate change logs for our releases. You can link an issue to this PR using the GitHub syntax. For example `Closes #123` indicates that this PR will close issue #123.
-->
Closes #.
# Rationale for this change
<!---
Why are you proposing this change? If this is already explained clearly in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand your changes and offer better suggestions for fixes.
-->
# What changes are included in this PR?
<!---
There is no need to duplicate the description in the issue here but it is sometimes worth providing a summary of the individual changes in this PR.
-->
# Are there any user-facing changes?
<!---
If there are user-facing changes then we may require documentation to be updated before approving the PR.
-->
<!---
If there are any breaking changes to public APIs, please add the `breaking change` label.
-->
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: github-unsubscribe@arrow.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
[GitHub] [arrow-rs] codecov-commenter commented on pull request #1906: Fix misaligned reference and logic error in crc32
Posted by GitBox <gi...@apache.org>.
codecov-commenter commented on PR #1906:
URL: https://github.com/apache/arrow-rs/pull/1906#issuecomment-1159593395
# [Codecov](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=h1&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation) Report
> Merging [#1906](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=desc&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation) (5d6bd7c) into [master](https://codecov.io/gh/apache/arrow-rs/commit/535cd20e3179409394d5b3c464d31ab9885a24e9?el=desc&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation) (535cd20) will **decrease** coverage by `0.00%`.
> The diff coverage is `100.00%`.
> :exclamation: Current head 5d6bd7c differs from pull request most recent head cba393e. Consider uploading reports for the commit cba393e to get more accurate results
```diff
@@ Coverage Diff @@
## master #1906 +/- ##
==========================================
- Coverage 83.42% 83.42% -0.01%
==========================================
Files 214 214
Lines 57025 57018 -7
==========================================
- Hits 47574 47567 -7
Misses 9451 9451
```
| [Impacted Files](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation) | Coverage Δ | |
|---|---|---|
| [parquet/src/util/hash\_util.rs](https://codecov.io/gh/apache/arrow-rs/pull/1906/diff?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation#diff-cGFycXVldC9zcmMvdXRpbC9oYXNoX3V0aWwucnM=) | `95.16% <100.00%> (-0.50%)` | :arrow_down: |
| [arrow/src/datatypes/datatype.rs](https://codecov.io/gh/apache/arrow-rs/pull/1906/diff?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation#diff-YXJyb3cvc3JjL2RhdGF0eXBlcy9kYXRhdHlwZS5ycw==) | `65.42% <0.00%> (-0.38%)` | :arrow_down: |
| [parquet\_derive/src/parquet\_field.rs](https://codecov.io/gh/apache/arrow-rs/pull/1906/diff?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation#diff-cGFycXVldF9kZXJpdmUvc3JjL3BhcnF1ZXRfZmllbGQucnM=) | `65.75% <0.00%> (ø)` | |
| [parquet/src/encodings/encoding.rs](https://codecov.io/gh/apache/arrow-rs/pull/1906/diff?src=pr&el=tree&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation#diff-cGFycXVldC9zcmMvZW5jb2RpbmdzL2VuY29kaW5nLnJz) | `93.65% <0.00%> (+0.19%)` | :arrow_up: |
------
[Continue to review full report at Codecov](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=continue&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation).
> **Legend** - [Click here to learn more](https://docs.codecov.io/docs/codecov-delta?utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation)
> `Δ = absolute <relative> (impact)`, `ø = not affected`, `? = missing data`
> Powered by [Codecov](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=footer&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation). Last update [535cd20...cba393e](https://codecov.io/gh/apache/arrow-rs/pull/1906?src=pr&el=lastupdated&utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation). Read the [comment docs](https://docs.codecov.io/docs/pull-request-comments?utm_medium=referral&utm_source=github&utm_content=comment&utm_campaign=pr+comments&utm_term=The+Apache+Software+Foundation).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: github-unsubscribe@arrow.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
[GitHub] [arrow-rs] tustvold commented on a diff in pull request #1906: Fix misaligned reference and logic error in crc32
Posted by GitBox <gi...@apache.org>.
tustvold commented on code in PR #1906:
URL: https://github.com/apache/arrow-rs/pull/1906#discussion_r901060762
##########
parquet/src/util/hash_util.rs:
##########
@@ -107,27 +107,18 @@ unsafe fn crc32_hash(bytes: &[u8], seed: u32) -> u32 {
#[cfg(target_arch = "x86_64")]
use std::arch::x86_64::*;
- let u32_num_bytes = std::mem::size_of::<u32>();
- let mut num_bytes = bytes.len();
- let num_words = num_bytes / u32_num_bytes;
- num_bytes %= u32_num_bytes;
-
- let bytes_u32: &[u32] = std::slice::from_raw_parts(
- &bytes[0..num_words * u32_num_bytes] as *const [u8] as *const u32,
- num_words,
- );
-
- let mut offset = 0;
let mut hash = seed;
- while offset < num_words {
- hash = _mm_crc32_u32(hash, bytes_u32[offset]);
- offset += 1;
+ for chunk in bytes
+ .chunks_exact(4)
+ .map(|chunk| u32::from_le_bytes(chunk.try_into().unwrap()))
+ {
+ hash = _mm_crc32_u32(hash, chunk);
}
- offset = num_words * u32_num_bytes;
- while offset < num_bytes {
- hash = _mm_crc32_u8(hash, bytes[offset]);
- offset += 1;
+ let remainder = bytes.len() % 4;
Review Comment:
You could use https://doc.rust-lang.org/std/slice/struct.ChunksExact.html#method.remainder
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: github-unsubscribe@arrow.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org
[GitHub] [arrow-rs] tustvold merged pull request #1906: Fix misaligned reference and logic error in crc32
Posted by GitBox <gi...@apache.org>.
tustvold merged PR #1906:
URL: https://github.com/apache/arrow-rs/pull/1906
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: github-unsubscribe@arrow.apache.org
For queries about this service, please contact Infrastructure at:
users@infra.apache.org