You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Fix regression in LZ4 compression performance since 10.6 (facebook#14017)
Summary:
In RocksDB 10.6 with facebook#13805, due to inaccurate testing of an async system, it went undetected at the time that LZ4 compression was using more CPU despite making a change to reuse stream objects which dramatically improved LZ4HC compression efficiency.
This change switches to using a basic LZ4 compress API which appears to be faster than all of these:
* Legacy behavior of creating LZ4_stream_t for each compression
* 10.6-10.7 behavior of re-using streams between compressions for the same file (with stream-as-WorkingArea)
* using LZ4's extState APIs without streams (with extState-as-WorkingArea) (data not shown in below results)
Also in this PR: more improvements to sst_dump --recompress, which is arguably the best SST construction benchmark right now since db_bench seems to be so noisy due to backgroun flush+compaction, even with no compaction (FIFO). Streamlined some output and added a SST read time test, mostly for decompression performance.
Pull Request resolved: facebook#14017
Test Plan:
Performance test using sst_dump --recompress with newer sst_dump back-ported to 10.5:
```
./sst_dump --command=recompress --compression_types=kLZ4Compression
test5.sst --compression_level_from=-6 --compression_level_to=-1
```
and with default compression level.
10.5:
```
Cx level: -6 Cx size: 61608137 Write usec: 880404
Cx level: -5 Cx size: 60793749 Write usec: 840903
Cx level: -4 Cx size: 58134030 Write usec: 836365
Cx level: -3 Cx size: 55193773 Write usec: 857113
Cx level: -2 Cx size: 54013891 Write usec: 855642
Cx level: -1 Cx size: 50400393 Write usec: 865194
Cx level: 32767 Cx size: 50400393 Write usec: 886310
```
Before this change (showing the regression, more time, from 10.6:
```
Cx level: -6 Cx size: 61608137 Write usec: 933448
Cx level: -5 Cx size: 60793749 Write usec: 893826
Cx level: -4 Cx size: 58134030 Write usec: 891138
Cx level: -3 Cx size: 55193773 Write usec: 898461
Cx level: -2 Cx size: 54013891 Write usec: 897485
Cx level: -1 Cx size: 50400393 Write usec: 936970
Cx level: 32767 Cx size: 50400393 Write usec: 958764
```
After this change (faster than both the above):
```
Cx level: -6 Cx size: 63641883 Write usec: 874190
Cx level: -5 Cx size: 58860032 Write usec: 834662
Cx level: -4 Cx size: 57150188 Write usec: 832707
Cx level: -3 Cx size: 58791894 Write usec: 850305
Cx level: -2 Cx size: 53145885 Write usec: 839574
Cx level: -1 Cx size: 49809139 Write usec: 845639
Cx level: 32767 Cx size: 49809139 Write usec: 875199
```
Similar tests with dictionary compression show essentially no difference (need to use stream APIs and reuse doesn't seem to matter). LZ4HC also unaffected (still improved vs. 10.5)
Reviewed By: hx235
Differential Revision: D83722880
Pulled By: pdillinger
fbshipit-source-id: 30149dd187686d5dd98321e6aa7d74bd7653a905
0 commit comments