Repository navigation
Performance of blake2 #5
Description
Activity
Well, here are the results of a benchmark for the code extracted from rust compiler:
test bench_16 ... bench: 475 ns/iter (+/- 9) = 33 MB/s test bench_1k ... bench: 2,177 ns/iter (+/- 141) = 470 MB/s test bench_256 ... bench: 536 ns/iter (+/- 10) = 477 MB/s test bench_64k ... bench: 131,659 ns/iter (+/- 3,418) = 497 MB/sThe 16 bytes case is probably slow because I'm not skipping some setup cost (you can observe benchmarks at the bottom of the gist linked above). Otherwise, it looks like rust has much better implementation than this crate (unless I'm testing it wrong).
Thank you for reporting this!
The original code from rust-crypto for blake2 is somewhat messy so I planned to work on it later, performance issues will be an additional incentive to do it. I will probably start working on it in January, but if you would like to try to improve this crate before that you are welcome!
Regarding SIMD, none of the crates currently use explicit SIMD instructions, they only use
fake-simdcrate to help LLVM optimiser and to prepare ground for future SIMD stabilization. I've tried to usesimd-alt(fork ofsimdwith minor API changes) and it gave a measurable improvement, but because it's a nigthly only I've decided not to use it for now.Well, I don't have time to invest into it right now. I'm just trying to choose good hasher that will serve me in the long term. So I'm going to use blake2 and hopefully will benefit from it even more in future.
Thank you for work on this library!
With a delay I've started rework of blake2 crate, you can see it in #17. For blake2b I am currently getting up to 720 MB/s on my machine and 470 MB/s for blake2s. I think several parts of the code can be optimized even further.
I think we can close this issue now. It's definitely should be possible to optimize code even further, but I don't see easy ways to do it for current implementation.
- added a commit that references this issue
on Mar 15, 2025
Hi,
It looks like performance of blake2 compared to sha512 is very similar (while still faster):
But https://blake2.net/ shows that performance should be about 3x faster.
Any ideas? It looks like blake2 does not use SIMD? Is it not needed? Is there a chance that blake2 crate here will be optimized better in future? Or it's just because recent processors already execute sha512 much faster?