并行计算MPI 埃拉筛及性能优化
最后还需自行实现 optimize4 来尽可能的去优化,但是不能开 O2 、改算法等。
Using MPI to optimize the Eratosthenes Sieve algorithm. No changes to the algorithm itself, but tons of tricks of parallelization.
Optimize the Eratosthenes Sieve algorithm using MPI. Without changing the algorithm itself, how can we make it faster using parallelization?
Remove even numbers/Remove broadcast/Optimize cache/memset/loop unrolling/__builtin_popcount/register&inline assembly
Reached 27.6x speedup on 16 cores
最后还需自行实现 optimize4 来尽可能的去优化,但是不能开 O2 、改算法等。