- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Hello,
I'm doing some performance benchmarks on Intel's mkl right now and I noticed some unexpected issues. When using dgemm_() function instead of cblas_dgemm() on the same matrix I get around 4 times fewer Gflop/s. It's the same either with the non-parallel and parallel version.
Did someone experience similiar things or could probably point me to a possible failure?
Thanks
I'm doing some performance benchmarks on Intel's mkl right now and I noticed some unexpected issues. When using dgemm_() function instead of cblas_dgemm() on the same matrix I get around 4 times fewer Gflop/s. It's the same either with the non-parallel and parallel version.
Did someone experience similiar things or could probably point me to a possible failure?
Thanks
Link Copied
0 Replies

Reply
Topic Options
- Subscribe to RSS Feed
- Mark Topic as New
- Mark Topic as Read
- Float this Topic for Current User
- Bookmark
- Subscribe
- Printer Friendly Page