- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
The HPL benchmark performance obtained on a host + 2 MIC cards is coming only 719GFlops. The Host system has 128 GB memory. The theoretical peak is 1.2TF + 1.2TF + 460GFLOPS = 2.8TF. Efficiency is just 35%(~). May I know how to optimize the hpl performance? I've used the OFFLOAD execution, with the executable xhpl_offload_intel64(manually compiled from source of Intel linpack 11.1.1)
Link Copied
1 Reply
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
try use another floating point models and use huge pages , which should increase ur performance as well , thats a factor of 2 or 3 if u use both. It also depends on the problem u just want to compute, do u have many memor acceses ?.
best regards

Reply
Topic Options
- Subscribe to RSS Feed
- Mark Topic as New
- Mark Topic as Read
- Float this Topic for Current User
- Bookmark
- Subscribe
- Printer Friendly Page