<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic The specific performance in Software Tuning, Performance Optimization &amp; Platform Monitoring</title>
    <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155558#M6926</link>
    <description>&lt;P&gt;The specific performance numbers might help narrow down the possible mechanisms....&lt;/P&gt;&lt;P&gt;What&amp;nbsp;compiler and compilation options were used? &amp;nbsp;What is the OS?&lt;/P&gt;&lt;P&gt;I would start by comparing the two systems with a set of smaller tests:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Single-thread performance bound to each socket on each node&lt;UL&gt;&lt;LI&gt;export OMP_NUM_THREADS=1;&amp;nbsp;numactl --membind=0 --cpunodebind=0 ./stream&lt;/LI&gt;&lt;LI&gt;export OMP_NUM_THREADS=1;&amp;nbsp;numactl --membind=1 --cpunodebind=1 ./stream&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;Multi-thread performance bound to each socket on each node, using 2..20 cores.&lt;UL&gt;&lt;LI&gt;If HyperThreading is enabled, set OMP_PROC_BIND=spread&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Thu, 12 Dec 2019 16:21:32 GMT</pubDate>
    <dc:creator>McCalpinJohn</dc:creator>
    <dc:date>2019-12-12T16:21:32Z</dc:date>
    <item>
      <title>Slower memory bandwidth on identical nodes reported by STREAM benchmark</title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155557#M6925</link>
      <description>&lt;P&gt;I am running stream benchmark on two identical nodes, but one node is reporting almost 5X slower performance compare to other node&lt;/P&gt;&lt;P&gt;Following is the node configuration&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;&lt;STRONG&gt;Processor&lt;/STRONG&gt;&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;2 X Intel(R) Xeon(R) CPU E5-2698 v4&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;2 X 20 Cores, 2.20GHz,&amp;nbsp;L1d cache:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;32 K,&amp;nbsp;L1i cache:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;32 K,&amp;nbsp;L2 cache:&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;256 K,&amp;nbsp;L3 cache:&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 51200 K&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;&lt;STRONG&gt;Memory&lt;/STRONG&gt;&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;128 GB, 2400 Hz, 4 memory channels (32GB each)&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;&amp;nbsp;&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;Please help me to identify the issue.&lt;/P&gt;&lt;P style="margin-left:0in; margin-right:0in"&gt;I have checked BIOS setting and drivers available, its identical for both nodes.&lt;/P&gt;</description>
      <pubDate>Thu, 12 Dec 2019 05:21:36 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155557#M6925</guid>
      <dc:creator>Londhe__Ashutosh</dc:creator>
      <dc:date>2019-12-12T05:21:36Z</dc:date>
    </item>
    <item>
      <title>The specific performance</title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155558#M6926</link>
      <description>&lt;P&gt;The specific performance numbers might help narrow down the possible mechanisms....&lt;/P&gt;&lt;P&gt;What&amp;nbsp;compiler and compilation options were used? &amp;nbsp;What is the OS?&lt;/P&gt;&lt;P&gt;I would start by comparing the two systems with a set of smaller tests:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Single-thread performance bound to each socket on each node&lt;UL&gt;&lt;LI&gt;export OMP_NUM_THREADS=1;&amp;nbsp;numactl --membind=0 --cpunodebind=0 ./stream&lt;/LI&gt;&lt;LI&gt;export OMP_NUM_THREADS=1;&amp;nbsp;numactl --membind=1 --cpunodebind=1 ./stream&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;LI&gt;Multi-thread performance bound to each socket on each node, using 2..20 cores.&lt;UL&gt;&lt;LI&gt;If HyperThreading is enabled, set OMP_PROC_BIND=spread&lt;/LI&gt;&lt;/UL&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 12 Dec 2019 16:21:32 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155558#M6926</guid>
      <dc:creator>McCalpinJohn</dc:creator>
      <dc:date>2019-12-12T16:21:32Z</dc:date>
    </item>
    <item>
      <title>Hello John,</title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155559#M6927</link>
      <description>&lt;P&gt;Hello John,&lt;/P&gt;&lt;P&gt;Following are the details you asked&lt;/P&gt;&lt;P&gt;compilation:&amp;nbsp;&lt;/P&gt;&lt;P&gt;gcc -fopenmp -O3 -DSTREAM_ARRAY_SIZE=60000000 stream.c -o Stream_60M.exe&lt;/P&gt;&lt;P&gt;gcc version : 6.2.0&lt;/P&gt;&lt;P&gt;OS: Linux&lt;/P&gt;&lt;P&gt;I will try the experiments you suggested and let you know.&lt;/P&gt;&lt;P&gt;Thanks for the feedback.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Fri, 13 Dec 2019 04:47:30 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155559#M6927</guid>
      <dc:creator>Londhe__Ashutosh</dc:creator>
      <dc:date>2019-12-13T04:47:30Z</dc:date>
    </item>
    <item>
      <title>Hello john,</title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155560#M6928</link>
      <description>&lt;P&gt;Hello john,&lt;/P&gt;&lt;P&gt;Issue resolved. It was due to faulty PSU which limiting the node performance.&lt;/P&gt;&lt;P&gt;Thanks for your help.&lt;/P&gt;</description>
      <pubDate>Fri, 13 Dec 2019 08:41:10 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155560#M6928</guid>
      <dc:creator>Londhe__Ashutosh</dc:creator>
      <dc:date>2019-12-13T08:41:10Z</dc:date>
    </item>
    <item>
      <title>Glad you found the problem! </title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155561#M6929</link>
      <description>&lt;P&gt;Glad you found the problem!&amp;nbsp;&lt;/P&gt;&lt;P&gt;This is an area that often causes problems in our supercomputing environment -- in many cases we would rather have a node fail than have it run slowly....&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Fri, 13 Dec 2019 17:00:12 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/Slower-memory-bandwidth-on-identical-nodes-reported-by-STREAM/m-p/1155561#M6929</guid>
      <dc:creator>McCalpinJohn</dc:creator>
      <dc:date>2019-12-13T17:00:12Z</dc:date>
    </item>
  </channel>
</rss>

