<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic main memory memcpy performance on Haswells in Software Tuning, Performance Optimization &amp; Platform Monitoring</title>
    <link>https://community.intel.com/t5/Software-Tuning-Performance/main-memory-memcpy-performance-on-Haswells/m-p/1121248#M6184</link>
    <description>&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp;I'm using my NetPIPE communication benchmark to measure the copy rate for various array sizes from main memory (no cache effects). &amp;nbsp;On sandy/ivy bridge I see nice curves for both gcc and icc _intel_fast_memcpy but on Haswells the performance is significantly lower in the mid-range for both between around 8 kB to 4 MB only achieving decent performance for very large array sizes. &amp;nbsp;To me it seems that the the code is just not tuned for Haswells but it seems odd that I'm seeing the same deficiency for both gcc and icc routines. &amp;nbsp;I've attached a graph showing both curves (these show copy rates, bandwidths would be 2x this). &amp;nbsp;Measurements avoid cache effects by moving the source and destination pointers through a very large memory buffer.&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Thu, 20 Oct 2016 23:17:25 GMT</pubDate>
    <dc:creator>Dave_T_1</dc:creator>
    <dc:date>2016-10-20T23:17:25Z</dc:date>
    <item>
      <title>main memory memcpy performance on Haswells</title>
      <link>https://community.intel.com/t5/Software-Tuning-Performance/main-memory-memcpy-performance-on-Haswells/m-p/1121248#M6184</link>
      <description>&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp;I'm using my NetPIPE communication benchmark to measure the copy rate for various array sizes from main memory (no cache effects). &amp;nbsp;On sandy/ivy bridge I see nice curves for both gcc and icc _intel_fast_memcpy but on Haswells the performance is significantly lower in the mid-range for both between around 8 kB to 4 MB only achieving decent performance for very large array sizes. &amp;nbsp;To me it seems that the the code is just not tuned for Haswells but it seems odd that I'm seeing the same deficiency for both gcc and icc routines. &amp;nbsp;I've attached a graph showing both curves (these show copy rates, bandwidths would be 2x this). &amp;nbsp;Measurements avoid cache effects by moving the source and destination pointers through a very large memory buffer.&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 20 Oct 2016 23:17:25 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Tuning-Performance/main-memory-memcpy-performance-on-Haswells/m-p/1121248#M6184</guid>
      <dc:creator>Dave_T_1</dc:creator>
      <dc:date>2016-10-20T23:17:25Z</dc:date>
    </item>
  </channel>
</rss>

