<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Hi Norbert, in OpenCL* for CPU</title>
    <link>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012860#M3087</link>
    <description>&lt;P&gt;Hi Norbert,&lt;/P&gt;

&lt;P&gt;Would it be possible to provide your benchmarks to us? If you do not want to post it in a public forum, you could send it as a private message.&lt;/P&gt;

&lt;P&gt;Thanks!&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Thu, 16 Jul 2015 15:53:55 GMT</pubDate>
    <dc:creator>Robert_I_Intel</dc:creator>
    <dc:date>2015-07-16T15:53:55Z</dc:date>
    <item>
      <title>Random memory read performance difference between GPU and CPU (I7-4770R)?</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012859#M3086</link>
      <description>&lt;P&gt;We are running a simple&amp;nbsp;code doing random reads and sequential write (i.e. gather operation) on both the CPU and GPU part of the I7-4770R (separately, one at a time) and experiencing 4x slower performance on the GPU compared to the CPU. When doing sequential reads and writes and even random writes, the performance is very similar indicating that both the internals of the chip as well as the memory controller allows the GPU to access the DRAM with the same speed the CPU does. However have no idea why random reads suffer a 4x performance penalty and this limits our application’s performance quite a lot. Would be good to know what the reason of this performance difference is and see whether there is some remedy for it.&lt;/P&gt;
&lt;P&gt;Here are also the numbers from our experiments. The metric is execution time, so the lower the better.&lt;/P&gt;
&lt;TABLE style="WIDTH: 1177px" border="0" cellspacing="0" cellpadding="0" width="1177"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD style="WIDTH: 715px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;　&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 52px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;&lt;STRONG&gt;MAP&lt;/STRONG&gt;&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;&lt;STRONG&gt;REDUCE&lt;/STRONG&gt;&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;&lt;STRONG&gt;GATHER&lt;/STRONG&gt;&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 68px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;&lt;STRONG&gt;SCATTER&lt;/STRONG&gt;&lt;/P&gt;&lt;/TD&gt;&lt;/TR&gt;
&lt;TR&gt;
&lt;TD style="WIDTH: 715px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;Intel i-4770r IrisPro-16G mem-4 Cores-OpenMP-CPU&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 52px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;24.73&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;13.65&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;36.34&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 68px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;231.67&lt;/P&gt;&lt;/TD&gt;&lt;/TR&gt;
&lt;TR&gt;
&lt;TD style="WIDTH: 715px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;Intel i-4770r IrisPro-16G mem-40 EU-OpenCL-GPU&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 52px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;23.55&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;16.29&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 60px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;167.03&lt;/P&gt;&lt;/TD&gt;
&lt;TD style="WIDTH: 68px; HEIGHT: 19px" nowrap=""&gt;
&lt;P align="center"&gt;270.7&lt;/P&gt;&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;</description>
      <pubDate>Thu, 16 Jul 2015 05:44:20 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012859#M3086</guid>
      <dc:creator>Norbert_Egi</dc:creator>
      <dc:date>2015-07-16T05:44:20Z</dc:date>
    </item>
    <item>
      <title>Hi Norbert,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012860#M3087</link>
      <description>&lt;P&gt;Hi Norbert,&lt;/P&gt;

&lt;P&gt;Would it be possible to provide your benchmarks to us? If you do not want to post it in a public forum, you could send it as a private message.&lt;/P&gt;

&lt;P&gt;Thanks!&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 16 Jul 2015 15:53:55 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012860#M3087</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-07-16T15:53:55Z</dc:date>
    </item>
    <item>
      <title>Hi Robert,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012861#M3088</link>
      <description>&lt;DIV&gt;Hi Robert,&lt;/DIV&gt;

&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;

&lt;DIV&gt;Thank you for the help. Please see the details here:&lt;/DIV&gt;

&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;

&lt;DIV&gt;
	&lt;P&gt;For map:&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;create an input &amp;nbsp;and an output array of 32M integer elements each&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;fill the input array with data.&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;Walk the input array in sequence , assigning value of each input array element to output array element in squence&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;Int a&lt;N&gt;, b&lt;N&gt;;&lt;/N&gt;&lt;/N&gt;&lt;/P&gt;

	&lt;P&gt;fill_data(a);&lt;/P&gt;

	&lt;P&gt;for(i=o; i&amp;lt;32*1024*1024; ++i)&lt;/P&gt;

	&lt;P&gt;&amp;nbsp; b&lt;I&gt; = a&lt;I&gt;;&lt;/I&gt;&lt;/I&gt;&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;For gather:&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;Create an input, and output and a index array of 32M elements each&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;fill the input array with data&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;fill the index array with random indices into the output array&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;walk the index array in sequence, &amp;nbsp;using the random index value to gather from input array for sequential assignment to output&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;int a&lt;N&gt;, b&lt;N&gt;, index&lt;N&gt;;&lt;/N&gt;&lt;/N&gt;&lt;/N&gt;&lt;/P&gt;

	&lt;P&gt;fill_data(a)&lt;/P&gt;

	&lt;P&gt;fill_random_index(index);&amp;nbsp; // fills with random value between 0 and 32M-1&lt;/P&gt;

	&lt;P&gt;for(i=0; i&amp;lt;32*1024*1024; ++i)&lt;/P&gt;

	&lt;P&gt;{&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&amp;nbsp; idx = index&lt;I&gt;;&lt;/I&gt;&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&amp;nbsp; b&lt;I&gt; = a[idx];&lt;/I&gt;&lt;/P&gt;

	&lt;P&gt;}&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;For scatter:&lt;/P&gt;

	&lt;P&gt;-&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;similar to gather, but &amp;nbsp;walkt the index array in sequence and using the random index value to scatter to output array from sequentially read input&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;int a&lt;N&gt;, b&lt;N&gt;, index&lt;N&gt;;&lt;/N&gt;&lt;/N&gt;&lt;/N&gt;&lt;/P&gt;

	&lt;P&gt;fill_data(a)&lt;/P&gt;

	&lt;P&gt;fill_random_index(index);&amp;nbsp; // fills with random value between 0 and 32M-1&lt;/P&gt;

	&lt;P&gt;for(i=0; i&amp;lt;32*1024*1024; ++i)&lt;/P&gt;

	&lt;P&gt;{&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&amp;nbsp; idx = index&lt;I&gt;;&lt;/I&gt;&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&amp;nbsp; b[idx] = a&lt;I&gt;;&lt;/I&gt;&lt;/P&gt;

	&lt;P&gt;}&lt;/P&gt;

	&lt;P&gt;&amp;nbsp;&lt;/P&gt;

	&lt;P&gt;On the GPU, it’s just single OpenCL kernels, on CPU, we use openmp with multiple cores.&lt;/P&gt;
&lt;/DIV&gt;

&lt;DIV&gt;Best regards,&lt;/DIV&gt;

&lt;DIV&gt;Norbert&lt;/DIV&gt;</description>
      <pubDate>Fri, 17 Jul 2015 02:05:13 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012861#M3088</guid>
      <dc:creator>Norbert_Egi</dc:creator>
      <dc:date>2015-07-17T02:05:13Z</dc:date>
    </item>
    <item>
      <title>Norbert,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012862#M3089</link>
      <description>&lt;P&gt;Norbert,&lt;/P&gt;

&lt;P&gt;Couple of things:&lt;/P&gt;

&lt;P&gt;1. What about reduce?&lt;/P&gt;

&lt;P&gt;2. If you provide the actual code, this would speed up things quite a bit.&lt;/P&gt;

&lt;P&gt;Thanks!&lt;/P&gt;</description>
      <pubDate>Fri, 17 Jul 2015 17:00:32 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Random-memory-read-performance-difference-between-GPU-and-CPU-I7/m-p/1012862#M3089</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-07-17T17:00:32Z</dc:date>
    </item>
  </channel>
</rss>

