<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Thanks for bringing this to in OpenCL* for CPU</title>
    <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074922#M4546</link>
    <description>&lt;P&gt;Thanks for bringing this to our attention and sorry for the delayed reply.&amp;nbsp; We will definitely want to look into this further as it could mean there are opportunities for improvement&amp;nbsp;in our implementation.&amp;nbsp; For now, to be honest, we have some&amp;nbsp;hunches but so far no&amp;nbsp;obvious answer for the performance difference.&lt;/P&gt;

&lt;P&gt;Are the&amp;nbsp;reductions a critical bottleneck for your application?&amp;nbsp; Or does the CLOGS approach meet your needs?&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Tue, 31 Jan 2017 08:21:20 GMT</pubDate>
    <dc:creator>Jeffrey_M_Intel1</dc:creator>
    <dc:date>2017-01-31T08:21:20Z</dc:date>
    <item>
      <title>builtin workgroup reduction performance</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074921#M4545</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;

&lt;P&gt;I'm working on writing a global reduction in OpenCL 2.0. I started with the implementation from CLOGS:&lt;BR /&gt;
	&lt;A href="https://sourceforge.net/p/clogs/wiki/Home/" target="_blank"&gt;https://sourceforge.net/p/clogs/wiki/Home/&lt;/A&gt;&lt;/P&gt;

&lt;P&gt;Essentially, the approach is just a series of workgroup wide reductions that are combined at the end. I thought I would try updating the implementation to use the OpenCL 2.0 workgroup built-in reduction, i.e.&amp;nbsp;work_group_reduce_add().&lt;/P&gt;

&lt;P&gt;I was surprised that the global reduction performs slower when the workgroup reductions are computed using the built-in&amp;nbsp;reduction. Specifically, I ran a test of 1000 reductions on randomly sized arrays (sizes in range 1 - 100000). The random numbers are provided the same seed, so they will be the same for different runs.&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 1em;"&gt;Using the &lt;/SPAN&gt;built-in&lt;SPAN style="font-size: 1em;"&gt; reduction, the combined total kernel time is ~39 ms. Using the CLOGS approach, the total kernel time is ~32 ms. I was also surprised to see that the kernel using the built-in&amp;nbsp;reduction used&amp;nbsp;&lt;/SPAN&gt;1796&amp;nbsp; bytes of local memory, while the CLOGS approach used only&amp;nbsp;1028 bytes (as reported by OpenCL code builder).&lt;/P&gt;

&lt;P&gt;We're interested in learning why the built-in reduction appears to be slower and use more resources than the simple CLOGS approach. Perhaps there are trade-offs we're not aware of? We get similar results on an AMD GPU.&lt;/P&gt;

&lt;P&gt;I've attached the kernel file that implements the global reduction (with both the built-in and CLOG approach to workgroup wide reductions).&lt;/P&gt;

&lt;P&gt;The GPU I am using is:&lt;BR /&gt;
	Intel HD5500&lt;BR /&gt;
	Windows 10&lt;BR /&gt;
	driver:&amp;nbsp;20.19.15.4531&lt;/P&gt;

&lt;P&gt;Thanks in advance!&lt;/P&gt;</description>
      <pubDate>Thu, 19 Jan 2017 20:27:58 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074921#M4545</guid>
      <dc:creator>Tyler_S_2</dc:creator>
      <dc:date>2017-01-19T20:27:58Z</dc:date>
    </item>
    <item>
      <title>Thanks for bringing this to</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074922#M4546</link>
      <description>&lt;P&gt;Thanks for bringing this to our attention and sorry for the delayed reply.&amp;nbsp; We will definitely want to look into this further as it could mean there are opportunities for improvement&amp;nbsp;in our implementation.&amp;nbsp; For now, to be honest, we have some&amp;nbsp;hunches but so far no&amp;nbsp;obvious answer for the performance difference.&lt;/P&gt;

&lt;P&gt;Are the&amp;nbsp;reductions a critical bottleneck for your application?&amp;nbsp; Or does the CLOGS approach meet your needs?&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 31 Jan 2017 08:21:20 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074922#M4546</guid>
      <dc:creator>Jeffrey_M_Intel1</dc:creator>
      <dc:date>2017-01-31T08:21:20Z</dc:date>
    </item>
    <item>
      <title>Thanks for the reply Jeffrey!</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074923#M4547</link>
      <description>&lt;P&gt;Thanks for the reply Jeffrey! We'd definitely be interested in hearing more details if/when they become available.&lt;/P&gt;

&lt;P&gt;Reductions are not our critical bottleneck and the CLOGS approach will work fine for now. We will definitely switch to the builtins if/when the performance catches up. I really like the idea of these builtins (much easier than trying to write efficient reductions/scans for different chips).&lt;/P&gt;

&lt;P&gt;Thanks again!&lt;/P&gt;</description>
      <pubDate>Tue, 31 Jan 2017 13:33:21 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074923#M4547</guid>
      <dc:creator>Tyler_S_2</dc:creator>
      <dc:date>2017-01-31T13:33:21Z</dc:date>
    </item>
    <item>
      <title>I have experienced similar</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074924#M4548</link>
      <description>&lt;P&gt;&lt;SPAN style="font-size: 1em;"&gt;I have experienced similar issues on other platforms, as well. As such, I had written a small benchmark for the evaluation of workgroup/subgroup reductions.&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 1em;"&gt;For more information you may check:&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;A href="https://github.com/ekondis/cl2-reduce-bench" target="_blank"&gt;https://github.com/ekondis/cl2-reduce-bench&lt;/A&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Sat, 18 Feb 2017 08:08:45 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074924#M4548</guid>
      <dc:creator>Elias_K_1</dc:creator>
      <dc:date>2017-02-18T08:08:45Z</dc:date>
    </item>
    <item>
      <title>Very interesting results</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074925#M4549</link>
      <description>&lt;P&gt;Very interesting results Elias, thanks for sharing!&amp;nbsp;&lt;/P&gt;

&lt;P&gt;If I'm reading your results correctly, the only platform where the&amp;nbsp;Workgroup function outperforms the shared memory implementation is the Intel CPU (and by a considerable ~6x amount). On some of the AMD chips, it looks like the hybrid approach gives some small performance benefits.&lt;/P&gt;

&lt;P&gt;I'll try running this on our Intel HD5500 and Iris 6100 sometime this week and let you know the results.&lt;/P&gt;</description>
      <pubDate>Mon, 20 Feb 2017 11:48:49 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/builtin-workgroup-reduction-performance/m-p/1074925#M4549</guid>
      <dc:creator>Tyler_S_2</dc:creator>
      <dc:date>2017-02-20T11:48:49Z</dc:date>
    </item>
  </channel>
</rss>

