<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>sujet Same kernel but huge performance difference under linux and windows dans OpenCL* for CPU</title>
    <link>https://community.intel.com/t5/OpenCL-for-CPU/Same-kernel-but-huge-performance-difference-under-linux-and/m-p/1019744#M3241</link>
    <description>&lt;P&gt;Hi,&amp;nbsp;&lt;/P&gt;

&lt;P&gt;I have managed to run my kernel on iGPU under Linux and Windows.&lt;/P&gt;

&lt;P&gt;Officially linux does not support to run kernel on iGPU but an OpenCL source project "beignet" come to help.&lt;/P&gt;

&lt;P&gt;So following is the performance result for my kernel (deblocking filter in HEVC), the performance (time in seconds) was not obtained by binding event to kernel launching in OpenCL as it also depends on the OpenCL runtime implementation under windows and linux, instead, it was obtained by the host side CPU profiling utilities.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; H2D &amp;nbsp; &amp;nbsp; Kernel &amp;nbsp; &amp;nbsp; D2H&lt;/P&gt;

&lt;P&gt;Linux &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 1.95, &amp;nbsp; &amp;nbsp;3.89, &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;1.56&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 13.0080003738403px; line-height: 17.7381820678711px;"&gt;Windows &amp;nbsp; &amp;nbsp; &amp;nbsp; 6.74, &amp;nbsp; &amp;nbsp;0.85, &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;1.44&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;I am not sure whether the beignet develop team use the same compiler &amp;nbsp;to the windows OpenCL compiler, but the performance of kernel differs too much under these two systems. Also the host to device copy take much more time on Windows, can not figure out why.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Any hints?&amp;nbsp;&lt;/P&gt;

&lt;P&gt;my configuration&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;hardware:&lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;CPU: i5-4570R, &amp;nbsp;iGPU (HD5200)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;OS: Win8.1&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;iGPU driver version&amp;nbsp;10.18.10.3960, latest INDE, Visual Studio 2013&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;Linux :14.04, &lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;kernel 3.13&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;Beignet Release v1.0&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;gcc 4.8.3&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Tue, 24 Feb 2015 10:08:48 GMT</pubDate>
    <dc:creator>Biao_W_</dc:creator>
    <dc:date>2015-02-24T10:08:48Z</dc:date>
    <item>
      <title>Same kernel but huge performance difference under linux and windows</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Same-kernel-but-huge-performance-difference-under-linux-and/m-p/1019744#M3241</link>
      <description>&lt;P&gt;Hi,&amp;nbsp;&lt;/P&gt;

&lt;P&gt;I have managed to run my kernel on iGPU under Linux and Windows.&lt;/P&gt;

&lt;P&gt;Officially linux does not support to run kernel on iGPU but an OpenCL source project "beignet" come to help.&lt;/P&gt;

&lt;P&gt;So following is the performance result for my kernel (deblocking filter in HEVC), the performance (time in seconds) was not obtained by binding event to kernel launching in OpenCL as it also depends on the OpenCL runtime implementation under windows and linux, instead, it was obtained by the host side CPU profiling utilities.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; H2D &amp;nbsp; &amp;nbsp; Kernel &amp;nbsp; &amp;nbsp; D2H&lt;/P&gt;

&lt;P&gt;Linux &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 1.95, &amp;nbsp; &amp;nbsp;3.89, &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;1.56&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 13.0080003738403px; line-height: 17.7381820678711px;"&gt;Windows &amp;nbsp; &amp;nbsp; &amp;nbsp; 6.74, &amp;nbsp; &amp;nbsp;0.85, &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;1.44&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;I am not sure whether the beignet develop team use the same compiler &amp;nbsp;to the windows OpenCL compiler, but the performance of kernel differs too much under these two systems. Also the host to device copy take much more time on Windows, can not figure out why.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Any hints?&amp;nbsp;&lt;/P&gt;

&lt;P&gt;my configuration&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;hardware:&lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;CPU: i5-4570R, &amp;nbsp;iGPU (HD5200)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;OS: Win8.1&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;iGPU driver version&amp;nbsp;10.18.10.3960, latest INDE, Visual Studio 2013&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;Linux :14.04, &lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;kernel 3.13&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;Beignet Release v1.0&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 16.3636360168457px;"&gt;gcc 4.8.3&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 24 Feb 2015 10:08:48 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Same-kernel-but-huge-performance-difference-under-linux-and/m-p/1019744#M3241</guid>
      <dc:creator>Biao_W_</dc:creator>
      <dc:date>2015-02-24T10:08:48Z</dc:date>
    </item>
    <item>
      <title>Hi Biao,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/Same-kernel-but-huge-performance-difference-under-linux-and/m-p/1019745#M3242</link>
      <description>&lt;P&gt;Hi Biao,&lt;/P&gt;

&lt;P&gt;Please submit Beignet bugs here: &lt;A href="https://bugs.freedesktop.org/enter_bug.cgi?product=Beignet"&gt;&lt;U&gt;&lt;FONT color="#0066cc"&gt;&lt;/FONT&gt;&lt;/U&gt;&lt;/A&gt;&lt;A href="https://bugs.freedesktop.org/enter_bug.cgi?product=Beignet" target="_blank"&gt;https://bugs.freedesktop.org/enter_bug.cgi?product=Beignet&lt;/A&gt;&amp;nbsp; or direct questions about Beignet performance to the mailing list: &lt;A href="http://lists.freedesktop.org/mailman/listinfo/beignet"&gt;&lt;U&gt;&lt;FONT color="#0066cc"&gt;&lt;/FONT&gt;&lt;/U&gt;&lt;/A&gt;&lt;A href="http://lists.freedesktop.org/mailman/listinfo/beignet" target="_blank"&gt;http://lists.freedesktop.org/mailman/listinfo/beignet&lt;/A&gt;&amp;nbsp; - we are not supporting it in this forum.&lt;/P&gt;

&lt;P&gt;As far as D2H and H2D times, you can avoid these copies by properly aligning the memory of your buffers, using CL_USE_HOST_PTR flag when creating the buffers, and aligning the size of your buffers to 64 bytes. See this excellent article by Adam Lake: &lt;A href="https://software.intel.com/en-us/articles/getting-the-most-from-opencl-12-how-to-increase-performance-by-minimizing-buffer-copies-on-intel-processor-graphics"&gt;https://software.intel.com/en-us/articles/getting-the-most-from-opencl-12-how-to-increase-performance-by-minimizing-buffer-copies-on-intel-processor-graphics&lt;/A&gt; on how to do it properly.&lt;/P&gt;</description>
      <pubDate>Tue, 24 Feb 2015 17:29:01 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/Same-kernel-but-huge-performance-difference-under-linux-and/m-p/1019745#M3242</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-02-24T17:29:01Z</dc:date>
    </item>
  </channel>
</rss>

