<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Hi Robert,  in OpenCL* for CPU</title>
    <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996650#M2779</link>
    <description>&lt;P&gt;Hi Robert,&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Thanks for your reply.&lt;/P&gt;

&lt;P&gt;In clinfo I see the Compute unit for this CPU is 8, does that mean each logic core (when hyperthreading enabled) is one compute unit?&lt;/P&gt;

&lt;P&gt;What if the CPU does not support AVX2, will the Intel compiler try to compile the kernel with the latest SIMD for low end CPUs?&lt;/P&gt;

&lt;P&gt;Finally, I have one kernel (SAO) from HEVC decoder. It is offloaded to CPU as OpenCL device. I also have a head coding SIMD AVX2 implementations for this kernel. The performance of OCL SAO on CPU is bad, compared to head coding SIMD.&amp;nbsp;&lt;SPAN style="font-size: 1em; line-height: 1.5;"&gt;Is there any way to measure the efficiency of Intel OpenCL compiled SIMD instructions? See attached picture.&lt;/SPAN&gt;&lt;/P&gt;</description>
    <pubDate>Tue, 03 Feb 2015 12:48:16 GMT</pubDate>
    <dc:creator>Biao_W_</dc:creator>
    <dc:date>2015-02-03T12:48:16Z</dc:date>
    <item>
      <title>OpenCL achieve 800% CPU utilization</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996648#M2777</link>
      <description>&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Hi all,&amp;nbsp;&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;I am curious about the CPU implementation of OpenCL for Intel processors.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;I run a small set of benchmark from clpeak on a i7-4770S (4 cores, hyperthreading enabled) under linux.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;it shows the CPU utilization can achieve almost 800% (using top), meaning all CPU resource are utilized.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;However, when I run the benchmark in clpeak individually, it shows maximum 400%.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Run benchmark consecutively can benefit from OpenCL runtime.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Is that mean when a workload is issued to OpenCL CPU runtime, it will not all of the cores but part of them.&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Besides, is OpenCL CPU runtime using SIMD to execute consecutive workitems?&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Appreciate in advance!&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Best,&lt;/P&gt;

&lt;P style="font-size: 13.0080003738403px; line-height: 13.0080003738403px;"&gt;Biao&lt;/P&gt;</description>
      <pubDate>Mon, 02 Feb 2015 22:45:31 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996648#M2777</guid>
      <dc:creator>Biao_W_</dc:creator>
      <dc:date>2015-02-02T22:45:31Z</dc:date>
    </item>
    <item>
      <title>Biao,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996649#M2778</link>
      <description>&lt;P&gt;Biao,&lt;/P&gt;

&lt;P&gt;OpenCL CPU runtime is very efficient at using all available CPU resources: all cores and all threads will be utilized. In addition, Kernels are compiled with AVX2 instructions in mind (a 256-bit form of SIMD). To achieve comparable performance in a regular C/C++ code, you will need to use 256-bit data types and intrinsics in addition to regular multithreading.&lt;/P&gt;</description>
      <pubDate>Mon, 02 Feb 2015 22:58:57 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996649#M2778</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-02-02T22:58:57Z</dc:date>
    </item>
    <item>
      <title>Hi Robert, </title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996650#M2779</link>
      <description>&lt;P&gt;Hi Robert,&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Thanks for your reply.&lt;/P&gt;

&lt;P&gt;In clinfo I see the Compute unit for this CPU is 8, does that mean each logic core (when hyperthreading enabled) is one compute unit?&lt;/P&gt;

&lt;P&gt;What if the CPU does not support AVX2, will the Intel compiler try to compile the kernel with the latest SIMD for low end CPUs?&lt;/P&gt;

&lt;P&gt;Finally, I have one kernel (SAO) from HEVC decoder. It is offloaded to CPU as OpenCL device. I also have a head coding SIMD AVX2 implementations for this kernel. The performance of OCL SAO on CPU is bad, compared to head coding SIMD.&amp;nbsp;&lt;SPAN style="font-size: 1em; line-height: 1.5;"&gt;Is there any way to measure the efficiency of Intel OpenCL compiled SIMD instructions? See attached picture.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 03 Feb 2015 12:48:16 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996650#M2779</guid>
      <dc:creator>Biao_W_</dc:creator>
      <dc:date>2015-02-03T12:48:16Z</dc:date>
    </item>
    <item>
      <title>Hi Biao,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996651#M2780</link>
      <description>&lt;P&gt;Hi Biao,&lt;/P&gt;

&lt;P&gt;Yes, each logic core is one compute unit. I will ask our product engineers about the last couple of questions. Typically, OpenCL kernel on the CPU device should not perform better than the hand coded AVX2 code, however I am not sure what the typical overhead for using OpenCL should be. On your diagram, which kernels are hand coded and which ones are OpenCL kernels?&lt;/P&gt;</description>
      <pubDate>Tue, 03 Feb 2015 15:57:58 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996651#M2780</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-02-03T15:57:58Z</dc:date>
    </item>
    <item>
      <title>Hi Robert, </title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996652#M2781</link>
      <description>&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Hi Robert,&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Actually there are two kernels inside the figure, but let us focus on SAO kernel only.&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;The black bar surround by a green cycle is the performance of hand coded SAO kernel while the blue bars surrounded by the other green cycle is the performance of OpenCL CPU implementation. The latter one has two parts, because I run the kernel for two subsets of the input data.&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 03 Feb 2015 16:10:38 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996652#M2781</guid>
      <dc:creator>Biao_W_</dc:creator>
      <dc:date>2015-02-03T16:10:38Z</dc:date>
    </item>
    <item>
      <title>Hi Biao,</title>
      <link>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996653#M2782</link>
      <description>&lt;P&gt;Hi Biao,&lt;/P&gt;

&lt;P&gt;Just got a response from our product engineer:&lt;/P&gt;

&lt;P&gt;&lt;FONT color="#000000" face="Times New Roman" size="3"&gt;What if the CPU does not support AVX2, will the Intel compiler try to compile the kernel with the latest SIMD for low end CPUs?&lt;/FONT&gt;&lt;/P&gt;

&lt;P style="margin: 0in 0in 0pt;"&gt;&lt;SPAN style="color: rgb(153, 51, 102); font-family: &amp;quot;Calibri&amp;quot;,sans-serif; font-size: 11pt;"&gt;Yes. OpenCL CPU run-time automatically detects and optimizes for supported vector extension. Supported CPUs can be found in release notes.&lt;/SPAN&gt;&lt;/P&gt;

&lt;P style="margin: 0in 0in 0pt;"&gt;&amp;nbsp;&lt;/P&gt;

&lt;P style="margin: 0in 0in 0pt;"&gt;&lt;FONT color="#000000"&gt;&lt;SPAN style="background: yellow; font-family: &amp;quot;Times New Roman&amp;quot;,serif; font-size: 12pt; mso-fareast-font-family: Calibri; mso-fareast-theme-font: minor-latin; mso-highlight: yellow; mso-ansi-language: EN-US; mso-fareast-language: EN-US; mso-bidi-language: AR-SA;"&gt;Is there any way to measure the efficiency of Intel OpenCL compiled SIMD instructions?&lt;/SPAN&gt;&lt;SPAN style="font-family: &amp;quot;Times New Roman&amp;quot;,serif; font-size: 12pt; mso-fareast-font-family: Calibri; mso-fareast-theme-font: minor-latin; mso-ansi-language: EN-US; mso-fareast-language: EN-US; mso-bidi-language: AR-SA;"&gt; &lt;/SPAN&gt;&lt;/FONT&gt;&lt;/P&gt;

&lt;P style="margin: 0in 0in 0pt;"&gt;&amp;nbsp;&lt;/P&gt;

&lt;P style="margin: 0in 0in 0pt;"&gt;&lt;SPAN style="color: rgb(153, 51, 102); font-family: &amp;quot;Calibri&amp;quot;,sans-serif; font-size: 11pt;"&gt;If you want to understand how efficiently OpenCL is uses HW, you can use profiler (we support Intel VTune). If the question about the quality of JIT code produced by OpenCL compiler, you can analyze JIT &amp;amp; LLVM IR using tools provided by OpenCL SDK. If the question is about why hand-coded application outperforms OpenCL app, then I suggest to use both (profiler + SDK tools) to analyze.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 03 Feb 2015 16:53:56 GMT</pubDate>
      <guid>https://community.intel.com/t5/OpenCL-for-CPU/OpenCL-achieve-800-CPU-utilization/m-p/996653#M2782</guid>
      <dc:creator>Robert_I_Intel</dc:creator>
      <dc:date>2015-02-03T16:53:56Z</dc:date>
    </item>
  </channel>
</rss>

