<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re:Parallel Version of Code not as efficient as Serial Version of Code in Intel® oneAPI DPC++/C++ Compiler</title>
    <link>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1255972#M956</link>
    <description>&lt;P&gt;Hi Nikhil,&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Please give us an update on the provided details.&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Warm Regards,&lt;/P&gt;&lt;P&gt;Abhishek&lt;/P&gt;&lt;BR /&gt;</description>
    <pubDate>Mon, 15 Feb 2021 06:00:19 GMT</pubDate>
    <dc:creator>AbhishekD_Intel</dc:creator>
    <dc:date>2021-02-15T06:00:19Z</dc:date>
    <item>
      <title>Parallel Version of Code not as efficient as Serial Version of Code</title>
      <link>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1252408#M939</link>
      <description>&lt;P&gt;Hi there!&lt;/P&gt;
&lt;P&gt;I wrote a code for the restriction operator used in multigrid algorithms. Code is given below:&lt;/P&gt;
&lt;P&gt;#include &amp;lt;iostream&amp;gt;&lt;BR /&gt;#include &amp;lt;CL/sycl.hpp&amp;gt;&lt;BR /&gt;#include &amp;lt;vector&amp;gt;&lt;/P&gt;
&lt;P&gt;using namespace sycl;&lt;/P&gt;
&lt;P&gt;std::vector &amp;lt;float&amp;gt; Restriction2D(std::vector &amp;lt;float&amp;gt;&amp;amp; vec_h) {&lt;BR /&gt;int vec_h_dim = int(std::sqrt(vec_h.size()));&lt;BR /&gt;int vec_2h_dim = int((vec_h_dim - 1) / 2);&lt;BR /&gt;std::vector&amp;lt;float&amp;gt; vec_2h(vec_2h_dim * vec_2h_dim, 0);&lt;BR /&gt;for (int i_2h = 1; i_2h &amp;lt;= vec_2h_dim; i_2h++) {&lt;/P&gt;
&lt;P&gt;for (int j_2h = 1; j_2h &amp;lt;= vec_2h_dim; j_2h++) {&lt;/P&gt;
&lt;P&gt;vec_2h[(i_2h - 1) * vec_2h_dim + j_2h - 1] = (1 / 16) * (vec_h[(2 * i_2h - 1 - 1) * vec_h_dim + 2 * j_2h - 1 - 1] + vec_h[(2 * i_2h - 1 - 1) * vec_h_dim + 2 * j_2h]&lt;BR /&gt;+ vec_h[2 * i_2h * vec_h_dim + 2 * j_2h - 1 - 1] + vec_h[2 * i_2h * vec_h_dim + 2 * j_2h] + 2 * (vec_h[(2 * i_2h - 1) * vec_h_dim + 2 * j_2h - 1 - 1] +&lt;BR /&gt;vec_h[(2 * i_2h - 1) * vec_h_dim + 2 * j_2h] + vec_h[(2 * i_2h - 1 - 1) * vec_h_dim + 2 * j_2h - 1] + vec_h[2 * i_2h * vec_h_dim + 2 * j_2h - 1]) +&lt;BR /&gt;4 * vec_h[(2 * i_2h - 1) * vec_h_dim + 2 * j_2h - 1]);&lt;BR /&gt;}&lt;BR /&gt;}&lt;BR /&gt;return vec_2h;&lt;BR /&gt;}&lt;/P&gt;
&lt;P&gt;std::vector &amp;lt;float&amp;gt; Restriction2D_parallel(std::vector &amp;lt;float&amp;gt;&amp;amp; vec_h) {&lt;BR /&gt;int vec_h_dim = int(std::sqrt(vec_h.size()));&lt;BR /&gt;int vec_2h_dim = int((vec_h_dim - 1) / 2);&lt;BR /&gt;std::vector&amp;lt;float&amp;gt; vec_2h(vec_2h_dim * vec_2h_dim, 0);&lt;BR /&gt;cl::sycl::queue q;&lt;BR /&gt;{&lt;BR /&gt;buffer &amp;lt;float, 2&amp;gt; vec_2h_buf(vec_2h.data(), range&amp;lt;2&amp;gt;{vec_2h_dim, vec_2h_dim});&lt;BR /&gt;buffer &amp;lt;float, 2&amp;gt; vec_h_buf(vec_h.data(), range&amp;lt;2&amp;gt;{vec_h_dim, vec_h_dim});&lt;/P&gt;
&lt;P&gt;//float* host_vector_2h = malloc_host&amp;lt;float&amp;gt;(vec_2h_dim, q);&lt;BR /&gt;q.submit([&amp;amp;](handler&amp;amp; h) {&lt;BR /&gt;accessor vec_2h_acc{ vec_2h_buf , h };&lt;BR /&gt;accessor vec_h_acc{ vec_h_buf , h };&lt;BR /&gt;program p(q.get_context());&lt;BR /&gt;p.build_with_kernel_type&amp;lt;class Restriction&amp;gt;();&lt;BR /&gt;&lt;BR /&gt;h.parallel_for&amp;lt;class Restriction&amp;gt;(p.get_kernel&amp;lt;class Restriction&amp;gt;(),range&amp;lt;2&amp;gt;{ vec_2h_dim, vec_2h_dim}, [=](id&amp;lt;2&amp;gt;idx) {&lt;BR /&gt;int i_2h = idx[0]; //0 to vec_2h_dim -1&lt;BR /&gt;int j_2h = idx[1]; //0 to vec_2h_dim -1&lt;BR /&gt;vec_2h_acc[i_2h][j_2h] = (1 / 16) * (vec_h_acc[2 * i_2h - 1][2 * j_2h - 1] + vec_h_acc[2 * i_2h - 1][2 * j_2h + 1]&lt;BR /&gt;+ vec_h_acc[2 * i_2h + 1][2 * j_2h - 1] + vec_h_acc[2 * i_2h + 1][2 * j_2h + 1] + 2 * (vec_h_acc[2 * i_2h][2 * j_2h - 1] +&lt;BR /&gt;vec_h_acc[2 * i_2h][2 * j_2h + 1] + vec_h_acc[2 * i_2h - 1][2 * j_2h] + vec_h_acc[2 * i_2h + 1][2 * j_2h]) +&lt;BR /&gt;4 * vec_h_acc[2 * i_2h][2 * j_2h]);&lt;BR /&gt;});&lt;BR /&gt;});&lt;BR /&gt;//q.wait();&lt;BR /&gt;}&lt;BR /&gt;return vec_2h;&lt;BR /&gt;}&lt;/P&gt;
&lt;P&gt;int main() {&lt;BR /&gt;std::size_t size = 11108889;&lt;BR /&gt;std::vector&amp;lt;float&amp;gt; test_vec(size, 0.0);&lt;BR /&gt;for (int i = 0; i &amp;lt; test_vec.size(); i++) {&lt;BR /&gt;test_vec[i] = i / 4.0;&lt;BR /&gt;}&lt;BR /&gt;std::vector&amp;lt;float&amp;gt;test_vec_restricted = Restriction2D_parallel(test_vec);&lt;BR /&gt;//std::cout &amp;lt;&amp;lt; test_vec_restricted.size();&lt;BR /&gt;return 0;&lt;/P&gt;
&lt;P&gt;}&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;While running the Restriction2D and Restriction2D_parallel, the serial version of the code seems to perform better than the parallel version. I have also attached the results of HPC Vtune analysis for both.&lt;/P&gt;
&lt;P&gt;Can someone explain it to me why is this happening? What knowledge am I lacking here?&lt;/P&gt;</description>
      <pubDate>Tue, 02 Feb 2021 15:34:14 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1252408#M939</guid>
      <dc:creator>Nikhil_T</dc:creator>
      <dc:date>2021-02-02T15:34:14Z</dc:date>
    </item>
    <item>
      <title>Re:Parallel Version of Code not as efficient as Serial Version of Code</title>
      <link>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1253482#M943</link>
      <description>&lt;P&gt;Hi Nikhil,&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Thanks for reaching out to us.&lt;/P&gt;&lt;P&gt;From your code, we can see that you are trying to perform simple operations on vectors and finally adding them to get the desired result, and as it will take a constant time to access the elements and do some simple operations you can very well relate your code as a vector add. And as the complexity of the loops is also not exceeding O(size) your code will take less than a second to complete sequentially. &lt;/P&gt;&lt;P&gt;So this small workload is not quite ideal to compare for sequential and parallel executions. This is the reason for the difference in performance between parallel and sequential executions.&lt;/P&gt;&lt;P&gt;To get good performance stats between sequential and parallel execution you can try increasing your workload and can make your code more compute intensive. &lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Hope the provided details will help you to get more clarity on your issue.&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Warm Regards,&lt;/P&gt;&lt;P&gt;Abhishek&lt;/P&gt;&lt;BR /&gt;</description>
      <pubDate>Fri, 05 Feb 2021 13:06:21 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1253482#M943</guid>
      <dc:creator>AbhishekD_Intel</dc:creator>
      <dc:date>2021-02-05T13:06:21Z</dc:date>
    </item>
    <item>
      <title>Re:Parallel Version of Code not as efficient as Serial Version of Code</title>
      <link>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1255972#M956</link>
      <description>&lt;P&gt;Hi Nikhil,&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Please give us an update on the provided details.&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Warm Regards,&lt;/P&gt;&lt;P&gt;Abhishek&lt;/P&gt;&lt;BR /&gt;</description>
      <pubDate>Mon, 15 Feb 2021 06:00:19 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1255972#M956</guid>
      <dc:creator>AbhishekD_Intel</dc:creator>
      <dc:date>2021-02-15T06:00:19Z</dc:date>
    </item>
    <item>
      <title>Re:Parallel Version of Code not as efficient as Serial Version of Code</title>
      <link>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1258036#M969</link>
      <description>&lt;P&gt;Hi Nikhil,&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;We haven't heard back from you for a long time, so we are assuming that the provided solution had helped you in solving your issue. So we are no longer monitoring this thread.&lt;/P&gt;&lt;P&gt;Please post a new thread if you have any other issues.&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;&lt;/P&gt;&lt;P&gt;Warm Regards,&lt;/P&gt;&lt;P&gt;Abhishek&lt;/P&gt;&lt;BR /&gt;</description>
      <pubDate>Mon, 22 Feb 2021 07:08:43 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-oneAPI-DPC-C-Compiler/Parallel-Version-of-Code-not-as-efficient-as-Serial-Version-of/m-p/1258036#M969</guid>
      <dc:creator>AbhishekD_Intel</dc:creator>
      <dc:date>2021-02-22T07:08:43Z</dc:date>
    </item>
  </channel>
</rss>

