<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic I optimized communication in Software Archive</title>
    <link>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097820#M67526</link>
    <description>&lt;P&gt;I optimized communication between host and coprocessor to a minimum. I copy to the coprocessor only nesesary data to computation.&amp;nbsp;Every timestepe i transfer three arrays that containing &amp;nbsp;near to 1000 elements&amp;nbsp;(as shown in a source code in previous post).&amp;nbsp;&lt;/P&gt;

&lt;PRE class="brush:cpp;"&gt;/*...*/
 
char* transfer;
char* offload;
const int n = 1000;
 
for(int i=0; i&amp;lt;2000; i++)
{
    #pragma offload_transfer target(mic : 0) \
        in( tab1 : length(n) alloc_if(0) free_if(0) ) \
        in( tab2 : length(n) alloc_if(0) free_if(0) ) \
        in( tab3 : length(n) alloc_if(0) free_if(0) ) \
        signal(transfer)
    
	// First part of computation on Host
	
	#pragma offload_transfer target(mic : 0) wait(transfer) \
		nocopy(tab1) nocopy(tab2) nocopy(tab3) signal(offload)
	{
        // Computation on Intel Xeon Phi
    }
 
    // Second part of computation on Host
	
    #pragma offload_wait target(mic : 0) wait(offload)
}  

/*...*/&lt;/PRE&gt;

&lt;P&gt;This code shown my idea of asynchronous transfer. First part of computation on a CPU takes more time than asynchronous transfer. I think that the aggregated time of calling first pragma ( offload_transfer) and second pragma (offload wait() ... signal() ) is more time-consuming than data transfer presented in previous post.&lt;/P&gt;

&lt;P&gt;Thanks for your reply, Mr.&amp;nbsp;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Dempsey. :)&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Thu, 21 Jan 2016 16:55:00 GMT</pubDate>
    <dc:creator>H__Kamil</dc:creator>
    <dc:date>2016-01-21T16:55:00Z</dc:date>
    <item>
      <title>Offload transfer question</title>
      <link>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097818#M67524</link>
      <description>&lt;P&gt;Hi, i have a question about transfer data from host do coprocessor.&amp;nbsp;Look at samplce code below.&amp;nbsp;Are data transferred asynchronously to coprocessor?&amp;nbsp;I would like to overlap transfer and computation performed on Intel Xeon Phi with computation carried out by CPU.&amp;nbsp;When i use combination of offload transfer signal() and offload wait() performance of computation is a lower than in code presented below.&lt;/P&gt;

&lt;PRE class="brush:cpp;"&gt;/*...*/

char* offload;
const int n = 1000;

for(int i=0; i&amp;lt;2000; i++)
{
	#pragma offload target(mic : 0) \
		in( tab1 : length(n) alloc_if(0) free_if(0) ) \
		in( tab2 : length(n) alloc_if(0) free_if(0) ) \
		in( tab3 : length(n) alloc_if(0) free_if(0) ) \
		signal(offload)
	{
		// Computation on Intel Xeon Phi
	}

	// Computation on Host

	#pragma offload_wait target(mic : 0) wait(signal)
}	
/*...*/&lt;/PRE&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 21 Jan 2016 13:37:28 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097818#M67524</guid>
      <dc:creator>H__Kamil</dc:creator>
      <dc:date>2016-01-21T13:37:28Z</dc:date>
    </item>
    <item>
      <title>What happens when you</title>
      <link>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097819#M67525</link>
      <description>&lt;P&gt;What happens when you increase n and/or increase the computation on the Xeon Phi?&lt;/P&gt;

&lt;P&gt;Meaning, there is a cost of setting up the signal, managing the signal&amp;nbsp;and for the wait on the signal. When the runtime on the Xeon Phi (per offload) is less than about 2x that of the &lt;STRONG&gt;additional &lt;/STRONG&gt;overhead, it may not be worth it to use the asynchronous offload. You can create a test with the above sketch code varying n and/or varying the computation load per n on the Xeon Phi to get some metrics as to when it is advantageous to use asynchronous offloading&amp;nbsp;for your application.&lt;/P&gt;

&lt;P&gt;Jim Dempsey&lt;/P&gt;</description>
      <pubDate>Thu, 21 Jan 2016 15:28:24 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097819#M67525</guid>
      <dc:creator>jimdempseyatthecove</dc:creator>
      <dc:date>2016-01-21T15:28:24Z</dc:date>
    </item>
    <item>
      <title>I optimized communication</title>
      <link>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097820#M67526</link>
      <description>&lt;P&gt;I optimized communication between host and coprocessor to a minimum. I copy to the coprocessor only nesesary data to computation.&amp;nbsp;Every timestepe i transfer three arrays that containing &amp;nbsp;near to 1000 elements&amp;nbsp;(as shown in a source code in previous post).&amp;nbsp;&lt;/P&gt;

&lt;PRE class="brush:cpp;"&gt;/*...*/
 
char* transfer;
char* offload;
const int n = 1000;
 
for(int i=0; i&amp;lt;2000; i++)
{
    #pragma offload_transfer target(mic : 0) \
        in( tab1 : length(n) alloc_if(0) free_if(0) ) \
        in( tab2 : length(n) alloc_if(0) free_if(0) ) \
        in( tab3 : length(n) alloc_if(0) free_if(0) ) \
        signal(transfer)
    
	// First part of computation on Host
	
	#pragma offload_transfer target(mic : 0) wait(transfer) \
		nocopy(tab1) nocopy(tab2) nocopy(tab3) signal(offload)
	{
        // Computation on Intel Xeon Phi
    }
 
    // Second part of computation on Host
	
    #pragma offload_wait target(mic : 0) wait(offload)
}  

/*...*/&lt;/PRE&gt;

&lt;P&gt;This code shown my idea of asynchronous transfer. First part of computation on a CPU takes more time than asynchronous transfer. I think that the aggregated time of calling first pragma ( offload_transfer) and second pragma (offload wait() ... signal() ) is more time-consuming than data transfer presented in previous post.&lt;/P&gt;

&lt;P&gt;Thanks for your reply, Mr.&amp;nbsp;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Dempsey. :)&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Thu, 21 Jan 2016 16:55:00 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Offload-transfer-question/m-p/1097820#M67526</guid>
      <dc:creator>H__Kamil</dc:creator>
      <dc:date>2016-01-21T16:55:00Z</dc:date>
    </item>
  </channel>
</rss>

