<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Is the assignment at line 83 in Software Archive</title>
    <link>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033547#M43040</link>
    <description>&lt;P&gt;Is the assignment at line 83 to array c what you were intending?&lt;/P&gt;
&lt;P&gt;It seems&amp;nbsp;perhaps the&amp;nbsp;use of "i" might not have been intended since that is associated with &lt;STRONG&gt;REPEATNTIMES &lt;/STRONG&gt;whereas the arrays are sized based on &lt;STRONG&gt;ROWS &lt;/STRONG&gt;and &lt;STRONG&gt;COLWIDTH; &lt;/STRONG&gt;thus&amp;nbsp;I wonder&amp;nbsp;if line 83 maybe should be:&lt;/P&gt;
&lt;P&gt;c&lt;K&gt;&lt;J&gt; += a&lt;K&gt;&lt;J&gt; * b&lt;K&gt;&lt;J&gt;;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/P&gt;
&lt;P&gt;I see a difference in execution times with dynamic allocation when running natively on the coprocessor so I will investigate further and post again after I know more.&lt;/P&gt;</description>
    <pubDate>Fri, 24 Oct 2014 15:19:54 GMT</pubDate>
    <dc:creator>Kevin_D_Intel</dc:creator>
    <dc:date>2014-10-24T15:19:54Z</dc:date>
    <item>
      <title>Dynamic allocation problems on Xeon Phi</title>
      <link>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033546#M43039</link>
      <description>&lt;P&gt;I am creating a simple matrix multiplication procedure, operating on the Intel Xeon Phi architecture.&lt;/P&gt;

&lt;P&gt;I am using, aligned data. However, if the matrices are allocated using dynamic memory (posix_memalign), the computation incurs in a severe slow down, i.e. for TYPE=float and 512x512 matrices takes ~0.55s in the dynamic case while in the other case ~0.07s.&lt;/P&gt;

&lt;P&gt;On a different architecture (Intel Xeon E5-2650 @ 2.00GHz), the problem changes because the static allocated case doesn't calculate the matrix (it gives me all zeros when i print a random position of C, I think because the #pragma simd. Anyway, the dynamic allocating case takes about 0.08s.&lt;/P&gt;

&lt;P&gt;Here is the code, i also attached the optimization reports of static &amp;amp; dynamic cases:&lt;/P&gt;

&lt;PRE class="brush:cpp;"&gt;#define ROW 512
#define COLWIDTH 512
#define REPEATNTIMES 512

#include &amp;lt;sys/time.h&amp;gt;
#include &amp;lt;stdio.h&amp;gt;
#include &amp;lt;math.h&amp;gt;
#include &amp;lt;stdlib.h&amp;gt;

#define FTYPE float
#define ALIGNMENT 128
double clock_it(void)
{
        double duration = 0.0;
        struct timeval start;

        gettimeofday(&amp;amp;start, NULL);
        duration = (double)(start.tv_sec + start.tv_usec/1000000.0);
        return duration;
}

int main()
{
        double execTime = 0.0;
        double startTime, endTime;

    int k, size1, size2, i, j;

#ifdef STACK
	printf("Using Stack!\n");
        FTYPE a[ROW][COLWIDTH];
        FTYPE b[ROW][COLWIDTH];
        FTYPE c[ROW][COLWIDTH];
        for(i=0; i&amp;lt;ROW; i++){
                for(j=0; j&amp;lt;COLWIDTH; j++){
                        a&lt;I&gt;&lt;J&gt; = 1.0f;
                        b&lt;I&gt;&lt;J&gt; = 1.0f;
                        c&lt;I&gt;&lt;J&gt; = 0.0f;
                }
        }
#else
     	printf("Using Heap!\n");
        FTYPE **a;
        posix_memalign((void **) &amp;amp;a, ALIGNMENT, sizeof(FTYPE*)*ROW);
        FTYPE **b;
        posix_memalign((void **) &amp;amp;b, ALIGNMENT, sizeof(FTYPE*)*ROW);
        FTYPE **c;
        posix_memalign((void **) &amp;amp;c, ALIGNMENT, sizeof(FTYPE*)*ROW);
        for(i=0; i&amp;lt;ROW; i++){
                posix_memalign((void **) &amp;amp;a&lt;I&gt;, ALIGNMENT, sizeof(FTYPE)*COLWIDTH);
                posix_memalign((void **) &amp;amp;b&lt;I&gt;, ALIGNMENT, sizeof(FTYPE)*COLWIDTH);
                posix_memalign((void **) &amp;amp;c&lt;I&gt;, ALIGNMENT, sizeof(FTYPE)*COLWIDTH);
                for(j=0; j&amp;lt;COLWIDTH; j++){
                        a&lt;I&gt;&lt;J&gt; = 1.0f;
                        b&lt;I&gt;&lt;J&gt; = 1.0f;
                        c&lt;I&gt;&lt;J&gt; = 0.0f;
                }
        }
#endif
      	size1 = ROW;
        size2 = COLWIDTH;
        printf("\nROW:%d COL: %d\n",ROW,COLWIDTH);

        //start timing the matrix multiply code
        startTime = clock_it();
        #ifndef STACK
        __assume_aligned(a, ALIGNMENT);
        __assume_aligned(b, ALIGNMENT);
        __assume_aligned(c, ALIGNMENT);
        #endif
	#pragma vector aligned
        for (i = 0; i &amp;lt; REPEATNTIMES; i++) {
                #pragma vector aligned
                for (k = 0; k &amp;lt; size1; k++) {
                        #pragma simd
                        #pragma vector aligned
                        for (j = 0;j &amp;lt; size2; j++) {
                                #ifndef STACK
                                        __assume_aligned(a&lt;I&gt;, ALIGNMENT);
                                        __assume_aligned(b&lt;K&gt;, ALIGNMENT);
                                        __assume_aligned(c&lt;I&gt;, ALIGNMENT);
                                #endif
                                c&lt;I&gt;&lt;J&gt; += a&lt;I&gt;&lt;K&gt; * b&lt;K&gt;&lt;J&gt;;
                        }
                }
        }

	endTime = clock_it();
        execTime = endTime - startTime;

        printf("Execution time is %2.3f seconds\n", execTime);
        printf("GigaFlops = %f\n", (((double)REPEATNTIMES * (double)COLWIDTH * (double)ROW * 2.0) / (double)(execTime))/1000000000.0);
        printf("Random c_i,j %f\n", c[rand()%512][rand()%512]);
	return 0;
}
&lt;/J&gt;&lt;/K&gt;&lt;/K&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/I&gt;&lt;/K&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/I&gt;&lt;/I&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/J&gt;&lt;/I&gt;&lt;/PRE&gt;

&lt;P&gt;Any help is appreciated!&lt;/P&gt;</description>
      <pubDate>Thu, 23 Oct 2014 18:19:15 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033546#M43039</guid>
      <dc:creator>Luca_A_</dc:creator>
      <dc:date>2014-10-23T18:19:15Z</dc:date>
    </item>
    <item>
      <title>Is the assignment at line 83</title>
      <link>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033547#M43040</link>
      <description>&lt;P&gt;Is the assignment at line 83 to array c what you were intending?&lt;/P&gt;
&lt;P&gt;It seems&amp;nbsp;perhaps the&amp;nbsp;use of "i" might not have been intended since that is associated with &lt;STRONG&gt;REPEATNTIMES &lt;/STRONG&gt;whereas the arrays are sized based on &lt;STRONG&gt;ROWS &lt;/STRONG&gt;and &lt;STRONG&gt;COLWIDTH; &lt;/STRONG&gt;thus&amp;nbsp;I wonder&amp;nbsp;if line 83 maybe should be:&lt;/P&gt;
&lt;P&gt;c&lt;K&gt;&lt;J&gt; += a&lt;K&gt;&lt;J&gt; * b&lt;K&gt;&lt;J&gt;;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/P&gt;
&lt;P&gt;I see a difference in execution times with dynamic allocation when running natively on the coprocessor so I will investigate further and post again after I know more.&lt;/P&gt;</description>
      <pubDate>Fri, 24 Oct 2014 15:19:54 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033547#M43040</guid>
      <dc:creator>Kevin_D_Intel</dc:creator>
      <dc:date>2014-10-24T15:19:54Z</dc:date>
    </item>
    <item>
      <title>Quote:Kevin Davis (Intel)</title>
      <link>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033548#M43041</link>
      <description>&lt;P&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;Kevin Davis (Intel) wrote:&lt;BR /&gt;&lt;P&gt;&lt;/P&gt;

&lt;P&gt;Is the assignment at line 83 to array c what you were intending?&lt;/P&gt;

&lt;P&gt;It seems&amp;nbsp;perhaps the&amp;nbsp;use of "i" might not have been intended since that is associated with &lt;STRONG&gt;REPEATNTIMES &lt;/STRONG&gt;whereas the arrays are sized based on &lt;STRONG&gt;ROWS &lt;/STRONG&gt;and &lt;STRONG&gt;COLWIDTH; &lt;/STRONG&gt;thus&amp;nbsp;I wonder&amp;nbsp;if line 83 maybe should be:&lt;/P&gt;

&lt;P&gt;c&lt;K&gt;&lt;J&gt; += a&lt;K&gt;&lt;J&gt; * b&lt;K&gt;&lt;J&gt;;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/J&gt;&lt;/K&gt;&lt;/P&gt;

&lt;P&gt;&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;/P&gt;

&lt;P&gt;Thanks for answering. Maybe I didn't understood your question, but I was intending exactly that (the order i-k-j is only the usual optimization of the naive i-j-k algorithm for matrix multiplication). The names for the loop boundaries you're referring to are legacy names inherited by the intel sample vectorization code in the composerxe folder, from which I started after my old code started to run very slowly.&lt;/P&gt;

&lt;P&gt;&lt;/P&gt;&lt;BLOCKQUOTE&gt;Kevin Davis (Intel) wrote:&lt;BR /&gt;&lt;P&gt;&lt;/P&gt;

&lt;P&gt;I see a difference in execution times with dynamic allocation when running natively on the coprocessor so I will investigate further and post again after I know more.&lt;/P&gt;

&lt;P&gt;&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;/P&gt;

&lt;P&gt;I'm starting to think that my icc compiler isn't working well, is it possibile that it doesn't align the data?&lt;/P&gt;

&lt;P&gt;Am I missing something?&lt;BR /&gt;
	&lt;BR /&gt;
	P.S. the #define ALIGNMENT on top is set to 64, not to 128, I pasted wrong.&lt;/P&gt;</description>
      <pubDate>Fri, 24 Oct 2014 15:58:00 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033548#M43041</guid>
      <dc:creator>Luca_A_</dc:creator>
      <dc:date>2014-10-24T15:58:00Z</dc:date>
    </item>
    <item>
      <title>I was just noting where ROWS</title>
      <link>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033549#M43042</link>
      <description>&lt;P&gt;I was just noting where &lt;STRONG&gt;ROWS =&lt;/STRONG&gt; &lt;STRONG&gt;REPEATNTIMES &lt;/STRONG&gt;there isn't a concern, but where &lt;STRONG&gt;REPEATNTIMES &lt;/STRONG&gt;&amp;gt; &lt;STRONG&gt;ROWS&lt;/STRONG&gt;&lt;STRONG&gt; &lt;/STRONG&gt;the loop accesses beyond the array row dimension for a and c.&lt;/P&gt;</description>
      <pubDate>Fri, 24 Oct 2014 17:26:23 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/Dynamic-allocation-problems-on-Xeon-Phi/m-p/1033549#M43042</guid>
      <dc:creator>Kevin_D_Intel</dc:creator>
      <dc:date>2014-10-24T17:26:23Z</dc:date>
    </item>
  </channel>
</rss>

