<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Hello, Zvika! in Intel® Integrated Performance Primitives</title>
    <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179732#M27142</link>
    <description>&lt;P&gt;Hello, Zvika!&lt;/P&gt;&lt;P&gt;The "srcStep" and "dstStep" parameters are in bytes. The Ipp8u data type is 8-bit (1 byte) long, so the code in example works as intended. The Ipp32f data type is 32-bit, i.e.&amp;nbsp;4 bytes, long. So you should multiply the distance in elements by sizeof(Ipp32f).&lt;/P&gt;&lt;P&gt;Here is an example:&lt;/P&gt;
&lt;PRE class="brush:cpp; class-name:dark;"&gt;Ipp32f src[8 * 4] = { 1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8 };
Ipp32f dst[4 * 4];
IppiSize srcRoi = { 4, 4 };
ippiTranspose_32f_C1R(src, 8*sizeof(Ipp32f), dst, 4*sizeof(Ipp32f), srcRoi);&lt;/PRE&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Always glad to help.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Best regards,&lt;/P&gt;
&lt;P&gt;Ivan Galanin.&lt;/P&gt;</description>
    <pubDate>Wed, 10 Jul 2019 07:07:02 GMT</pubDate>
    <dc:creator>Ivan_G_Intel1</dc:creator>
    <dc:date>2019-07-10T07:07:02Z</dc:date>
    <item>
      <title>ippiTranspose_32f_C1R: Wrong output</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179731#M27141</link>
      <description>&lt;P&gt;Hello,&lt;/P&gt;&lt;P&gt;Based on Intel's example, I used the following code:&lt;/P&gt;&lt;P&gt;Ipp32f&amp;nbsp;src[8*4] = {1, 2, 3, 4, 8, 8, 8, 8,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 1, 2, 3, 4, 8, 8, 8, 8,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 1, 2, 3, 4, 8, 8, 8, 8,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; 1, 2, 3, 4, 8, 8, 8, 8};&lt;/P&gt;&lt;P&gt;Ipp32f&amp;nbsp;dst[4*4];&lt;/P&gt;&lt;P&gt;IppiSize srcRoi = { 4, 4 };&lt;/P&gt;&lt;P&gt;ippiTranspose_32f_C1R ( src, 8, dst, 4, srcRoi );&lt;/P&gt;&lt;P&gt;The output is:&lt;/P&gt;&lt;P&gt;{ 1, 2, 3, 4,&lt;/P&gt;&lt;P&gt;&amp;nbsp; 8, 8, 2, 0xCCCCCCCC,&lt;/P&gt;&lt;P&gt;&amp;nbsp; 0xCCCCCCC, 0xCCCCCC, 0xCCCCCCCC, 0xCCCCCCCC,&lt;/P&gt;&lt;P&gt;&amp;nbsp; 0xCCCCCCC, 0xCCCCCC, 0xCCCCCCCC, 0xCCCCCCCC}&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Can you please explain what is wrong in my code ?&lt;/P&gt;&lt;P&gt;The original code is using Ipp8u.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Thank you,&lt;/P&gt;&lt;P&gt;Zvika&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 05:12:47 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179731#M27141</guid>
      <dc:creator>ZVere</dc:creator>
      <dc:date>2019-07-10T05:12:47Z</dc:date>
    </item>
    <item>
      <title>Hello, Zvika!</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179732#M27142</link>
      <description>&lt;P&gt;Hello, Zvika!&lt;/P&gt;&lt;P&gt;The "srcStep" and "dstStep" parameters are in bytes. The Ipp8u data type is 8-bit (1 byte) long, so the code in example works as intended. The Ipp32f data type is 32-bit, i.e.&amp;nbsp;4 bytes, long. So you should multiply the distance in elements by sizeof(Ipp32f).&lt;/P&gt;&lt;P&gt;Here is an example:&lt;/P&gt;
&lt;PRE class="brush:cpp; class-name:dark;"&gt;Ipp32f src[8 * 4] = { 1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8,
&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp; &amp;nbsp;&amp;nbsp;&amp;nbsp;1, 2, 3, 4, 8, 8, 8, 8 };
Ipp32f dst[4 * 4];
IppiSize srcRoi = { 4, 4 };
ippiTranspose_32f_C1R(src, 8*sizeof(Ipp32f), dst, 4*sizeof(Ipp32f), srcRoi);&lt;/PRE&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Always glad to help.&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Best regards,&lt;/P&gt;
&lt;P&gt;Ivan Galanin.&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 07:07:02 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179732#M27142</guid>
      <dc:creator>Ivan_G_Intel1</dc:creator>
      <dc:date>2019-07-10T07:07:02Z</dc:date>
    </item>
    <item>
      <title>Hi Ivan,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179733#M27143</link>
      <description>&lt;P&gt;Hi Ivan,&lt;/P&gt;&lt;P&gt;Thank you very much. Works great !&lt;/P&gt;&lt;P&gt;In case I need an&lt;STRONG&gt; in-place transpose&lt;/STRONG&gt;, with:&lt;STRONG&gt; ippiTranspose_32f_C1IR&lt;/STRONG&gt;, I can use only src [4 * 4], roi [4, 4]&lt;/P&gt;&lt;P&gt;Am I right ?&lt;/P&gt;&lt;P&gt;Best regards,&lt;/P&gt;&lt;P&gt;Zvika&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 07:17:10 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179733#M27143</guid>
      <dc:creator>ZVere</dc:creator>
      <dc:date>2019-07-10T07:17:10Z</dc:date>
    </item>
    <item>
      <title>Zvika,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179734#M27144</link>
      <description>&lt;P&gt;Zvika,&lt;/P&gt;&lt;P&gt;For in-place operations,&amp;nbsp;roiSize.width&amp;nbsp;must be equal to&amp;nbsp;roiSize.height (meaning roi must be a square). src width and height equality is not required.&lt;/P&gt;&lt;P&gt;So, even for in-place operation src [8&amp;nbsp;* 4] is correct,&amp;nbsp;but roi [8 * 4] will be incorrect.&lt;/P&gt;&lt;P&gt;Always glad to help.&lt;/P&gt;&lt;P&gt;Best regards,&lt;BR /&gt;Ivan Galanin.&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 07:33:13 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179734#M27144</guid>
      <dc:creator>Ivan_G_Intel1</dc:creator>
      <dc:date>2019-07-10T07:33:13Z</dc:date>
    </item>
    <item>
      <title>Hi Ivan,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179735#M27145</link>
      <description>&lt;P&gt;Hi Ivan,&lt;/P&gt;&lt;P&gt;My src has 512 rows X&amp;nbsp;4096 columns of float Ipp32f . A row is consecutive in RAM.&amp;nbsp;&lt;/P&gt;&lt;P&gt;I want to transpose the left up corner of this matrix which is: 379 rows X 2400 columns.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Ipp32f src [4096 * 512];&lt;/P&gt;&lt;P&gt;Ipp32f dst[4096 * 512];&lt;/P&gt;&lt;P&gt;IppiSize srcRoi = { 2400, 379&amp;nbsp;};&lt;/P&gt;&lt;P&gt;ippiTranspose_32f_C1R(src,&amp;nbsp; 4096*sizeof(Ipp32f), dst, 4096*sizeof(Ipp32f), srcRoi);&lt;/P&gt;&lt;P&gt;I got exception in&amp;nbsp;ippiTranspose_32f_C1R.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Can you please tell what is wrong in my code ?&lt;/P&gt;&lt;P&gt;Thank you,&lt;/P&gt;&lt;P&gt;Zvika&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 09:14:02 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179735#M27145</guid>
      <dc:creator>ZVere</dc:creator>
      <dc:date>2019-07-10T09:14:02Z</dc:date>
    </item>
    <item>
      <title>Hi Ivan, All,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179736#M27146</link>
      <description>&lt;P&gt;Hi Ivan, All,&lt;/P&gt;&lt;P&gt;I think I found my mistake:&lt;/P&gt;&lt;P&gt;In order to transpose 2D matrix with:&amp;nbsp;&lt;/P&gt;&lt;P&gt;#define COLS 2400&lt;/P&gt;&lt;P&gt;#define ROWS 379&lt;/P&gt;&lt;P&gt;ippiTranspose_32f_C1R(src, COLS * sizeof(Ipp32f), dst, &lt;STRONG&gt;ROWS &lt;/STRONG&gt;* sizeof(Ipp32f), srcRoi);&lt;/P&gt;&lt;P&gt;Thank you,&lt;/P&gt;&lt;P&gt;Zvika&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 10 Jul 2019 10:58:00 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179736#M27146</guid>
      <dc:creator>ZVere</dc:creator>
      <dc:date>2019-07-10T10:58:00Z</dc:date>
    </item>
    <item>
      <title>Hello, Zvika,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179737#M27147</link>
      <description>&lt;P&gt;Hello, Zvika,&lt;/P&gt;&lt;P&gt;The parameters&amp;nbsp;srcStep and&amp;nbsp;dstStep are distances, in bytes, between the starting points of consecutive lines in the source and destination images/matrices&amp;nbsp;respectively (i.e. columns count multiplied by the size of a data type). In the previous message your code was:&amp;nbsp;&lt;/P&gt;
&lt;PRE class="brush:cpp; class-name:dark;"&gt;ippiTranspose_32f_C1R(src,  4096*sizeof(Ipp32f), dst, 4096*sizeof(Ipp32f), srcRoi);&lt;/PRE&gt;

&lt;P&gt;Which is correct if source and destination images/matrices have 4096 columns.&lt;BR /&gt;The ROI can be smaller than the whole image/matrix, but in this case you still have to use&amp;nbsp;image/matrix width in a calculation of srcStep and dstStep.&lt;BR /&gt;&lt;BR /&gt;The problem is that your destination height is smaller than the ROI width (it's not about memory amount you use, but how you use it and how you and the function handle it). Formally, the "srcRoi.width" number of columns are copied to rows in a resulting image/matrix (2400 into 512).&lt;BR /&gt;Because of that the result of the transpose is unpredictable for different ROI dimensions and can result in an error (as it does).&lt;/P&gt;
&lt;P&gt;Here is an example&amp;nbsp;that works:&lt;/P&gt;

&lt;PRE class="brush:cpp; class-name:dark;"&gt;	Ipp32f* src = ippsMalloc_32f(4096 * 512 * sizeof(Ipp32f));
&amp;nbsp;   //standard memory allocation, same for the source
	Ipp32f* dst = ippsMalloc_32f(4096 * 512 * sizeof(Ipp32f));
	IppiSize srcRoi = { 2400, 379 };

    //handle the memory as if the destination image dimensions were 512 columns and 4096 rows
	ippiTranspose_32f_C1R(src, 4096 * sizeof(Ipp32f), dst, 512 * sizeof(Ipp32f), srcRoi);

	ippsFree(src);
	ippsFree(dst);&lt;/PRE&gt;

&lt;P&gt;here is more optimized version that can be more illustrative:&lt;/P&gt;

&lt;PRE class="brush:cpp; class-name:dark;"&gt;	int srcStep, dstStep;
	// The values of srcStep and dstStep are calculated automatically
	// by ippiMalloc functions, if you check them, they will be
	// 4096 * sizeof(Ipp32f) and 
	// 512 * sizeof(Ipp32f) respectively
	Ipp32f* src = ippiMalloc_32f_C1(4096, 512, &amp;amp;srcStep);
	Ipp32f* dst = ippiMalloc_32f_C1(512, 4096, &amp;amp;dstStep);
	IppiSize srcRoi = { 2400, 379 };

	ippiTranspose_32f_C1R(src, srcStep, dst, dstStep, srcRoi);

	ippiFree(src);
	ippiFree(dst);&lt;/PRE&gt;

&lt;P&gt;So, the destination height must be greater than or equal to your ROI width.&lt;/P&gt;
&lt;P&gt;Always glad to help.&lt;/P&gt;
&lt;P&gt;Best regards,&lt;BR /&gt;Ivan Galanin.&lt;/P&gt;</description>
      <pubDate>Fri, 12 Jul 2019 11:09:00 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179737#M27147</guid>
      <dc:creator>Ivan_G_Intel1</dc:creator>
      <dc:date>2019-07-12T11:09:00Z</dc:date>
    </item>
    <item>
      <title>Hi Ivan,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179738#M27148</link>
      <description>&lt;P&gt;Hi Ivan,&lt;/P&gt;&lt;P&gt;Thank you very much for detailed explanation.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Currently, there is no "ippiTranspose_32f_C2R".&amp;nbsp;&lt;/P&gt;&lt;P&gt;We need it to transpose a 2D complex float matrix. Our signal processing is heavily based on complex numbers.&amp;nbsp;&lt;/P&gt;&lt;P&gt;I wrote such transpose but my code is "naive". I'm sure it can be optimized.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Can Intel consider developing an "ippiTranspose_32f_C2R" ?&lt;/P&gt;&lt;P&gt;Or - If I will get the&amp;nbsp;ippiTranspose_32f_C1R&amp;nbsp; maybe I can port the code to&amp;nbsp;ippiTranspose_32f_C2R.&amp;nbsp;&lt;/P&gt;&lt;P&gt;Best regards,&lt;/P&gt;&lt;P&gt;Zvika&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Fri, 12 Jul 2019 16:02:00 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179738#M27148</guid>
      <dc:creator>ZVere</dc:creator>
      <dc:date>2019-07-12T16:02:00Z</dc:date>
    </item>
    <item>
      <title>Hi Zvi,</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179739#M27149</link>
      <description>&lt;P&gt;Hi Zvi,&lt;/P&gt;&lt;P&gt;in IPP notation your request will be "to develop ippiTranspose_32fc_C1R", that is the same as ippiTranspose_64f_C1R&amp;nbsp;- that we will consider as a feature request.&amp;nbsp;We deprecated and removed "complex" image support in IPP 9.0, therefore the flavor will be 64f_C1R - it can be easily type-casted to 32fc_C1R.&lt;/P&gt;&lt;P&gt;regards, Igor&lt;/P&gt;</description>
      <pubDate>Mon, 15 Jul 2019 08:05:29 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/ippiTranspose-32f-C1R-Wrong-output/m-p/1179739#M27149</guid>
      <dc:creator>Igor_A_Intel</dc:creator>
      <dc:date>2019-07-15T08:05:29Z</dc:date>
    </item>
  </channel>
</rss>

