<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Speeding up IPP FFT in Intel® Integrated Performance Primitives</title>
    <link>https://community.intel.com/t5/Intel-Integrated-Performance/Speeding-up-IPP-FFT/m-p/904692#M13270</link>
    <description>...my mistake, I was using different normalizations, the performances are comparable now.&lt;BR /&gt;&lt;BR /&gt;Apart from this mistake, any advice or trick speeding up IPP-FFT is very wellcome!&lt;BR /&gt;&lt;BR /&gt;Francesco&lt;BR /&gt;</description>
    <pubDate>Mon, 16 Feb 2009 18:01:14 GMT</pubDate>
    <dc:creator>fbasile</dc:creator>
    <dc:date>2009-02-16T18:01:14Z</dc:date>
    <item>
      <title>Speeding up IPP FFT</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/Speeding-up-IPP-FFT/m-p/904691#M13269</link>
      <description>I'm benchmarking different implementations of 1D-FFT, and I've written a simple code for testing IPP:&lt;BR /&gt;&lt;SPAN class="sectionBodyText"&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;SPAN class="sectionBodyText"&gt;#include &lt;IOSTREAM&gt;&lt;BR /&gt;#include &lt;IPPS.H&gt;&lt;BR /&gt;#include &lt;SYS&gt;&lt;BR /&gt;&lt;BR /&gt;int main() {&lt;BR /&gt; for (unsigned int myOrder = 10; myOrder &amp;lt; 21; myOrder++) {&lt;BR /&gt; const unsigned int myLength = (1 &amp;lt;&amp;lt; myOrder);&lt;BR /&gt; std::cout &amp;lt;&amp;lt; "Length = " &amp;lt;&amp;lt; myLength &amp;lt;&amp;lt; std::endl;&lt;BR /&gt;&lt;BR /&gt; IppsFFTSpec_C_32fc* mySpec;&lt;BR /&gt; ippsFFTInitAlloc_C_32fc(&amp;amp;mySpec, myOrder, IPP_FFT_NODIV_BY_ANY, ippAlgHintFast);&lt;BR /&gt;&lt;BR /&gt; int myBufferSize = 0;&lt;BR /&gt; ippsFFTGetBufSize_C_32fc(mySpec, &amp;amp;myBufferSize);&lt;BR /&gt; Ipp8u *myBuffer = ippsMalloc_8u(myBufferSize);&lt;BR /&gt;&lt;BR /&gt; Ipp32fc *myA = ippsMalloc_32fc(myLength);&lt;BR /&gt; Ipp32fc *myB = ippsMalloc_32fc(myLength);&lt;BR /&gt; for(unsigned int n = 0; n &amp;lt; myLength; ++n) {&lt;BR /&gt; myA&lt;N&gt;.re = rand()/RAND_MAX;&lt;BR /&gt; myA&lt;N&gt;.im = rand()/RAND_MAX;&lt;BR /&gt; }&lt;BR /&gt;&lt;BR /&gt; struct timeval myStartingTime;&lt;BR /&gt; struct timeval myEndingTime;&lt;BR /&gt; gettimeofday(&amp;amp;myStartingTime, 0);&lt;BR /&gt;&lt;BR /&gt; for (unsigned int myRepetition = 0; myRepetition &amp;lt; 1000; ++myRepetition) {&lt;BR /&gt; ippsFFTFwd_CToC_32fc(myA, myB, mySpec, myBuffer);&lt;BR /&gt; }&lt;BR /&gt;&lt;BR /&gt; gettimeofday(&amp;amp;myEndingTime, 0);&lt;BR /&gt; double myElapsedSeconds = myEndingTime.tv_sec - myStartingTime.tv_sec +&lt;BR /&gt; (myEndingTime.tv_usec - myStartingTime.tv_usec) / 1000000.0;&lt;BR /&gt; std::cout &amp;lt;&amp;lt; "\t\testimated Mflops: "&lt;BR /&gt; &amp;lt;&amp;lt; 2.5 * myOrder * myLength / myElapsedSeconds / 1000&lt;BR /&gt; &amp;lt;&amp;lt; std::endl;&lt;BR /&gt;&lt;BR /&gt; ippsFree(myA);&lt;BR /&gt; ippsFree(myB);&lt;BR /&gt; ippsFree(myBuffer);&lt;BR /&gt; ippsFFTFree_C_32fc(mySpec);&lt;BR /&gt; }&lt;BR /&gt;}&lt;BR /&gt;&lt;/N&gt;&lt;/N&gt;&lt;/SYS&gt;&lt;/IPPS.H&gt;&lt;/IOSTREAM&gt;&lt;/SPAN&gt;&lt;BR /&gt;And I compile it with icc.&lt;BR /&gt;&lt;BR /&gt;It seems to be drammatically slower then FFTW3 implementation, BUT I'm not happy with this result: I've found benchmarks (http://www.fftw.org/speed/CoreDuo-3.0GHz-icc64/) showing performances to be similar.&lt;BR /&gt;&lt;BR /&gt;I'm afraid (pretty sure) I'm not fully exploiting IPP. How should I speed up this code?&lt;BR /&gt;&lt;BR /&gt;Thanks&lt;BR /&gt;francesco</description>
      <pubDate>Fri, 13 Feb 2009 16:47:36 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/Speeding-up-IPP-FFT/m-p/904691#M13269</guid>
      <dc:creator>fbasile</dc:creator>
      <dc:date>2009-02-13T16:47:36Z</dc:date>
    </item>
    <item>
      <title>Re: Speeding up IPP FFT</title>
      <link>https://community.intel.com/t5/Intel-Integrated-Performance/Speeding-up-IPP-FFT/m-p/904692#M13270</link>
      <description>...my mistake, I was using different normalizations, the performances are comparable now.&lt;BR /&gt;&lt;BR /&gt;Apart from this mistake, any advice or trick speeding up IPP-FFT is very wellcome!&lt;BR /&gt;&lt;BR /&gt;Francesco&lt;BR /&gt;</description>
      <pubDate>Mon, 16 Feb 2009 18:01:14 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-Integrated-Performance/Speeding-up-IPP-FFT/m-p/904692#M13270</guid>
      <dc:creator>fbasile</dc:creator>
      <dc:date>2009-02-16T18:01:14Z</dc:date>
    </item>
  </channel>
</rss>

