<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Hi I repeated the experiment in Software Archive</title>
    <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024121#M39398</link>
    <description>&lt;P&gt;Hi I repeated the experiment many times.&lt;/P&gt;

&lt;P&gt;The numbers are still in favour of xeon.&lt;/P&gt;

&lt;P&gt;3.9 ms for xeon&lt;/P&gt;

&lt;P&gt;4.0 for xeon phi.&lt;/P&gt;

&lt;P&gt;Any idea how to improve on it?&lt;/P&gt;</description>
    <pubDate>Mon, 19 Oct 2015 10:35:04 GMT</pubDate>
    <dc:creator>aketh_t_</dc:creator>
    <dc:date>2015-10-19T10:35:04Z</dc:date>
    <item>
      <title>First touch time greater than parallel time</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024118#M39395</link>
      <description>&lt;P&gt;Hi all,&lt;/P&gt;

&lt;P&gt;I was looking to parallelize my code for speedup.&lt;/P&gt;

&lt;P&gt;As xeon phi was a NUMA core I used the first touch placement of the data.&lt;/P&gt;

&lt;P&gt;while xeon phi is performing better than xeon no doubt, the problem is that totaltime(time for first touch+looptime) is greater.&lt;/P&gt;

&lt;P&gt;How do I resolve this issue?&lt;/P&gt;

&lt;P&gt;This code when integrated into the main code(cannot post it here) will call state function many times from various different places. So is it possible that even if I dont first touch as I have in the code attached below this overhead is just a onetime problem?&lt;/P&gt;

&lt;P&gt;The code attached below as state_test_offload is for MIC and state_test is for Xeon host.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 09:34:20 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024118#M39395</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T09:34:20Z</dc:date>
    </item>
    <item>
      <title>During the call of state in</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024119#M39396</link>
      <description>&lt;P&gt;During the call of state in the actual program, state get called at different places with TEMPK and SALTK having values different at each call.&lt;BR /&gt;
	&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 09:38:12 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024119#M39396</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T09:38:12Z</dc:date>
    </item>
    <item>
      <title>As xeon phi was a NUMA core I</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024120#M39397</link>
      <description>&lt;BLOCKQUOTE&gt;
	&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;As xeon phi was a NUMA core I used the first touch placement of the data.&lt;/SPAN&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Actually, KNC does not have much NUMA sensitivity since it is a single chip, so the time for any core to reach the memory controllers is similar. (Unlike a multi-socket Xeon where the time to reach a memory controller on another chip is large because the request has to cross the external QPI coherence fabric).&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;However for portability using first-touch still makes sense.&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;The overhead is almost certainly the cost of the kernel actually instantiating the page (mapping the virtual address to real, physical, memory) when it is first written to. (Allocating a physical page and zeroing it). You will pay that cost whichever thread first writes to the memory, but, you'll only pay it once. (Since after the page is written it will have memory allocated to it).&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;So&lt;/SPAN&gt;&lt;/P&gt;

&lt;OL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Using a first touch policy is sensible for performance portability&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;Overall it shouldn't cost any more than touching all the data in one place (you have to pay the cost of instantiating the page somewhere!)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;You should be able to time the different loops and observe that the first loopm which writes to all the pages is slow, but subsequent loops are faster. (You can also observe this by looking at the number of page-faults, since the copy-on-write which instantiates the page is entered as a result of a page-fault).&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;You can, of course, reduce the number of page-faults by using big pages (I haven't measured whether &amp;nbsp;that reduces the overall time significantly or not; if the time is mostly spent in zeroing the memory the number of page faults won't make much difference...)&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;HTH&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 09:53:34 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024120#M39397</guid>
      <dc:creator>James_C_Intel2</dc:creator>
      <dc:date>2015-10-19T09:53:34Z</dc:date>
    </item>
    <item>
      <title>Hi I repeated the experiment</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024121#M39398</link>
      <description>&lt;P&gt;Hi I repeated the experiment many times.&lt;/P&gt;

&lt;P&gt;The numbers are still in favour of xeon.&lt;/P&gt;

&lt;P&gt;3.9 ms for xeon&lt;/P&gt;

&lt;P&gt;4.0 for xeon phi.&lt;/P&gt;

&lt;P&gt;Any idea how to improve on it?&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 10:35:04 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024121#M39398</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T10:35:04Z</dc:date>
    </item>
    <item>
      <title>sorry its actually 2.5ms for</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024122#M39399</link>
      <description>&lt;P&gt;sorry its actually 2.5ms for xeon.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 10:39:27 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024122#M39399</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T10:39:27Z</dc:date>
    </item>
    <item>
      <title>I mean I added a loop over</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024123#M39400</link>
      <description>&lt;P&gt;I mean I added a loop over call state function, when I mean repeated many times.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 10:54:39 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024123#M39400</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T10:54:39Z</dc:date>
    </item>
    <item>
      <title>I don't think this has</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024124#M39401</link>
      <description>&lt;P&gt;I don't think this has anything to do with first touch.&lt;/P&gt;

&lt;P&gt;Have you looked at the time to transfer data?&amp;nbsp;https://software.intel.com/en-us/node/524675 should help.&lt;/P&gt;

&lt;P&gt;p.s. Forcibly setting the number of threads to 240 is a bad idea (better than forcing 244, but still bad). By default the runtime will do the right thing, and, even better, it will also do the right thing in the future when you run on a machine (like Knights Landing) which has a different optimal number of threads.&lt;/P&gt;

&lt;P&gt;In this context, it's also worth experimenting with KMP_PLACE_THREADS (&amp;nbsp;https://software.intel.com/en-us/node/512745 ) to see whether using only one or two threads/core improves your performance (KMP_PLACE_THREADS=1t or KMP_PLACE_THREADS=2t plus whatever offload prefix you need to have the envirable propagated, of course).&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 11:19:21 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024124#M39401</guid>
      <dc:creator>James_C_Intel2</dc:creator>
      <dc:date>2015-10-19T11:19:21Z</dc:date>
    </item>
    <item>
      <title>Right now I am not concerned</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024125#M39402</link>
      <description>&lt;P&gt;Right now I am not concerned with transfer time. I think I can asynchronously do a few routines.&lt;/P&gt;

&lt;P&gt;By Default how do I know how many threads to set? it would just use whats in the OMP_NUM_THREADS or PHI_OMP_NUM_THREADS right?&lt;BR /&gt;
	&lt;BR /&gt;
	Will try the last approach.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 11:29:53 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024125#M39402</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T11:29:53Z</dc:date>
    </item>
    <item>
      <title>By Default how do I know how</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024126#M39403</link>
      <description>&lt;BLOCKQUOTE&gt;
	&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;By Default how do I know how many threads to set? it would just use whats in the OMP_NUM_THREADS or PHI_OMP_NUM_THREADS right?&lt;/SPAN&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;You don't need to set anything. Just let the combination of the offload code and the OpenMP runtime do the right thing. As soon as you force values into your code&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;

&lt;UL&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;It becomes non-portable. (Either to someone else's machine where they may have a different number of cores in their KNC, or to the machine you buy in a few years, after you've forgotten that there's this magic constant in the code).&lt;/SPAN&gt;&lt;/LI&gt;
	&lt;LI&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;It becomes much harder to do a scaling study (you'd have to recompile the code to try each different number of cores)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;You can change the number of threads that will be used by setting OMP_NUM_THREADS for the host side and the same envirable with the appropriate prefix (I don't know if that is "PHI_" or "MIC_". or what; you can choose it somehow :-)). Similarly use KMP_PLACE_THREADS (with the appropriate prefix).&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;By all means use omp_get_max_threads() in the offload to see what you're going to use, (in the serial offload code), and print it, but actively forcing 240 as a constant is nor recommended.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 11:44:06 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024126#M39403</guid>
      <dc:creator>James_C_Intel2</dc:creator>
      <dc:date>2015-10-19T11:44:06Z</dc:date>
    </item>
    <item>
      <title>I ran the code with the the</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024127#M39404</link>
      <description>&lt;P&gt;I ran the code with the the setting of threads removed its performing badly?&lt;/P&gt;

&lt;P&gt;its now 60ms&lt;/P&gt;

&lt;P&gt;Should I set PHI_OMP_NUM_THREADS or unset that?&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 11:48:34 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024127#M39404</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T11:48:34Z</dc:date>
    </item>
    <item>
      <title>Assuming you have set MIC_ENV</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024128#M39405</link>
      <description>&lt;P&gt;Assuming you have set MIC_ENV_PREFIX=PHI, you should try PHI_KMP_PLACE_THREADS settings as James suggested.&amp;nbsp; This implies a consistent setting of num_threads.&amp;nbsp; In my experience, where MIC native OpenMP reached peak performance with KMP_PLACE_THREADS=59c,2t and OMP_PROC_BIND=close, offload didn't benefit beyond PHI_KMP_PLACE_THREADS=59c,1t.&amp;nbsp; It's particularly important, as James said, not to force worker threads onto the core which runs MPSS and data transfers.&amp;nbsp; Running more than 1 thread per MIC core typically doesn't benefit, even in native mode, without thread placement so as to share the cache effectively.&amp;nbsp; Page initialization overhead, as James mentioned, also contributes to offload not being able to use as many threads effectively.&lt;/P&gt;

&lt;P&gt;If you set OMP_NUM_THREADS for host, MIC used to inherit that setting when you didn't set an appropriate value, so you have no chance if you don't set appropriate values for MIC.&lt;/P&gt;

&lt;P&gt;The micsmc gui is an important tool to visualize whether you have distributed threads efficiently on MIC.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 12:25:18 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024128#M39405</guid>
      <dc:creator>TimP</dc:creator>
      <dc:date>2015-10-19T12:25:18Z</dc:date>
    </item>
    <item>
      <title>hi,</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024129#M39406</link>
      <description>&lt;P&gt;hi,&lt;/P&gt;

&lt;P&gt;This is what I did&lt;/P&gt;

&lt;P&gt;1)unset PHI_OMP_NUM_THREADS&lt;/P&gt;

&lt;P&gt;2)tried for PHI_KMP_PLACE_THREADS=30c,2t seems to work decently 17ms&lt;/P&gt;

&lt;P&gt;3)59c,2t is bad even with OMP_PROC_BIND performance isnt as good as xeon 23ms&lt;/P&gt;

&lt;P&gt;any idea if I could find out hotspots/inneffciency using Vtune?&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 12:56:39 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024129#M39406</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T12:56:39Z</dc:date>
    </item>
    <item>
      <title>which metric should I be</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024130#M39407</link>
      <description>&lt;P&gt;which metric should I be looking at?&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 13:02:24 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024130#M39407</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T13:02:24Z</dc:date>
    </item>
    <item>
      <title>any idea if I could find out</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024131#M39408</link>
      <description>&lt;BLOCKQUOTE&gt;
	&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;any idea if I could find out hotspots/inneffciency using Vtune?&lt;/SPAN&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;You can. Even the &lt;A href="https://software.intel.com/en-us/intel-vtune-amplifier-xe"&gt;VTune product landing page &lt;/A&gt;is showing that&amp;nbsp;&lt;/SPAN&gt;and my fourth Google hit for the query&amp;nbsp;&lt;STRONG&gt;intel vtune amplifier openmp tuning i&lt;/STRONG&gt;&lt;SPAN style="font-size: 1em; line-height: 1.5;"&gt;s&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN style="font-size: 1em; line-height: 1.5;"&gt;&lt;A href="https://software.intel.com/en-us/articles/how-to-analyze-openmp-applications-using-intel-vtune-amplifie-xe-2015" target="_blank"&gt;https://software.intel.com/en-us/articles/how-to-analyze-openmp-applications-using-intel-vtune-amplifie-xe-2015&lt;/A&gt; which seems to have a step by step example for you.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 13:11:14 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024131#M39408</guid>
      <dc:creator>James_C_Intel2</dc:creator>
      <dc:date>2015-10-19T13:11:14Z</dc:date>
    </item>
    <item>
      <title>sorry for being so imprecise</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024132#M39409</link>
      <description>&lt;P&gt;sorry for being so imprecise on my query.&lt;/P&gt;

&lt;P&gt;I apologise.&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 13:30:52 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024132#M39409</guid>
      <dc:creator>aketh_t_</dc:creator>
      <dc:date>2015-10-19T13:30:52Z</dc:date>
    </item>
    <item>
      <title>I apologise</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024133#M39410</link>
      <description>&lt;BLOCKQUOTE&gt;
	&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;I apologise&lt;/SPAN&gt;&lt;/P&gt;
&lt;/BLOCKQUOTE&gt;

&lt;P&gt;&lt;SPAN style="font-size: 12px; line-height: 18px;"&gt;No problem, everyone's Google-fu runs out sometimes :-)&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 13:34:47 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024133#M39410</guid>
      <dc:creator>James_C_Intel2</dc:creator>
      <dc:date>2015-10-19T13:34:47Z</dc:date>
    </item>
    <item>
      <title>I think this may shed some</title>
      <link>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024134#M39411</link>
      <description>&lt;P&gt;I think this may shed some light on where the time is spent:&lt;/P&gt;

&lt;PRE class="brush:fortran;"&gt;!dir$ offload begin target(mic:0)

begin_time = omp_get_wtime()

call omp_set_num_threads(240)

!$omp parallel
dum_time = omp_get_wtime() ! dummy do something for firts parallel region
!$omp end parallel

first_time = omp_get_wtime() 

...
!$omp end parallel&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp;&amp;nbsp; 
end_time = omp_get_wtime()

!dir$ end offload&amp;nbsp;&amp;nbsp; 
&amp;nbsp;
print *,"OpenMP thread pool init time is",first_time - begin_time
print *,"loop time is", end_time - start_time
print *,"total compute time is",end_time - first_time
print *,"total time is", end_time - begin_time
&lt;/PRE&gt;

&lt;P&gt;You will find that starting the initial thread pool on the KNC has significant overhead. Therefore, when programming in offload model on&amp;nbsp; KNC, it is generally beneficial at program start (host side) to call a subroutine (or place in line as first few statements), code that ostensibly does no work other than initialize the KNC OpenMP thread team.&lt;/P&gt;

&lt;P&gt;RE: First Touch on KNC&lt;/P&gt;

&lt;P&gt;As James stated,on KNC&amp;nbsp;any code you insert for "first touch" latency is not used for NUMA locality is not beneficial, but doesn't hurt, and when multi-socket KNL arrives, it will be beneficial. What does matter with respect to latency, is after process start, while the process may have many GB of Virtual Memory, only that memory that has been "first touched" (page granularity) since process start is mapped to physical RAM (and/or Page File). Therefore the first time allocation from heap may encounter none/one/several/many page faults as each page in the allocation is "first touched" during use. This overhead is quite significant.&lt;/P&gt;

&lt;P&gt;On KNC you still have "first touch" latency as James stated with respect to first time a page in the process VM is touched after allocation from the heap. This is a one-time latency overhead. In writing applications you might want to take this into consideration when timing.&lt;/P&gt;

&lt;P&gt;In a real-world performance critical application you typically are not interested in:&lt;/P&gt;

&lt;P&gt;Time from double-click on program Icon to program completion time.&lt;/P&gt;

&lt;P&gt;What is typically of concern is the time that an internal process loop takes _after_ everything has been initialized. As an example you have a simulation that will iterate millions of times. Your interest is in total time, not time for first iteration.&lt;/P&gt;

&lt;P&gt;If you go to &lt;A href="http://www.lotsofcores.com/"&gt;http://www.lotsofcores.com/&lt;/A&gt; and scroll down a little way you will see a graphical video comparison of a simulation run using two different programming strategies. If you Pause the simulation and drag the horizontal scroll bar to time 0:00:07 you will find&lt;/P&gt;

&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper" image-alt="Time_07.jpg"&gt;&lt;img src="https://community.intel.com/t5/image/serverpage/image-id/8136i3FF98FCBDBC773C5/image-size/large?v=v2&amp;amp;px=999&amp;amp;whitelist-exif-data=Orientation%2CResolution%2COriginalDefaultFinalSize%2CCopyright" role="button" title="Time_07.jpg" alt="Time_07.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;

&lt;P&gt;On the left pane is a typical tiled implementation using rectangular tiles. The different sizes of the tiles painted blue&amp;nbsp;indicate the different completion states of each tile as run by different threads (240). The relatively large amount of skew in tile size is attributable to OpenMP thread pool initialization, as well as thread skew on getting threads started in parallel region. Right right pane shows an alternate algorithm using columnar tiles. At the 7 second time step, both sides, some threads have yet to produce output.&lt;/P&gt;

&lt;P&gt;At time 0:01:03, after "first touch" and several iterations:&lt;/P&gt;

&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper" image-alt="Time_103.jpg"&gt;&lt;img src="https://community.intel.com/t5/image/serverpage/image-id/8137i79B26C06D9AE2402/image-size/large?v=v2&amp;amp;px=999&amp;amp;whitelist-exif-data=Orientation%2CResolution%2COriginalDefaultFinalSize%2CCopyright" role="button" title="Time_103.jpg" alt="Time_103.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;

&lt;P&gt;The left pane shows less skew in thread completion of individual tiles (relatively same size). The right pane also illustrates the thread skew, but in this case, the different algorithm thread skew time becomes immaterial. You will notice in the right pane, left side of red zone, one of the threads is lagging behind (black stripe). This may possibly be due to preemption of the thread to run other processes inside the KNC.&lt;/P&gt;

&lt;P&gt;The first screenshot clearly illustrates the initialization overhead that you must be aware of when timing your code.&lt;/P&gt;

&lt;P&gt;Jim Dempsey&lt;/P&gt;</description>
      <pubDate>Mon, 19 Oct 2015 14:56:47 GMT</pubDate>
      <guid>https://community.intel.com/t5/Software-Archive/First-touch-time-greater-than-parallel-time/m-p/1024134#M39411</guid>
      <dc:creator>jimdempseyatthecove</dc:creator>
      <dc:date>2015-10-19T14:56:47Z</dc:date>
    </item>
  </channel>
</rss>

