<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Error using more than 1 node in Intel® MPI Library</title>
    <link>https://community.intel.com/t5/Intel-MPI-Library/Error-using-more-than-1-node/m-p/1035541#M4233</link>
    <description>&lt;P&gt;HELP!&lt;BR /&gt;
	I have some code to simply solve Ax=b using an iterative method. I read in a partitioned file and load A and b into each processes unique A and b. Then iteratively try to solve. Everything is great until I run on more than 1 node. I have a 64 partition example which runs great on 1 node(20 physical cores) but trying to run on 40 60 or 80 cores there is erroneous behavior ( get wrong results or MPI hangs). I am going crazy over this. The unpartitioned Ax=b works great, 8 partitions and 16 work great 32 works on 2 nodes consistently but takes extra time for some reason. and 64 will hang or hit iteration limits and give incorrect results.&lt;/P&gt;

&lt;P&gt;I'm using the Eigen matrix library . doing a simple ISend and Recv multiple times. And doing MPI_AllReduce.&lt;/P&gt;

&lt;P&gt;I've tried using MPICH, openmpi, intel-mpi, they all compile without warnings but all suffer from the same issue. openmpi seems to suffer worse on my machine, causing some extra warnings to be generated during run time.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Im using&amp;nbsp;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;Cray CS300-LC Linux Cluster&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;2560 compute cores (2.8 GHz Intel Xeon E5-2680 v2), 128 nodes&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;15,360 coprocessor cores (Intel Xeon Phi 5110P), two per node&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;8 Terabytes of RAM&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;FDR InfiniBand Network&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;Any help is appreciated!&lt;/SPAN&gt;&lt;/P&gt;

&lt;DL&gt;
	&lt;DD&gt;&amp;nbsp;&lt;/DD&gt;
	&lt;DD&gt;&amp;nbsp;&lt;/DD&gt;
&lt;/DL&gt;

&lt;P&gt;Thanks,&lt;/P&gt;

&lt;P&gt;Brian&lt;/P&gt;</description>
    <pubDate>Mon, 21 Sep 2015 01:02:53 GMT</pubDate>
    <dc:creator>Brian_S_3</dc:creator>
    <dc:date>2015-09-21T01:02:53Z</dc:date>
    <item>
      <title>Error using more than 1 node</title>
      <link>https://community.intel.com/t5/Intel-MPI-Library/Error-using-more-than-1-node/m-p/1035541#M4233</link>
      <description>&lt;P&gt;HELP!&lt;BR /&gt;
	I have some code to simply solve Ax=b using an iterative method. I read in a partitioned file and load A and b into each processes unique A and b. Then iteratively try to solve. Everything is great until I run on more than 1 node. I have a 64 partition example which runs great on 1 node(20 physical cores) but trying to run on 40 60 or 80 cores there is erroneous behavior ( get wrong results or MPI hangs). I am going crazy over this. The unpartitioned Ax=b works great, 8 partitions and 16 work great 32 works on 2 nodes consistently but takes extra time for some reason. and 64 will hang or hit iteration limits and give incorrect results.&lt;/P&gt;

&lt;P&gt;I'm using the Eigen matrix library . doing a simple ISend and Recv multiple times. And doing MPI_AllReduce.&lt;/P&gt;

&lt;P&gt;I've tried using MPICH, openmpi, intel-mpi, they all compile without warnings but all suffer from the same issue. openmpi seems to suffer worse on my machine, causing some extra warnings to be generated during run time.&amp;nbsp;&lt;/P&gt;

&lt;P&gt;Im using&amp;nbsp;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;Cray CS300-LC Linux Cluster&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;2560 compute cores (2.8 GHz Intel Xeon E5-2680 v2), 128 nodes&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;15,360 coprocessor cores (Intel Xeon Phi 5110P), two per node&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;8 Terabytes of RAM&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;FDR InfiniBand Network&lt;/SPAN&gt;&lt;/P&gt;

&lt;P&gt;&amp;nbsp;&lt;/P&gt;

&lt;P&gt;&lt;SPAN style="line-height: 1.5; font-size: 1em;"&gt;Any help is appreciated!&lt;/SPAN&gt;&lt;/P&gt;

&lt;DL&gt;
	&lt;DD&gt;&amp;nbsp;&lt;/DD&gt;
	&lt;DD&gt;&amp;nbsp;&lt;/DD&gt;
&lt;/DL&gt;

&lt;P&gt;Thanks,&lt;/P&gt;

&lt;P&gt;Brian&lt;/P&gt;</description>
      <pubDate>Mon, 21 Sep 2015 01:02:53 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-MPI-Library/Error-using-more-than-1-node/m-p/1035541#M4233</guid>
      <dc:creator>Brian_S_3</dc:creator>
      <dc:date>2015-09-21T01:02:53Z</dc:date>
    </item>
    <item>
      <title>The Intel® Trace Analyzer and</title>
      <link>https://community.intel.com/t5/Intel-MPI-Library/Error-using-more-than-1-node/m-p/1035542#M4234</link>
      <description>&lt;P&gt;The Intel® Trace Analyzer and Checker has a Message Checking capability.&amp;nbsp; I would recommend using this to examine your code.&amp;nbsp; See &lt;A href="https://software.intel.com/en-us/articles/intel-trace-analyzer-and-collector-for-linux-intel-mpi-correctness-checking-library"&gt;https://software.intel.com/en-us/articles/intel-trace-analyzer-and-collector-for-linux-intel-mpi-correctness-checking-library&lt;/A&gt; for details on how to use it.&lt;/P&gt;</description>
      <pubDate>Fri, 16 Oct 2015 15:03:38 GMT</pubDate>
      <guid>https://community.intel.com/t5/Intel-MPI-Library/Error-using-more-than-1-node/m-p/1035542#M4234</guid>
      <dc:creator>James_T_Intel</dc:creator>
      <dc:date>2015-10-16T15:03:38Z</dc:date>
    </item>
  </channel>
</rss>

