<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Measuring Memory Bus Traffic in Analyzers</title>
    <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846739#M2825</link>
    <description>&lt;DIV style="margin:0px;"&gt;&lt;/DIV&gt;
&lt;P&gt;As I recall, it was never possible to get more than 60% of theoretical bus bandwidth on Nocona, under realistic conditions. It looks like you already know more than I.&lt;/P&gt;</description>
    <pubDate>Thu, 13 Nov 2008 07:26:33 GMT</pubDate>
    <dc:creator>TimP</dc:creator>
    <dc:date>2008-11-13T07:26:33Z</dc:date>
    <item>
      <title>Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846733#M2819</link>
      <description>&lt;P&gt;I want to use Vtune to measure the traffic on memory bus when a program is running.&lt;/P&gt;
&lt;P&gt;I searched this forum and found some clues, however, they are not that clear.&lt;/P&gt;
&lt;P&gt;Probably, the following link gives some hint, and I hope somebody could say more about this topic.&lt;/P&gt;
&lt;P&gt;&lt;A href="http://software.intel.com/en-us/forums/showthread.php?t=44055" target="_blank"&gt;http://software.intel.com/en-us/forums/showthread.php?t=44055&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;
&lt;P&gt;BTW, I'm playing with a box with 4 Xeon CPUs.&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;
&lt;P&gt;Thanks in advance.&lt;/P&gt;</description>
      <pubDate>Tue, 11 Nov 2008 21:46:39 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846733#M2819</guid>
      <dc:creator>jqdu</dc:creator>
      <dc:date>2008-11-11T21:46:39Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846734#M2820</link>
      <description>&lt;DIV style="margin:0px;"&gt;&lt;/DIV&gt;
&lt;P&gt;I suggest you to readhttp://softwarecommunity.intel.com/isn/downloads/softwareproducts/pdfs/cycle_accounting.pdf&lt;/P&gt;</description>
      <pubDate>Wed, 12 Nov 2008 02:53:17 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846734#M2820</guid>
      <dc:creator>Peter_W_Intel</dc:creator>
      <dc:date>2008-11-12T02:53:17Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846735#M2821</link>
      <description>&lt;DIV style="margin:0px;"&gt;
&lt;DIV id="quote_reply" style="width: 100%; margin-top: 5px;"&gt;
&lt;DIV style="margin-left:2px;margin-right:2px;"&gt;Quoting - &lt;A href="https://community.intel.com/en-us/profile/338213"&gt;Zhen Yu Wang (Intel)&lt;/A&gt;&lt;/DIV&gt;
&lt;DIV style="background-color:#E5E5E5; padding:5px;border: 1px; border-style: inset;margin-left:2px;margin-right:2px;"&gt;&lt;EM&gt;
&lt;P&gt;I suggest you to readhttp://softwarecommunity.intel.com/isn/downloads/softwareproducts/pdfs/cycle_accounting.pdf&lt;/P&gt;
&lt;/EM&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;P&gt;Hi, this documents only covers Core 2 processors. I'm playing with a box with 4 Xeon processors. Some of the events are not available on my machine. Also, I'm not sure if two events have the same semantic because different naming methods are used for these two families.&lt;/P&gt;
&lt;P&gt;Do you have any idea about Xeon processors ?&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 12 Nov 2008 08:53:54 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846735#M2821</guid>
      <dc:creator>jqdu</dc:creator>
      <dc:date>2008-11-12T08:53:54Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846736#M2822</link>
      <description>&lt;DIV style="margin:0px;"&gt;
&lt;DIV id="quote_reply" style="width: 100%; margin-top: 5px;"&gt;
&lt;DIV style="margin-left:2px;margin-right:2px;"&gt;Quoting - &lt;A href="https://community.intel.com/en-us/profile/407368"&gt;jqdu&lt;/A&gt;&lt;/DIV&gt;
&lt;DIV style="background-color:#E5E5E5; padding:5px;border: 1px; border-style: inset;margin-left:2px;margin-right:2px;"&gt;&lt;EM&gt;
&lt;DIV style="margin:0px;"&gt;&lt;/DIV&gt;
&lt;P&gt;Hi, this documents only covers Core 2 processors. I'm playing with a box with 4 Xeon processors. Some of the events are not available on my machine. Also, I'm not sure if two events have the same semantic because different naming methods are used for these two families.&lt;/P&gt;
&lt;P&gt;Do you have any idea about Xeon processors ?&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;
&lt;/EM&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;P&gt;Xeon is a more inclusive term, covering Core 2, as well as those which came before and after. Possibly, if you would be more specific, someone could help.&lt;/P&gt;</description>
      <pubDate>Wed, 12 Nov 2008 15:11:41 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846736#M2822</guid>
      <dc:creator>TimP</dc:creator>
      <dc:date>2008-11-12T15:11:41Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846737#M2823</link>
      <description>&lt;DIV style="margin:0px;"&gt;
&lt;DIV id="quote_reply" style="width: 100%; margin-top: 5px;"&gt;
&lt;DIV style="margin-left:2px;margin-right:2px;"&gt;Quoting - &lt;A href="https://community.intel.com/en-us/profile/367365"&gt;tim18&lt;/A&gt;&lt;/DIV&gt;
&lt;DIV style="background-color:#E5E5E5; padding:5px;border: 1px; border-style: inset;margin-left:2px;margin-right:2px;"&gt;&lt;EM&gt;
&lt;DIV style="margin:0px;"&gt;&lt;/DIV&gt;
&lt;P&gt;Xeon is a more inclusive term, covering Core 2, as well as those which came before and after. Possibly, if you would be more specific, someone could help.&lt;/P&gt;
&lt;/EM&gt;&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;/DIV&gt;
&lt;P&gt;Sorry for this. It's a Xeon Nocona processor.&lt;/P&gt;
&lt;P&gt;The supported CPU Events for this platform are as below:&lt;BR /&gt;&lt;BR /&gt;128-bit MMX Instructions Retired&lt;BR /&gt;1st Level Cache Load Misses Retired&lt;BR /&gt;2nd Level Cache Load Misses Retired&lt;BR /&gt;2nd Level Cache Read Misses&lt;BR /&gt;2nd-Level Cache Read References&lt;BR /&gt;2nd-Level Cache Reads Hit Exclusive&lt;BR /&gt;2nd-Level Cache Reads Hit Modified&lt;BR /&gt;2nd-Level Cache Reads Hit Shared&lt;BR /&gt;3rd-Level Cache Read Misses&lt;BR /&gt;3rd-Level Cache Read References&lt;BR /&gt;3rd-Level Cache Reads Hit Exclusive&lt;BR /&gt;3rd-Level Cache Reads Hit Modified&lt;BR /&gt;3rd-Level Cache Reads Hit Shared&lt;BR /&gt;64-bit MMX Instructions Retired&lt;BR /&gt;64k/4M Aliasing Conflicts&lt;BR /&gt;All UC Underway from The Processor (AT-E)&lt;BR /&gt;All UC from the Processor&lt;BR /&gt;All WC Underway from The Processor (AT-E)&lt;BR /&gt;All WC from the Processor&lt;BR /&gt;All WCB Evictions (TI)&lt;BR /&gt;All calls&lt;BR /&gt;All conditionals&lt;BR /&gt;All indirect branches&lt;BR /&gt;All returns&lt;BR /&gt;Branches Retired&lt;BR /&gt;Bus Accesses Underway from All Agents (AT-E)&lt;BR /&gt;Bus Accesses Underway from The Processor (AT-E)&lt;BR /&gt;Bus Accesses from All agents&lt;BR /&gt;Bus Accesses from the Processor&lt;BR /&gt;Bus Data Ready from the Processor (TI)&lt;BR /&gt;Bus Reads Underway from The Processor (AT-E)&lt;BR /&gt;Clockticks&lt;BR /&gt;DTLB Load Misses Retired&lt;BR /&gt;DTLB Load and Store Misses Retired&lt;BR /&gt;DTLB Page Walks (TI)&lt;BR /&gt;DTLB Store Misses Retired&lt;BR /&gt;IO Reads Chunk (BSQ) (AT-E)&lt;BR /&gt;IO Writes Chunk (BSQ) (AT-E)&lt;BR /&gt;ITLB Misses&lt;BR /&gt;ITLB Page Walks (TI)&lt;BR /&gt;Instructions Completed&lt;BR /&gt;Instructions Retired&lt;BR /&gt;Loads Retired&lt;BR /&gt;MOB Loads Replays (AT-E)&lt;BR /&gt;MOB Loads Replays Retired&lt;BR /&gt;Machine Clear Count&lt;BR /&gt;Memory Order Machine Clear&lt;BR /&gt;Mispredicted Branches Retired&lt;BR /&gt;Mispredicted calls&lt;BR /&gt;Mispredicted conditionals&lt;BR /&gt;Mispredicted indirect branches&lt;BR /&gt;Mispredicted returns&lt;BR /&gt;Non-Halted Clockticks&lt;BR /&gt;Non-prefetch Bus Accesses from the Processor&lt;BR /&gt;Non-prefetch Reads Underway from The Processor (AT-E)&lt;BR /&gt;Packed Double-precision Floating-point Streaming SIMD Extension Instructions Retired&lt;BR /&gt;Packed Single-precision Floating-point Streaming SIMD Extension Instructions Retired&lt;BR /&gt;Reads Invalidate Full - RFO (BSQ) (AT-E)&lt;BR /&gt;Reads Non-prefetch Full (BSQ) (AT-E)&lt;BR /&gt;Reads Non-prefetch from the Processor&lt;BR /&gt;Reads from the Processor&lt;BR /&gt;Scalar Double-Precision Floating-Point Streaming SIMD Extension Instructions Retired&lt;BR /&gt;Scalar Single-precision Floating-point Streaming SIMD Extension Instructions Retired&lt;BR /&gt;Self-Modifying Code Clear&lt;BR /&gt;Speculative Instructions Completed&lt;BR /&gt;Speculative Microcode uops&lt;BR /&gt;Speculative TC-built uops&lt;BR /&gt;Speculative TC-delivered uops&lt;BR /&gt;Speculative Uops Retired&lt;BR /&gt;Split Load Replays&lt;BR /&gt;Split Loads Retired (AT-E)&lt;BR /&gt;Split Store Replays&lt;BR /&gt;Split Stores Retired (AT-E)&lt;BR /&gt;Stalled Cycles of Store Buffer Resources (Non-Standard)&lt;BR /&gt;Stalls of Store Buffer Resources (Non-Standard)&lt;BR /&gt;Stores Retired&lt;BR /&gt;Streaming SIMD Extensions Input Assists (TI)&lt;BR /&gt;TC flushes&lt;BR /&gt;TC to ROM Transfers&lt;BR /&gt;Tagged Mispredicted Branches Retired&lt;BR /&gt;Trace Cache Build Mode&lt;BR /&gt;Trace Cache Deliver Mode&lt;BR /&gt;Trace Cache Misses&lt;BR /&gt;UC Reads Chunk (BSQ) (AT-E)&lt;BR /&gt;UC Reads Chunk Split (BSQ) (AT-E)&lt;BR /&gt;UC Reads Chunk Underway (BSQ) (AT-E)&lt;BR /&gt;UC Write Partial (BSQ) (AT-E)&lt;BR /&gt;Uops Retired&lt;BR /&gt;WB Writes Full Underway(BSQ) (AT-E)&lt;BR /&gt;WCB Full Evictions (TI)&lt;BR /&gt;Write WC Full (BSQ) (AT-E)&lt;BR /&gt;Write WC Partial (BSQ) (AT-E)&lt;BR /&gt;Write WC Partial Underway (BSQ) (AT-E)&lt;BR /&gt;Writes Underway from The Processor (AT-E)&lt;BR /&gt;Writes WB Full (BSQ) (AT-E)&lt;BR /&gt;Writes from the Processor&lt;BR /&gt;x87 Input Assists&lt;BR /&gt;x87 Instructions Retired&lt;BR /&gt;x87 Output Assists&lt;/P&gt;</description>
      <pubDate>Wed, 12 Nov 2008 15:25:21 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846737#M2823</guid>
      <dc:creator>jqdu</dc:creator>
      <dc:date>2008-11-12T15:25:21Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846738#M2824</link>
      <description>&lt;DIV style="margin: 0px; height: auto;"&gt;&lt;/DIV&gt;
&lt;P&gt;Hi, Tim&lt;/P&gt;
&lt;P&gt;Now I'm using "Bus Data Ready from the Processor (TI)" to estimate the traffic on FSB.&lt;/P&gt;
&lt;P&gt;This event is explained in the Help Doc of Vtune as follows:&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;"This event counts the number of front-side bus clocks that the bus is   transmitting data driven by the processor core, including full reads|writes   and partial reads|writes and implicit writebacks."&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;The following link also mentioned that this event could be used to estimate the traffic.&lt;/P&gt;
&lt;P&gt;&lt;A href="http://software.intel.com/en-us/forums/showthread.php?t=44398" target="_blank"&gt;http://software.intel.com/en-us/forums/showthread.php?t=44398&lt;/A&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;"The Bus Data Ready event is most often used, in my limited experience, to determine the physical data rate. For Xeon/P4, this rate may be much higher than the useful data rate, particularly in cases of inefficient use of Read/Write Combining buffers. I don't know if you could correlate last level cache miss retired events with Bus Data Ready events, perhaps in cases where WCB use is efficient." &lt;/EM&gt;by tim18&lt;/P&gt;
&lt;P&gt;Here are some data from my experiment.&lt;/P&gt;
&lt;P&gt;During 0.5s, Vtune collected around 50,000,000 events on a FSB with a base frequency of 200MHz, quad pumped to 800MHz.&lt;/P&gt;
&lt;P&gt;Therefore, the theoretical FSB bandwidth should be 800M * 8 bytes/sec = 6.4GB/s&lt;/P&gt;
&lt;P&gt;However, the largest traffic I can produce with my program is about 50M * 2 * 8 bytes/sec * 4 = 3.2GB/s. Here 4 is "quad pumping". The program I used is like Stream Bench, which creats and writes to a large array.&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;
&lt;P&gt;Could you/anybody let me know how to use "Bus Data Ready from the Processor (TI)" to estimate the traffic on FSB?&lt;/P&gt;
&lt;P&gt;Or, are there any other events I can use to do this estimation?&lt;/P&gt;
&lt;P&gt;&lt;/P&gt;
&lt;P&gt;Thanks!&lt;/P&gt;</description>
      <pubDate>Wed, 12 Nov 2008 22:06:43 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846738#M2824</guid>
      <dc:creator>jqdu</dc:creator>
      <dc:date>2008-11-12T22:06:43Z</dc:date>
    </item>
    <item>
      <title>Re: Measuring Memory Bus Traffic</title>
      <link>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846739#M2825</link>
      <description>&lt;DIV style="margin:0px;"&gt;&lt;/DIV&gt;
&lt;P&gt;As I recall, it was never possible to get more than 60% of theoretical bus bandwidth on Nocona, under realistic conditions. It looks like you already know more than I.&lt;/P&gt;</description>
      <pubDate>Thu, 13 Nov 2008 07:26:33 GMT</pubDate>
      <guid>https://community.intel.com/t5/Analyzers/Measuring-Memory-Bus-Traffic/m-p/846739#M2825</guid>
      <dc:creator>TimP</dc:creator>
      <dc:date>2008-11-13T07:26:33Z</dc:date>
    </item>
  </channel>
</rss>

