<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic GPU 1100 Max temperature and nonfatal errors when running xpui-smi? in GPU Compute Software</title>
    <link>https://community.intel.com/t5/GPU-Compute-Software/GPU-1100-Max-temperature-and-nonfatal-errors-when-running-xpui/m-p/1634614#M1586</link>
    <description>&lt;P&gt;We are creating an agent to pull data from XPUs.&amp;nbsp; I have a node with Ubuntu 22.04.5 LTS OS&lt;/P&gt;&lt;P&gt;root@n022:~# uname -a&lt;BR /&gt;Linux n022 5.15.0-122-generic #132-Ubuntu SMP Thu Aug 29 13:45:52 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Hardware:&amp;nbsp;&lt;BR /&gt;root@n022:~# clinfo -l&lt;BR /&gt;Platform #0: Intel(R) OpenCL Graphics&lt;BR /&gt;+-- Device #0: Intel(R) Data Center GPU Max 1100&lt;BR /&gt;`-- Device #1: Intel(R) Data Center GPU Max 1100&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Driver:&lt;/P&gt;&lt;P&gt;root@n022:~# dkms status | grep -i i9&lt;BR /&gt;AUXILIARY_BUS is enabled for 5.15.0-122-generic.&lt;BR /&gt;intel-i915-dkms/1.23.10.72.231129.76, 5.15.0-122-generic, x86_64: installed&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Temperature sensor information returns N/A.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;root@n022:~# xpu-smi dump -d "-1" -m1,2,3,4,5,18 -i 1 -n1&lt;BR /&gt;Timestamp, DeviceId, GPU Power (W), GPU Frequency (MHz), GPU Core Temperature (Celsius Degree), GPU Memory Temperature (Celsius Degree), GPU Memory Utilization (%), GPU Memory Used (MiB)&lt;BR /&gt;11:12:08.956, 0, 52.37, 0, N/A, N/A, 0.05, 28.13&lt;BR /&gt;11:12:08.956, 1, 49.37, 0, N/A, N/A, 0.05, 27.99&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Is this something that is planned to be fixed soon?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Lastly, when I run xpu-smi, I get the below non-fatal errors in syslog.&amp;nbsp; Are these known issues?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The example code works well. Host has Xeon Max 8480+.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;[84460.492225] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected NONFATAL error GFX_MSTR_INTR:0x08000000&lt;BR /&gt;[84460.504560] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected DEV_ERR_STAT_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.516620] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC NONFATAL error&lt;BR /&gt;[84460.526912] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_MASTER_REG_NONFATAL:0x00000002&lt;BR /&gt;[84460.540884] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_SLAVE_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.554187] i915 0000:9a:00.0: [drm] *ERROR* GT0 [INTERRUPT] Invalid HBM SS3: Channel7 SOC NONFATAL error&lt;BR /&gt;[84460.565586] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected NONFATAL error GFX_MSTR_INTR:0x08000000&lt;BR /&gt;[84460.577907] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected DEV_ERR_STAT_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.589952] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC NONFATAL error&lt;BR /&gt;[84460.600246] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_MASTER_REG_NONFATAL:0x00000002&lt;BR /&gt;[84460.614211] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_SLAVE_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.627510] i915 0000:9a:00.0: [drm] *ERROR* GT0 [INTERRUPT] Invalid HBM SS3: Channel7 SOC NONFATAL error&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Brgds,&lt;/P&gt;&lt;P&gt;ToreL&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
    <pubDate>Wed, 02 Oct 2024 09:23:06 GMT</pubDate>
    <dc:creator>Tore</dc:creator>
    <dc:date>2024-10-02T09:23:06Z</dc:date>
    <item>
      <title>GPU 1100 Max temperature and nonfatal errors when running xpui-smi?</title>
      <link>https://community.intel.com/t5/GPU-Compute-Software/GPU-1100-Max-temperature-and-nonfatal-errors-when-running-xpui/m-p/1634614#M1586</link>
      <description>&lt;P&gt;We are creating an agent to pull data from XPUs.&amp;nbsp; I have a node with Ubuntu 22.04.5 LTS OS&lt;/P&gt;&lt;P&gt;root@n022:~# uname -a&lt;BR /&gt;Linux n022 5.15.0-122-generic #132-Ubuntu SMP Thu Aug 29 13:45:52 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Hardware:&amp;nbsp;&lt;BR /&gt;root@n022:~# clinfo -l&lt;BR /&gt;Platform #0: Intel(R) OpenCL Graphics&lt;BR /&gt;+-- Device #0: Intel(R) Data Center GPU Max 1100&lt;BR /&gt;`-- Device #1: Intel(R) Data Center GPU Max 1100&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Driver:&lt;/P&gt;&lt;P&gt;root@n022:~# dkms status | grep -i i9&lt;BR /&gt;AUXILIARY_BUS is enabled for 5.15.0-122-generic.&lt;BR /&gt;intel-i915-dkms/1.23.10.72.231129.76, 5.15.0-122-generic, x86_64: installed&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Temperature sensor information returns N/A.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;root@n022:~# xpu-smi dump -d "-1" -m1,2,3,4,5,18 -i 1 -n1&lt;BR /&gt;Timestamp, DeviceId, GPU Power (W), GPU Frequency (MHz), GPU Core Temperature (Celsius Degree), GPU Memory Temperature (Celsius Degree), GPU Memory Utilization (%), GPU Memory Used (MiB)&lt;BR /&gt;11:12:08.956, 0, 52.37, 0, N/A, N/A, 0.05, 28.13&lt;BR /&gt;11:12:08.956, 1, 49.37, 0, N/A, N/A, 0.05, 27.99&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Is this something that is planned to be fixed soon?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Lastly, when I run xpu-smi, I get the below non-fatal errors in syslog.&amp;nbsp; Are these known issues?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The example code works well. Host has Xeon Max 8480+.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;[84460.492225] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected NONFATAL error GFX_MSTR_INTR:0x08000000&lt;BR /&gt;[84460.504560] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected DEV_ERR_STAT_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.516620] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC NONFATAL error&lt;BR /&gt;[84460.526912] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_MASTER_REG_NONFATAL:0x00000002&lt;BR /&gt;[84460.540884] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_SLAVE_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.554187] i915 0000:9a:00.0: [drm] *ERROR* GT0 [INTERRUPT] Invalid HBM SS3: Channel7 SOC NONFATAL error&lt;BR /&gt;[84460.565586] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected NONFATAL error GFX_MSTR_INTR:0x08000000&lt;BR /&gt;[84460.577907] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected DEV_ERR_STAT_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.589952] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC NONFATAL error&lt;BR /&gt;[84460.600246] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_MASTER_REG_NONFATAL:0x00000002&lt;BR /&gt;[84460.614211] i915 0000:9a:00.0: [drm] *ERROR* [Hardware Error]: GT0 detected SOC_GLOBAL_ERR_STAT_SLAVE_REG_NONFATAL:0x00010000&lt;BR /&gt;[84460.627510] i915 0000:9a:00.0: [drm] *ERROR* GT0 [INTERRUPT] Invalid HBM SS3: Channel7 SOC NONFATAL error&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Brgds,&lt;/P&gt;&lt;P&gt;ToreL&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 02 Oct 2024 09:23:06 GMT</pubDate>
      <guid>https://community.intel.com/t5/GPU-Compute-Software/GPU-1100-Max-temperature-and-nonfatal-errors-when-running-xpui/m-p/1634614#M1586</guid>
      <dc:creator>Tore</dc:creator>
      <dc:date>2024-10-02T09:23:06Z</dc:date>
    </item>
  </channel>
</rss>

