<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>GPU Scheduling on myvirtualcloud.net</title><link>https://myvirtualcloud.net/tags/gpu-scheduling/</link><description>Recent content in GPU Scheduling on myvirtualcloud.net</description><generator>Hugo</generator><language>en-us</language><copyright>Andre Leibovici</copyright><lastBuildDate>Thu, 10 Sep 2026 09:00:00 +1200</lastBuildDate><atom:link href="https://myvirtualcloud.net/tags/gpu-scheduling/index.xml" rel="self" type="application/rss+xml"/><item><title>Handing one node's GPUs to HAMi without touching the live models</title><link>https://myvirtualcloud.net/handing-one-nodes-gpus-to-hami-without-touching-the-live-models/</link><pubDate>Thu, 10 Sep 2026 09:00:00 +1200</pubDate><guid>https://myvirtualcloud.net/handing-one-nodes-gpus-to-hami-without-touching-the-live-models/</guid><description>&lt;p&gt;The box has four Blackwell RTX PRO 6000s in it. Two are Server Edition at 600 W, two are Max-Q Workstation Edition at 300 W. Kubernetes advertised four identical &lt;code&gt;nvidia.com/gpu&lt;/code&gt; and handed out whichever card it felt like.&lt;/p&gt;&#10;&lt;p&gt;For a single notebook that is fine. For tensor parallelism across two cards it is not. A mixed pair paces every collective at the slower card, so a TP=2 deployment that lands on one SE and one Max-Q quietly runs at Max-Q speed forever. That was the blocker on a research AI platform we run for a New Zealand university: three GPU nodes, two of them carrying eight H200s each and serving a live model to researchers, one RTX node for development and benchmarking.&lt;/p&gt;</description></item></channel></rss>