<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>et0</title><description>Notes from building and breaking infrastructure.</description><link>https://stg-4a3c.et0.dev/</link><item><title>Building a Private GLM-5.2 Inference Service on Eight B200s</title><link>https://stg-4a3c.et0.dev/posts/how-i-built-a-long-context-inference-service-on-eight-nvidia-b200-gpus/</link><guid isPermaLink="true">https://stg-4a3c.et0.dev/posts/how-i-built-a-long-context-inference-service-on-eight-nvidia-b200-gpus/</guid><description>For a few weeks, I had access to a rented server with eight NVIDIA B200 GPUs. The server had enough memory to load GLM-5.2, but loading the model was never the goal. I wanted three or four colleagues </description><pubDate>Wed, 05 Aug 2026 20:18:20 GMT</pubDate></item></channel></rss>