~/et0
Notes from building and breaking infrastructure.
Posts 01
- 01 Building a Private GLM-5.2 Inference Service on Eight B200s
For a few weeks, I had access to a rented server with eight NVIDIA B200 GPUs. The server had enough memory to load GLM-5.2, but loading the model was never the goal. I wanted three or four colleagues