Local GPU Inference and GLM 5.2 Rigs
Jul 20, 2026 · 693 messages · 87 active members
@jcartu detailed a 4x RTX 6000 Pro Blackwell (96GB each) rig running hybrid NF3/NVFP4 GLM 5.2 at 90–110 tps, 3500 tps prefill, ~500k context, drawing 3kW under load at ~60–70k EUR build cost. @momodxb pushed back: hybrid…
Read full digest →