Two Flash Models on Two DGX Sparks
A behavioral audit of GLM-5.3-Flash and Qwen3.8-Flash-Next, deployed day-0 on a 2× DGX Spark cluster. Two self-built test programs run identically against both live endpoints: a 40-probe behavioral battery with written expectations, and a blind-scored creative-engineering eval — both models build a voxel Shore Temple in a single HTML file. Every number in the note comes from a measured artifact, published here raw.
Both models were hours old when deployed; both ran tensor-parallel across a 200G ConnectX-7 fabric within a day of release, from community NVFP4 quants on patched serving stacks. Both are genuinely strong at the creative task — asked to build a voxel Shore Temple in a single HTML file, both produced scenes a human scored between 5 and 9 out of 10, blind.
The findings live around the generation: Qwen3.8-Flash-Next’s default reasoning setting frequently returns nothing at all inside a normal agent-loop budget, while GLM-5.3-Flash’s answers reliably arrive — but GLM warns conscientiously and then, twice, did what you asked anyway. Everything is stated as precisely as the data supports, and the data is on this page.