Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.
Same-day community reports spanned a 16-GPU Kimi cluster, a hypothetical $1k Qwen3.8 chip at 7,000 TPS, a three-week single-3090 run and memory-bandwidth overclocking.
TL;DR
- A 16-GPU GB10 cluster was reported running Kimi K3 (2.8T) at 30 tokens/s coding throughput and a 136 tokens/s concurrency peak.
- A separate thread floated a $1k Qwen3.8-27B 'Taalas' chip at 7,000 TPS, and another reported three weeks of Qwen 3.8 27B on a single RTX 3090.
- Memory stayed in focus: China's CXMT said a new memory-chip platform entered mass production, and a hobbyist reported 1.89 TB/s from an overclocked CMP 170hx.
Long-running and large-model serving dominated the day's local-inference reports. A 16-GPU GB10 cluster was reported running Kimi K3, a 2.8T-parameter model, at 30 tokens/s coding throughput with a 136 tokens/s concurrency peak, while a separate thread reported three weeks of continuous Qwen 3.8 27B use on a single RTX 3090. [1] [3]
Hardware speculation and tooling ran alongside. One thread asked whether readers would buy a $1k Qwen3.8-27B 'Taalas' chip able to run at 7,000 TPS; another recommended ExllamaV3 for Flash-Next. A Hacker News post compared self-hosted inference orchestrators including LocalAI, exo, GPUStack and vLLM. [2] [4] [8]
Benchmark threads put numbers on consumer cards: one user tested nine LLMs on the same web-development prompt for about eight hours on an RTX 3060 12GB, and a Hacker News submission collected a test of eight vision-language models on a single RTX 3090. [5] [6]
The constraint underneath remained memory. A thread reported 1.89 TB/s of memory bandwidth from an overclocked CMP 170hx, and another noted that China's CXMT said a new memory-chip platform had entered mass production. Both are single-source or headline-level and carry no controlled methodology. [7] [9]
Why it matters
These single-machine reports are the fastest public read on what open models actually need. Prefill and memory bandwidth, not headline decode speed, decide whether large-model setups are usable, and the supply sits with memory makers.
Editor's note
All figures are community reports from individual machines with no shared methodology, and the CXMT item is headline-level; the numbers are reported as claimed.