Video wird geladen...
Video konnte nicht geladen werden
15 concurrent terminal workloads on a local DGX Spark, all served by nvidia/nemotron-3-super 120B A12B NVFP4 through vLLM. 15/15 completed, 0 errors, 30.2s wall time: no fake dashboard, just local inference under concurrent load. fineprint: live local concurrency demo
11,375 Aufrufe • vor 4 Monaten •via X (Twitter)
0 Kommentare
Keine Kommentare verfügbar
Kommentare vom Original-Post werden hier angezeigt
