正在加载视频...
视频加载失败
DUAL INTEL ARC PRO B70 GPUS AT $1,900 RUN QWEN 3 CODER 30B AT 268 TOKENS PER SECOND VIA VLLM ON 8 CONCURRENT REQUESTS, A NEW INTEL TIER UNDER YOUR MAP THAT BEATS RTX 3090 STACKS ON THROUGHPUT PER DOLLAR 02:22 the operator reads off his benchmark sheet, "vllm... show more
29,137 次观看 • 3 个月前 •via X (Twitter)
40 条评论

Intel gets interesting the moment concurrency matters more than single-user speed

ye concurrency is where the real game starts

dual b70 for 1900 and beating 3090 stacks on concurrent stuff is wild. finally intel making sense for local runs

ye intel finally pulling weight on concurrent runs

intel also have good hardware

ye intel hardware finally pulling its weight

@starmexxx dude, that's wild. didn’t expect intel to push past 3090s like that. game changer for sure

keep spreading this

appreciate it bro, intel tier is sleeper alpha

Intel really excels in all these features

ye intel finally cooking, the timing is wild

Intel quietly becoming viable

ye intel quietly stacking wins under the radar

So single stream is like 32 tok/s?

this is whats gonna shift homelabs from 3090 stacks to mixed-vendor builds

'Slow' Intel cards beating NVIDIA stacks on throughput because someone optimized the serving layer. ML remains not a serious science and the jankiest shit somehow works.

Wow, that's pretty impressive performance for the price point. Intel is really stepping up.

What do you think about this @grok ?

now way in hell I would run LLM on a ǰèŵ.

spending less for more is wild

ye throughput per dollar finally flipped

but what drivers support it, why not stay in the CUDA lane? and the VRAM what 3x slower than RTX?

intel just got interesting again

Intel finally making sense for local AI. Dual B70 for $1900 pulling 268 t/s is actually wild.

local hardware setups are getting scary fast these days

nvidia shaking in their overpriced boots 😂

it's now more cost-effective to buy a graphics card than to pay for a monthly subscription

exactly bro, the gpu pays itself off in months

Hey @starmexxx who is the guy in this video? Does he have a youtube channel?

intel just flipped the script hard

ye intel coming in hot, throughput per dollar is wild

bookmarked this!

ty!

Throughput per dollar maybe, but the layout matters more than the cards. On 4x3090 the trick is NOT tensor-parallel: 4 data-parallel copies of Gemma 12B do 3,425 tok/s vs 1,007 for TP4. Most 3090 stacks are just configured wrong:

Tell me what OS version and kernel you are using please? My arc pro wedges between 60-90 seconds of agentic load. It’ll donthe benchmarks, but is unusable for real workloads.

throughput per dollar is the right axis. the Intel tax shows up later in driver maturity and the hours you spend debugging the stack

Hi @starmexxx grettings from Mexico, what CPU (Intel or AMD) and how much RAM do you have in your setup?? Also can you recommend me or suggest a good motherboard to support both Intel cards?

Throughput per dollar is becoming a serious AI metric. The next deployment wave will care less about demos and more about what can run reliably at usable cost.

268 t/s on Intel Arc via vLLM at $1,900-software beats hardware, again. I'm waiting for independent tests, but if this is confirmed, it changes the equation for small teams.

Battlemage stronk af 🦾🦾🦾🦾🦾
