正在加载视频...

视频加载失败

We turned Qwen3.8-27B into a multimodal decision model. It beat Pokémon FireRed’s elite four and champion with sub-100 ms decisions from live game state. With SGLang’s native /v1/decisions, you can now turn LLMs and VLMs into classification and scoring models. We also added /v1/systemone so Jev-like open models can...

221,988 次观看 • 1 天前 •via X (Twitter)

21 条评论

SGLang 的头像
SGLang1 天前

One of the most exciting parts of Pokémon is type matchups and knowing when to attack, switch, or heal. Qwen3.8-27B is a surprisingly good Pokémon player. It cleared the Elite Four and Champion in one go. Full run playback👇

SGLang 的头像
SGLang1 天前

You can now try SGLang’s new decision API with the nightly image. It will also be included in v0.5.21 this Thursday.

BlackwellBoy 的头像
BlackwellBoy1 天前

I wonder if this could help with quicker decision making for my little chits project

Stochastic Cowboy 的头像
Stochastic Cowboy1 天前

Very cool! Will definitely try this out. For the Pokemon demo is there any harness that is tracking game state, or is it purely a loop where the model is presented with the screenshot and outputs an action?

Sakura Yuki 的头像
Sakura Yuki1 天前

Does sub-100 ms include game-state capture, image preprocessing, and VLM prefill, or is it just /v1/decisions server time? That breakdown is the interesting part.

bhavesh 的头像
bhavesh1 天前

awesome! but FireRed is easy :) we should see if models can win truly difficult battles

Day | AI + Software 的头像
Day | AI + Software1 天前

Woahhhh impressive might try out SG Lang for the first time

Trust⭕️ 的头像
Trust⭕️1 天前

Woah

Alan Ma 的头像
Alan Ma23 小时前

goated team

Will 的头像
Will1 天前

Yes! I love it.

Greg Konush (。◕‿◕。) 的头像
Greg Konush (。◕‿◕。)1 天前

p99 under concurrent load?

Amar Singh 的头像
Amar Singh1 天前

crazy to think what the future of video games are going to be, pokemon is fine but it's going to be impossible to determine if someone is a bot in online games

Hardik Jindal 的头像
Hardik Jindal1 天前

curious if you have added a lora adapter for qwen3.8, similar to kev 27b 🤔

Azeez 的头像
Azeez1 天前

This is so so sick 🔥 FireRed is my favorite Pokemon game (close second Soul Silver)

Shreyas Karnik 的头像
Shreyas Karnik1 天前

really cool!

Bin Chen 的头像
Bin Chen1 天前

can it be used in other task, eg: given two chat history capture, and ask if there is new message?

Jintao Zhang 张晋涛 的头像
Jintao Zhang 张晋涛1 天前

🐐

John Rood 的头像
John Rood1 天前

the artifact I want off this run: the decisive turns, the spots where a wrong move loses the battle, and what it picked there. a playthrough is mostly forgiving, so the clear is really a story about the handful of moments that weren't.

季风 的头像
季风1 天前

cool!

SocialWatch 的头像
SocialWatch1 天前

This is so cool!

Louis 的头像
Louis23 小时前

cool

相关视频

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 次观看 • 3 年前

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 次观看 • 2 年前

this is f*cking gold 20 GitHub repos with 500K+ combined stars that will level up your JEV workflow AGENTS > jev-ultrafast: Jev picks every click and DOM target, a small LLM only types > hermes-jev-skills: routing, memory, compaction and skill picks in one pack > typesafe-computer-use: OCR reads your Mac screen, Jev picks the next click MEMORY > fast-jev-compaction: scores every tool call keep, truncate or drop instead of summarizing > jevmem: project memory for Claude Code, Cursor and Codex, updated every turn > jev-second-brain: your Obsidian vault, Jev judges which notes duplicate, revise or contradict SAFETY > jev-guard: a gate before every tool call, Jev scores the risk, you set allow, ask or deny TOOLS > skills: the official TypeSafe skill for Claude Code and Codex > system-one-adapter-python: dry-run your questions on an ordinary LLM before you burn a Jev key > jev-mcp: claim checks, screening and ranking as MCP tools > typesafe-mcp: plug Jev into any MCP client > json-render: Vercel's generative UI, where Jev picks the components OPEN MODELS > SemIf-OpenJev: semantic ifs from frozen open models kev: Jev-like models on Qwen that run on your MacBook > laya-mlx: the Laya decision engine on Apple silicon > clm: an open System One model with Choice, Noul and Score > jevlike: train your own Jev-like model TRADING > jev-trader: one buy or sell decision per Monad block START HERE > awesome-jev: the biggest map of everything built on Jev > awesome-jev-by-typesafe: use cases, patterns and starter code bookmark it before your next build

NO1ennn

18,445 次观看 • 2 天前