Jointly announcing EAGLE-3 with SGLang: Setting a new record... in LLM inference acceleration! - 5x🚀than vanilla (on HF) - 1.4x🚀than EAGLE-2 (on HF) - A record of ~400 TPS on LLama 3.1 8B with a single H100 (on SGLang) - 1.65x🚀in latency even for large bs=64 (on SGLang) - A new scaling law: more training data, better speedup - Apache 2.0 Paper: Code: SGLang version: ⚒️Takeaway: Introducing training-time test, a novel draft model training technique: we replace feature prediction with direct token prediction and shift from top-layer-only features to multi-layer feature fusion. This approach unlocks a new scaling law previously undiscovered in EAGLE and EAGLE-2. 🙏Acknowledge: We would like to thank the SGLang team (zhyncs Lianmin Zheng Ying Sheng James Liu, Ke Bao, and others LMSYS Org) for their merge and careful evaluation of EAGLE-3 on SGLang. 🤝Want to collaborate? We're a small academic group with limited GPU resources. If you're interested in supporting our next version of EAGLE or would like us to train a preliminary version tailored to a specific model, please get in touch! Joint work with Yuhui Li, Fangyun Wei, and Chao Zhangshow more

Hongyang Zhang
42,200 Aufrufe • vor 1 Jahr
Introduce EAGLE, a new method for fast LLM decoding... based on compression: - 3x🚀than vanilla - 2x🚀 than Lookahead (on its benchmark) - 1.6x🚀 than Medusa (on its benchmark) - provably maintains text distribution - trainable (in 1~2 days) and testable on RTX 3090s Playground: Blog: Code: ⚒️First Principle: Compression! Yi Ma We find that the sequence of second-top-layer features is compressible, making the prediction of subsequent feature vectors from previous ones easy by a small model. 🙏Acknowledge: This project is greatly inspired by the Medusa team (Tianle Cai @yli3521 Zhengyang Geng Hongwu Peng Tri Dao), the Lookahead team (Hao Zhang LMSYS Org), and others. Joint work with Yuhui Li and Chao Zhangshow more

Hongyang Zhang
118,855 Aufrufe • vor 2 Jahren
🚀 Self-speculation brings 6.75x real speedup for LLM generation... with SGLang inference! Same model drafts future tokens in Diffusion mode → then verifies them in AR (causal) mode. One model and one KV cache. Just different attention masks. Thanks to perfect alignment, we get 2× longer acceptance lengths than MTP techniques (Eagle-3, MTP, dFlash). We run 2 forward passes… but the 2× higher acceptance means we break even - and with zero overhead from extra drafter, KV cache, or LM head that comes with MTP - those are not free. Last week we released Nemotron-Labs-Diffusion + Tri-mode LLMs! We did continued pre-training on Ministral-3 models by switching attention patterns (block causal bidirectional). Result: one model that runs AR mode, Diffusion mode, and Self-Speculation. Diffusion mode already shows high benchmark accuracy - excited to see what happens when someone beats left-to-right acceptance! 🔥 Github: Paper: SGLang inference: Try the models on HF:show more

Pavlo Molchanov
66,554 Aufrufe • vor 2 Monaten
World modeling and imitation learning have largely been considered... two disparate worlds. In our recent work, Unified World Models, just accepted to #RSS2025, Chuning Zhu provides a dead-simple unifying solution: just train a joint diffusion model over actions and future states, but with *decoupled* diffusion time steps across these modalities. Manipulating these decoupled time steps then allows for marginalization or conditioning on actions or states; a single model can serve as a policy, forward dynamics model, video prediction model, or inverse dynamics model by simply setting diffusion timesteps carefully. The resulting model can leverage video datasets along with robot training data much more effectively, and shows improved robustness, generalization, and flexibility. This is exciting because it is frustratingly simple, scalable, and shows strong improvement on real-world robotics problems. Please refer to Chuning Zhu 's excellent thread for more details! More details/code can be found on our website and in the paper -show more

Abhishek Gupta
11,430 Aufrufe • vor 1 Jahr
🚀 Attention Foxxi Community! 🚀 We're thrilled to invite... you to test our new Metaverse game #Bitverse and earn points for an upcoming Foxxi token airdrop & WL for #Bitminer ! Here's how you can participate: Application Rules: 1. Like, retweet & Comment on this post. 2. Follow @foxxiofficial 3. Join Discord : 3. Fill out the form Testing Rules: - Up to 100 people will get an invitation to test the BitVerse on the previously announced date (approx in 2 weeks). - All testers will earn points for the Foxxi token airdrop. - Test for at least 30 minutes any time on the event day. - Take a snapshot or record a video inside the BitVerse. - Tweet about your experience with the screenshot/video. Bonus Points: - Be present at the start of the event. - Fill out the feedback form with detailed feedback and bug reports. Join us and be part of the future of gaming and digital ownership! #bitmap #ordinals #Bitcoin #Airdrop 🎮✨🚀show more

JUNGL.
36,380 Aufrufe • vor 2 Jahren
HUB IS BACK 🌀 Despite attempts to silence us,... the Cult lives on, thriving in the face of adversity. Here’s a breakdown of what happened and what comes next ⤵️ As previously shared by @zastrahub, the Hub 𝕏 account has been targeted by several malicious attacks, with a major one striking on March 15, which ultimately led to its suspension. In the midst of this uncertainty, the team was reactive enough to fight back on all fronts. • Deployed Hub AI Agent Training on Discord to continue feeding the data layer. • Managed relationships with creators and advertisers efficiently. • Identified the problem and communicated with the X team. After submitting our report to X, we successfully regained control of the Hub account. We appreciate the 𝕏 team’s cooperation and will continue working closely with them in the future. For now, the Hub AI Agent remains on Discord, with one question posted daily in the "AI Training" section. We're approaching 500,000 replies in under a week, feeding Hub’s Data Layer at full speed. Meanwhile, we’re working on the release of Hub AI V2: a major evolution that will introduce powerful new social mechanisms and take the experience to the next level. Entities tried to shut us down. They made us unstoppable. To everyone who supported us, rallied behind us, and kept the fire alive: thank you. You are the Cult. You are the movement. We’re just getting started ⚡️show more

Hub
161,153 Aufrufe • vor 1 Jahr
Pyth Price Feeds are blasting off 🚀 Blast has... entered into orbit as a new Ethereum Layer 2 and the first of its kind to offer native yield for ETH and stablecoins. Blast is now live on mainnet. Learn more about Pyth’s deployment on Blast: ℹ️ About Blast Blast is the latest advancement in Ethereum Layer 2 solutions, delivering native yield for ETH and stablecoins. It accelerates and economizes transactions, with the backing of industry leaders like Paradigm, Standard Crypto, and eGirl Capital. 🔮 Pyth's Data-Powered Vision on Blast Over 15 apps have launched on the Blast and are harnessing Pyth’s low-latency, high-resolution price data: meathook—a gateway to 100+ crypto assets with high-leverage options. 100x—a high-speed perpetual DEX experience. Aark Digital—1000x perpetual DEX powered by LST/LRT. Blast Futures—a platform integrating perpetuals with native yield. Bloom—a leveraged trading DEX for rebasing assets. Curvance—a modular multi-chain money market with boosted yield. Deriblast—blends trading with gaming to create a unique experience. Easy X—a reimagined perpetual protocol for diverse asset exposure. Fragment—a new foundation for liquidity and lending protocols. HMX 🐉—a decentralized perpetual protocol with versatile collateral options. Juice Finance—an innovative approach to cross-margin DeFi. @Laser_on_Blast—a liquidity layer for on-chain banking on Blast. Orbit Protocol 🥮—a decentralized protocol for asset lending and borrowing. SynFutures—a decentralized derivatives trading protocol. YOLO GAMES—the go-to for high-stakes Degen Gaming. Zest 👾⚡️Genesis Version⚡️—a collateralized stablecoin with 100% capital efficiency. Pac Finance—a new pioneering DeFi hub on Blast. Seismic Finance—a new Blast native lending market. Thanks to the Pyth oracle, Blast is charting a new course for DeFi—one where accuracy and speed are not just nice-to-have features, but fundamentals that redefine users’ expectations and standards for on-chain finance.show more

Pyth Network 🔮
202,443 Aufrufe • vor 2 Jahren
What if you kept asking an LLM to "make... it better"? In some recent work at FAIR, we investigate how we can efficiently use RL to fine-tune LLMs to iteratively self-improve on their previous solutions at inference-time. Training for iterated self-improvement can be costly. The naive approach to training for K self-improvement steps leads to K times the number of rollout steps per episode. We introduce Exploratory Iteration (ExIt), an RL-based automatic curriculum method that bootstraps diverse training distributions of self-improvement tasks by upcycling the LLM's own responses at previous turns as the starting points for both self-improvement and *self-divergence.* In order to decide what task to train on next, the curriculum prioritizes sampling of partial turn histories that led to higher return variance in its GRPO group (a learnability score that comes for free). This automatic curriculum over the bootstrapped task space teaches the model how to perform iterated self-improvement while only ever training the model on single-step self-improvement tasks. We look at ExIt's impact in both single-turn (contest math problems) and multi-turn (BFCLv3 multi-turn tasks), as well as MLE-bench, where the LLM is run in a search scaffold to produce solutions to real Kaggle competitions. Across these eval settings, we find ExIt produces models with greater capacity for inference-time self-improvement compared to GRPO. Notably, ExIt models can self-improve on test tasks for many more steps than the typical solution depth encountered during training, including a 22% improvement in MLE-bench performance compared to GRPO.show more

Minqi Jiang
41,099 Aufrufe • vor 10 Monaten
Batch Normalization by hand ✍️ ~ 7 steps walkthrough... below Batch normalization is common practice for improving training and achieving faster convergence. It sounds simple. But it is often misunderstood. 🤔 Does batch normalization involve trainable parameters, tunable hyper-parameters, or both? 🤔 Is batch normalization applied to inputs, features, weights, biases, or outputs? 🤔 How is batch normalization different from layer normalization? So I drew and calculated one entirely by hand. Goal: normalize a mini-batch of 4 examples to mean 0 and variance 1, then let the network scale it back. = 1. Given = A mini-batch of 4 training examples, each with 3 features. = 2. Linear layer = Let us multiply by the weights and add the biases. Batch norm sits after this, which answers the second question: what gets normalized is features, not inputs, weights or biases. = 3. ReLU = We apply the activation, and -2 becomes 0. Negative values are suppressed before any statistic is taken. = 4. Batch statistics = Let us compute the sum, mean, variance and standard deviation, one row at a time. A row is a feature and the four columns are the four examples, so every number here measures one feature against the rest of the batch. That is the "batch" in batch normalization, and it is exactly what layer normalization does not do. The statistics are rounded to whole numbers, which is what keeps the rest of the page doable in pen. = 5. Shift to mean 0 = We subtract the mean, in green. The four values in each feature now average to zero. = 6. Scale to variance 1 = Let us divide by the standard deviation, in orange. Each feature now has variance one, whatever scale it arrived at. = 7. Scale and shift = We multiply by a linear transformation and pass the result on. The diagonal and the last column are trainable, so having just forced every feature to mean 0 and variance 1, we hand the network the means to undo it. The outputs: Mean of each feature = [2, 1, 2] Std dev of each feature = [1, 1, 2] To the next layer = [2, -2, 2, 0], [-3, 3, 6, -3], [2, 0, 1, 2] The answers: 🤔 Both. The scale and shift are trainable, the statistics are not. Epsilon and the momentum on the running statistics are the hyper-parameters, and one mini-batch by hand needs neither. 🤔 Features, after the linear layer, not inputs, weights or biases. 🤔 Batch norm measures across the batch, one feature at a time. Layer norm measures across the features, one example at a time. 💾 Save this post!show more

Tom Yeh
20,518 Aufrufe • vor 12 Tagen
Prediction Market is one of the leading Web3 niches... in 2024! With about $4 Billion in trading volume, $192 Million in TVL and millions of users in 2024, prediction protocols have been showing tremendous growth and attracting user adoption in Web3. PolyMarket seems to be leading the pack, with over $175M in TVL, and backing from Vitalik Buterin. However, a majority of Prediction Markets, including PolyMarket, currently lacks the flexibility and capital efficiency required for seamless transactions. Also, they are all majorly focused on driving Web2 users thus neglecting the need for markets that cater for short-term, high-risk investments from Web3 Degens. To tackle these flaws, there's a need for a revolutionary contender that understands the need of Web3 Chads That's where Predict Hub comes in PredictHub is a prediction Market that transforms real world events into opportunities for everyone to participate and forecast. Launching on Arbitrum, PredictHub is already catching the attention of major players in the Space by offering something PolyMarket and others don't - Flexibility and Incentives. By offering fast market updates and innovative prediction category like ETF Forecasts, PredictHub is changing how we interact with Prediction Markets. But then, here's where it gets more interesting; PredictHub offer users a unique point system, where you don't just make predictions, you also earn rewards. The more you Predict, the more you earn. These rewards are 2-fold: Nova and Orbit Points. Nova Points are earned by traders based on their trading activity and their leaderboard ranking. Orbit Points, on the other hand, are earned by users who provide liquidity, based on their LP size and duration. Other Point systems include PolyMarket User Points, Leaderboard Bonus and Market Multipliers. These rewards offer users more competitive edge than other prediction markets. Apart from these rewards, PredictHub features a unique 3-tier referral system, rewarding users with even more as you invite your friends. The more friends you bring, the greater the rewards. On top of these, PredictHub focuses on USDC and a wide-range of yield bearing assets like GLP, gUSDC, and sUSDe, enabling users to optimise their earning while holding assets across Networks. Exciting, right? PredictHub is in its Testnet phase and you can start earning Points Right away 🔅 Here's how to Get Started on PredictHub: 1. Go to 2. Request Faucet 3. Start making predictions and earning Points Easy-Peasy ✅ More Info can be gotten from Predict Hub All eyes are on PredictHub as the fix for the flaws of Prediction Market Protocols. With its unique approach targeting untapped niches that most existing prediction markets have yet to explore, I believe the Protocol has the potential to become a breakout success I will be placing good Predictions to Position 🚀🚀🚀show more

InfoSpace OG
19,674 Aufrufe • vor 1 Jahr
🌟 Familia!!🇵🇷Exciting Partnership Announcement! 🐸🌿 Dear Concho Community, We... are thrilled to announce a significant partnership in our conservation efforts for the Puerto Rican Crested Toad, also known as El Sapo Concho. Dr. Sondra I Vega Castillo, a respected biologist and professor at the University of Puerto Rico and part of the Puerto Rican Crested Toad Work Group, has reached out to us with exciting news. The Puerto Rican Crested Toad Working Group is a dedicated team of experts working tirelessly on education, management, reproduction, and conservation initiatives for the Sapo Concho. With their decades of experience in ensuring the survival of this species against various threats, we are honored to collaborate with them on this vital mission. Mr. Quique Rivera from Acho Studio connected with Dr. Vega Castillo, expressing our interest in supporting the conservation efforts for the Sapo Concho. We are committed to making a positive impact on the species and are grateful for this opportunity to work hand in hand with such esteemed professionals. As part of our commitment to conservation, we will be partnering with the Puerto Rican Crested Toad Conservancy ( a non-profit organization coordinating conservation initiatives for the Sapo Concho. All donations will be directed through this organization to ensure transparency and effectiveness in their use. We are immensely grateful for this opportunity to collaborate and support such a crucial cause alongside Dr. Vega Castillo and the Puerto Rican Crested Toad Working Group. Together, we can make a real difference in protecting this endangered species that holds deep cultural significance. More details to follow. If you would like to donate please send USDC to this wallet: QmqXGV5kTxoLW1ZfXk8wAMrFPkwWV8TundVU9Ajxj8r Thank you for your unwavering support! Sincerely, Team $Concho 🐸💙show more

Concho
65,045 Aufrufe • vor 1 Jahr
A preview of what's next, visualized with Rerun and... PlayCanvas supersplat ✨ (Also, feel free to send me a DM 📩; I’ll be in San Francisco from July 21–29, and I'm looking to meet like-minded folks!) I'm convinced that Gaussian Splats will be an integral part of any data engine as an underlying representation. So I've started putting together a repo that: 1. Given a single image, perform image outpainting 🖼️🖌️ 2. Estimate a monocular depth map on the outpainted image 📏 3. Train a Gaussian Splat initialized from the monocular depth 🎓✨ 4. Warp to new views, perform inpainting on the missing masks -> Train new splat 🔄🎨 This is going to be integrated into exo-egoforge, but I wanted to start with the simple single-image version before moving to a multi-video implementation There's some weirdness in the final rerun visualization, but the trained splat looks great 🎉! This is all based on the very cool VistaDream paper ( .github.io/) More on this next week!show more

Pablo Vela
26,036 Aufrufe • vor 1 Jahr
Fine-tune DeepSeek-OCR on your own language! (100% local) DeepSeek-OCR... is a 3B-parameter vision model that achieves 97% precision while using 10× fewer vision tokens than text-based LLMs. It handles tables, papers, and handwriting without killing your GPU or budget. Why it matters: Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow. DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents. The best part? You can easily fine-tune it for your specific use case on a single GPU. I used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate. ↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you're working with. I've shared the complete guide in the next tweet - all the code, notebooks, and environment setup ready to run with a single click. Everything is 100% open-source!show more

Akshay 🚀
126,122 Aufrufe • vor 8 Monaten
I'll always root for a team that open-sources its... best work, and Robbyant just did it properly. Robbyant, Ant Group's embodied-AI company, released LingBot-Vision, a vision foundation model for robots, and the part I love is the data. They trained it on 161M images, filtered down from 2B raw ones and mostly pulled straight from the open web, with no human labels, no edge detectors, no depth sensors anywhere in the loop. It learns the exact edges of objects from raw pixels. That's roughly a tenth of the data DINOv3 saw, and under a third of the training. And it shows in the results. On depth, working out how far away things are, the 1B model edges out a 7B on NYU-Depth. It also powers LingBot-Depth 2.0, which reads the surfaces cameras usually choke on, glass and mirrors, and halves indoor depth error. LingBot-Vision is fully open. Weights from the 1.1B flagship down to a tiny 21M version, code, and the paper. This is the timeline I want more of. Robbyantshow more

Chubby♨️
48,249 Aufrufe • vor 25 Tagen
Today we’re releasing an early version of a new... feature called Tip Cards. We built Tip Cards after seeing people all over the world tipping each other with Cash Links for their contributions online eg. posting good content, responding to questions, moderating group chats, and many others. Tip Cards simplify this experience by enabling every Code user to create their own personalized Tip Card and accept tips from anyone in the world. To create your Tip Card, simply connect your Twitter/X account in the Code app and your personalized Tip Card will be instantly generated for you. You can then share it as a link or get someone to scan it. (Video below of how to create your Tip Card) We’re rolling out this early version of Tip Cards to get feedback on this new payment type and test the Twitter APIs at scale. We plan to add more features to the tipping experience as we go, so your feedback is appreciated. Please try it out and let us know what you think. If you post your Tip Card and tag Code we’ll send you a tip!show more

Code
68,537 Aufrufe • vor 2 Jahren
Finding alpha shouldn't be so hard. During the bear... market, we dedicated ourselves to making DappRadar better than ever. Now, it's 213.43% easier to discover the next big thing! (Yes, I did the math 😂) We've overhauled navigation, merged rankings, improved layouts, and added powerful features. Today, we're topping it all off with a brand-new homepage that brings everything together! 🏠 Explore top dapps, chains, tokens, blockchain games, NFT collections, DeFi protocols, and so much more—all in one place. There's an entire universe waiting for you. 🌌 We redesigned our rankings, introduced new KPIs, and started using AI to make data and smart contracts more accessible to everyone. We didn't stop there. We've launched Hot Contracts, Quests, Airdrops guides, and many other updates. Our mission? To help both crypto newbies finding their way and seasoned degens hunting for the next moonshot, and to support builders in reaching their communities. 💎🚀 Unlock even more with exclusive features like Hot Contracts and advanced filters by staking $RADAR tokens and becoming PRO. It's our way of giving you the edge in the Web3 space! 🔓 If you haven't visited DappRadar in a while, give it a try—it's a whole NEW DappRadar. We're building this together, and your feedback means the world to us. Let me know what you'd like to see next! 🤝 Thanks for being part of our journey. The best is yet to come! 🚀show more

Dragos Dunica 📡
12,905 Aufrufe • vor 1 Jahr
[VAE] by Hand ✍️ A Variational Auto Encoder (VAE)... learns the structure (mean and variance) of hidden features and generates new data from the learned structure. In contrast, GANs only learn to generate new data to fool a discriminator; they may not necessarily know the underlying structure of the data. The International Conference on Learning Representations (ICLR) this year announced its first ever "Test of Time Award" to recognizes the VAE paper, published 10 years ago. This exercise demonstrates how to calculate a VAE by hand. [1] Given: ↳ Three training examples X1, X2, X3 ↳ Copy training examples to the bottom ↳ The purpose is to train the network to reconstruct the training examples. ↳ Since each target is a training example itself, we use the Greek word "auto" which means "self." This crucial step is what makes an autoencoder "auto." [2] Encoder: Layer 1 + ReLU ↳ Multiply inputs with weights and biases ↳ Apply ReLU, crossing out negative values (-1 -> 0) [3] Encoder: Mean and Variance ↳ Multiply features with two sets of weights and biases ↳ 🟩 The first set predicts the means (𝜇) of latent distributions ↳ 🟪 The second set predicts the standard deviation (𝜎) of latent distributions [4] Reparameterization Trick: Random Offset ↳ Sample epsilon ε from the normal distribution with mean = 0 and variance = 1. ↳ The purpose is to randomly pick a offset away from the mean. ↳ Multiply the standard deviation values with epsilon values. ↳ The purpose is to scale the offset by the standard deviation. [5] Reparameterization Trick: Mean + Offset ↳ Add the sampled offset to predicted mean ↳ The result are new parameters or features 🟨 as inputs to the Decoder. [6] Decoder: Layer 1 + ReLU ↳ Multiply input features with weights and biases ↳ Apply ReLU, crossing out negative values. Here, -4 is crossed out. [7] Decoder: Layer 2 ↳ Multiply features with weights and biases ↳ The output is Decoder's attempt to reconstruct the input data X from reparameterized distributions described by 𝜇 and 𝜎. [8]-[10] KL Divergence Loss [8] Loss Gradient: Mean 𝜇 ↳ We want 𝜇 to approach 0. ↳ A lot of math called SGVB simplifies the calculation of loss gradients to simply 𝜇 [9,10] Loss Gradient: Stdev 𝜎 ↳ We want 𝜎 to approach 1. ↳ A lot of math simplifies the calculation to 𝜎 - (1/ 𝜎) [11] Reconstruction Loss ↳ We want the reconstructed data Y (dark 🟧) to be the same as the input data X. ↳ Some math involving Mean Square Error simplifies the calculation to Y - X.show more

Tom Yeh
48,413 Aufrufe • vor 2 Jahren
BREAKING: As of last night, the DOJ’s “J6er online... registry” has been disabled from the internet! This has been a personal crusade I have worked on for many months. Yesterday I had a follow up call with Congressman Troy Nehls’s Congressman Troy E. Nehls office about this site (I’ve been pushing for its removal since Nov 6th), and was told yesterday afternoon it was being handled by our new US Attorney Eagle Ed Martin. This is a huge victory for J6ers. This site was one of countless weapons of harassment used by the federal government to make life impossible for its targets from J6. The site included every accusation, and every charge leveled against people, including the ones they were not convicted of and were never substantiated in court, and would appear as the top ranking search result. In other words, every time a potential employer, landlord, new social or business contact, etc, would search somebody targeted for J6 they would read a dossier on each person filled with FBI and FOJ accusations and narratives that were never proven, along with links to documents with even more damaging allegations. That site is now joins the trash heap of Joe Biden’s, Merrick Garland’s, Christopher Ray’s, and Matthew Grave’s legacy of corruption. Thank you, Troy Nehls, Ed Martin, and all who worked to get this taken down!show more

Brandon Straka #WalkAway
121,645 Aufrufe • vor 1 Jahr
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft paper... page: Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces an instruction-tuned Video Pretraining (VPT) model for Minecraft called STEVE-1, demonstrating that the unCLIP approach, utilized in DALL-E 2, is also effective for creating instruction-following sequential decision-making agents. STEVE-1 is trained in two steps: adapting the pretrained VPT model to follow commands in MineCLIP's latent space, then training a prior to predict latent codes from text. This allows us to finetune VPT through self-supervised behavioral cloning and hindsight relabeling, bypassing the need for costly human text annotations. By leveraging pretrained models like VPT and MineCLIP and employing best practices from text-conditioned image generation, STEVE-1 costs just $60 to train and can follow a wide range of short-horizon open-ended text and visual instructions in Minecraft. STEVE-1 sets a new bar for open-ended instruction following in Minecraft with low-level controls (mouse and keyboard) and raw pixel inputs, far outperforming previous baselines. We provide experimental evidence highlighting key factors for downstream performance, including pretraining, classifier-free guidance, and data scaling. All resources, including our model weights, training scripts, and evaluation tools are made available for further research.show more

AK
144,783 Aufrufe • vor 3 Jahren