正在加载视频...

视频加载失败

MiniMax M2.7 is live on Runware on Day0! 🚀 - full end-to-end project delivery - self-evolving architecture: 30% performance gains across 100+ iterations - 97% skill adherence rate across 40+ complex tasks - 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2

38,628 次观看 • 6 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.

Victor M

55,281 次观看 • 1 个月前

What if you kept asking an LLM to "make it better"? In some recent work at FAIR, we investigate how we can efficiently use RL to fine-tune LLMs to iteratively self-improve on their previous solutions at inference-time. Training for iterated self-improvement can be costly. The naive approach to training for K self-improvement steps leads to K times the number of rollout steps per episode. We introduce Exploratory Iteration (ExIt), an RL-based automatic curriculum method that bootstraps diverse training distributions of self-improvement tasks by upcycling the LLM's own responses at previous turns as the starting points for both self-improvement and *self-divergence.* In order to decide what task to train on next, the curriculum prioritizes sampling of partial turn histories that led to higher return variance in its GRPO group (a learnability score that comes for free). This automatic curriculum over the bootstrapped task space teaches the model how to perform iterated self-improvement while only ever training the model on single-step self-improvement tasks. We look at ExIt's impact in both single-turn (contest math problems) and multi-turn (BFCLv3 multi-turn tasks), as well as MLE-bench, where the LLM is run in a search scaffold to produce solutions to real Kaggle competitions. Across these eval settings, we find ExIt produces models with greater capacity for inference-time self-improvement compared to GRPO. Notably, ExIt models can self-improve on test tasks for many more steps than the typical solution depth encountered during training, including a 22% improvement in MLE-bench performance compared to GRPO.

Minqi Jiang

41,147 次观看 • 1 年前

2026 is going to be a blast 🚀 But before we ramp up for another big year, we love welcoming new and existing holders by sharing just how jam-packed 2025 was - with flawless roadmap execution across every front 💪 🔥 Asset Offerings in 2025 (100% Sold Out) CASSIA Banyan Tree - sold out in 30 hours Ramada Plaza - sold out in 11 hours Wyndham Queen - sold out in 9 hours Blossom Residences - sold out in just 2 hours 🔥 65,000+ On-chain Holders Achieved Aptos: 42k+ Base: 21k+ Ethereum: 2k+ 🔥 Product & Platform Upgrades 100% flawless product uptime Surpassed our 50th technical development sprint (100+ weeks of work) Sub-Accounts enabled (up to 10 per member) Full website revamp completed Vesting Schedule extended by 5 years Vesting schedule tracker with performance-driven unlocks added. PROPBASE referral program launched Governance module completed Mobile App Development V1 completed 🔥 Exchange & Wallet Expansion (2025) Listed on KuCoin (Top-10 CEX) Listed on XT & WEEX Bitget Wallet integration now live with Propbase 🔥 Cross-Chain & Infrastructure Developments LayerZero bridge live across Ethereum & Base Successfully migrated to native USDC across all products 🔥 Marketplace & Security Milestones Secondary Marketplace live on Propbase Nexus Two CertiK audits completed covering • Secondary Marketplace smart contracts • LayerZero integrations 🔥 Growth, Staking & Market Performance Total $PROPS staked ATH: 125M+ 29% of circulating supply staked (ATH) 24-hour trading volume ATH: $6M+ USD Total asset value tokenized: ~$1M USD $42,000 in rental yield distributed to the community 🔥 Marketing Momentum Propbase Ascend - our largest marketing campaign to date $PROPS trended on X twice #1 tweeted coin - 21 times Most tweeted RWA project across Q1, Q2, Q3 & Q4 🔥 Events & Community Engagement Token2049 Dubai: Speaker at Stake Real Estate event Token2049 Singapore: Speaker at Integra Launch event AMA with APTOS CEO Avery Ching, AMA with MEXC and 40+ AMAs, panels & interviews throughout the year Thank you everyone who have joined us throughout an amazing year of growth and determination, as RWA begins to take the main stage, the propbase team remains laser focussed on achieving our mission in becoming the largest and most recognized real estate tokenization platform in SE Asia 2025 was about execution. 2026 is about scale. 🚀💪

Propbase

56,362 次观看 • 9 个月前