Загрузка видео...

Не удалось загрузить видео

На главную

🚨Gemini 3.5 Pro leak just dropped and the frontend game is actually different. Targeting July 17 release with: • 2M token context window (this might be the real game changer) • New “Deep Think” reasoning mode • Significantly better UI & frontend generation • Cleaner design taste + stronger...

91,230 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

New short course: LLMs as Operating Systems: Agent Memory, created with Letta, and taught by its founders Charles Packer and Sarah Wooders. An LLM's input context window has limited space. Using a longer input context also costs more and results in slower processing. So, managing what's stored in this context window is important. In the innovative paper MemGPT: Towards LLMs as Operating Systems, its authors (which include the instructors) proposed using an LLM agent to manage this context window. Their system uses a large persistent memory that stores everything that could be included in the input context, and an agent decides what is actually included. Take the example of building a chatbot that needs to remember what's been said earlier in a conversation (perhaps over many days of interaction with a user). As the conversation's length grows, the memory management agent will move information from the input context to a persistent searchable database; summarize information to keep relevant facts in the input context; and restore relevant conversation elements from further back in time. This allows a chatbot to keep what's currently most relevant in its input context memory to generate the next response. When I read the original MemGPT paper, I thought it was an innovative technique for handling memory for LLMs. The open-source Letta framework, which we'll use in this course, makes MemGPT easy to implement. It adds memory to your LLM agents and gives them transparent long-term memory. In detail, you’ll learn: - How to build an agent that can edit its own limited input context memory, using tools and multi-step reasoning - What is a memory hierarchy (an idea from computer operating systems, which use a cache to speed up memory access), and how these ideas apply to managing the LLM input context (where the input context window is a "cache" storing the most relevant information; and an agent decides what to move in and out of this to/from a larger persistent storage system) - How to implement multi-agent collaboration by letting different agents share blocks of memory This course will give you a sophisticated understanding of memory management for LLMs, which is important for chatbots having long conversations, and for complex agentic workflows. Please sign up here!

Andrew Ng

201,127 просмотров • 1 год назад

Gemini-1.5 Pro has its spotlight stolen today, and people are poking fun at Sora vs Google memes. Well, I think it's the biggest boost in LLM capability so far in 2024. v1.5's 10M token context (1) excels at retrieval; (2) generalizes zero-shot to extremely long instructions like full tutorials and codebases; and (3) works across modalities such as text, audio, and video. Here's a stunning example: v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning. I talked about the Myth of Context Length many times before: don't get too excited by claims of 1M or even 1B context tokens. LSTMs already achieved literally infinite context length 25 yrs ago! What truly matters is how well the model actually uses the context to solve real-world problems, and Gemini-1.5 has surpassed the SOTA with flying colors. The paper is also well-written with lots of solid quantitative analysis on in-context memorization and generalization. Paper: “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context” Congrats to Jeff Dean Oriol Vinyals Sundar Pichai and team!

Jim Fan

278,517 просмотров • 2 лет назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 3 месяцев назад

A deep dive into $SEXY and why I believe this is one of the most undervalued gamefi projects in the ecosystem: A simple beginning of a frictionless gaming experience fitting itself into a mobile game market which produced $123B in revenue in 2023 and an expected groth into $190B into 2030. The current game features a simple UI with strong incentives such a $1M prize pool in tokens lasting from June to September for early users pre-layer 3 buildout. Though simple in design, the transaction volume is telling with 300k user txns within 3 days of opening. The game is very much similar to the old micro transaction mini games such as clash of clans. Daily activities to keep users engaged and retained especially given the incentivization program. My thesis around this is the growth reward potential. Currently as the game sits in a more private beta (only users with codes can participate), we are at the simple version of the ecosystem. As expansion happens and there is more use items, more mini-games, more access, and more use for $SEXY, the value will grow. I imagine the R/R of investing in a game like clash of clans early on, imagine being able to invest in the value of gems in the early days and the value of these changed with user demand. Clash of clans did $360m in revenue in 2023. This model would recreate that volume all flowing into the token (which would be similar to gems). Thus, the thesis is simple: $SEXY ecosystem grows, more usage for active players to use the token within the mini games, more addictive tendencies lead to users wanting to gain more loot, more referrals to beat their friends in the ecosystem all leads to a flywheel of growth; especially, as users are token holders and their dominant strategy is to play the game for incentives and to invite their friends to grow the token value and ecosystem. Personally, I will be staking my $SEXY, collecting loot (imagine the loot can be tradable like early Runescape party hats), building my ranking for the incentives, and encouraging the ecosystem growth as a user in the flywheel. Thus, I believe as time passes, the value of the token continually increases under the assumption the ecosystem is growing with it. Speculation transitions from I think future buyers will find this token a good investment into I think players will buy this token to play the games within the ecosystem. Disclaimer: I am an early investor in ETHXY and so is WWVentures. All thesis generation may be biased.

Crypto Max

23,392 просмотров • 2 лет назад

Anthropic just changed how they think about Claude 5. Not with a new model. Not with a benchmark. With a completely different philosophy for building AI systems. Most people will miss it. They're still trying to write better prompts. Anthropic is optimizing something else entirely: Context. Here's what every AI builder should learn from it: 1. Stop telling AI exactly what to do. Start telling it what success looks like. Older models needed rigid instructions. Newer models perform better when you define the objective and let them make the decisions. The goal is no longer more control. It's more clarity. 2. Context is a budget, not a storage unit. Every extra sentence competes for the model's attention. A massive context file doesn't make AI smarter. It often makes reasoning worse. The best systems don't load everything. They load only what's relevant. 3. Great tools beat great prompts. Most people spend hours tweaking prompt wording. Anthropic is investing in better interfaces instead. Clear parameters. Structured inputs. Well-designed tools. If the interface removes ambiguity, the model makes better decisions before it even starts reasoning. 4. Show reality instead of describing it. Want a specific coding style? Provide the codebase. Want a certain design language? Share the mockup. Want consistent outputs? Give it tests to satisfy. Concrete references consistently outperform long written instructions. 5. Deliver information when it's needed. Not everything belongs in the initial context. Documentation. Style guides. Verification. Reviews. Load them only when they're relevant. Think of context like RAM, not a hard drive. 6. The real advantage is system design. Prompt engineering isn't disappearing. It's becoming one small part of a much bigger stack. Memory. Retrieval. Context architecture. Tool design. Evaluation. Every improvement compounds. The biggest takeaway from Anthropic's update? The next generation of AI won't be won by the people writing the longest prompts. It'll be won by the people building the smartest systems. Same models. Completely different results.

Evan Luthra

38,751 просмотров • 1 месяц назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 просмотров • 1 месяц назад