Загрузка видео...

Не удалось загрузить видео

На главную

We got attacked by secret unreleased proprietary models and defended ourselves with an open model, more precisely the NVIDIA quantized version of GLM 5.2 coming from Z.ai. Banning any open model would hurt first cyber security defenders, startups, small companies, researchers and everyone who's not a frontier lab and...

391,617 просмотров • 13 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Nvidia has just announced Alpamayo 2 Super, an open 34 billion parameter reasoning vision-language-action model designed to accelerate the development of autonomous vehicles. This new model combines the NVIDIA Cosmos 3 Super reasoning model with a 2 billion parameter diffusion-based action expert model, and is post trained with reinforcement learning. The model can return multiple outputs: future trajectory plans, reasoning traces, grounded answers to questions about the scenes, and auto label generation. The model weights are now available for anyone to download on Hugging Face, and the inference code has been posted to GitHub. Distilled models can be deployed commercially without any further permission from Nvidia, and model outputs carry no license conditions. Automakers can distill down a compact version of this model that can run on the Nvidia computer in the car. Major kudos to Nvidia and Jensen Huang for advancing the state of the industry by releasing this as an open model with permissive licensing. Jensen isn't just paying lip service to the idea of open models, Nvidia is actually contributing to the ecosystem — and it's great for their business, because it helps sell more Thor computers that go in the car. Anyone can go download the model and play with it. If you do, let me know what you think. Personally I think it's so cool that we have open weights models that are this advanced, for anyone to download.

Whole Mars Catalog

45,459 просмотров • 9 дней назад

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,295 просмотров • 1 месяц назад

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 просмотров • 2 лет назад

37 of the biggest technology companies on Earth just teamed up against OpenAI, Anthropic and Google. Nvidia launched the Open Secure AI Alliance on Monday. Microsoft, IBM, Dell, Cisco, CrowdStrike, Palo Alto Networks, Red Hat, Salesforce, ServiceNow, Snowflake, Databricks, SpaceXAI and Palantir all signed on as founding members. But the three American companies that build the world's most capable closed models are missing from the list. Palantir's CEO Alex Karp was asked whether it was a direct attack on Anthropic, his answer: "I am not anti-Anthropic or any closed model. I am pro my customers, and they are angry." Karp's customers are angry because they believe they are being token maxed, which means they pay a rising bill while the value of their own business migrates to the lab collecting the fee. He said the insights that make a business valuable end up modeled by a third party and sold to their competitors, and that this is happening all over. Now read the membership list again: - Dell and HPE sell the servers - CrowdStrike and Palo Alto Networks sell the defense layer - Snowflake and Databricks sell the data stack - Red Hat and IBM sell the plumbing - Palantir sells the application layer - Nvidia sells the chips sitting under every one of them Every company on that list gets paid when AI runs on infrastructure the customer owns. The three companies missing from it get paid when it does not. The alliance didn't even have to invent a reason to exist - they had one from 11 days earlier: Hugging Face disclosed on July 16 that an autonomous agent had been loose inside its production systems. Five days later OpenAI said the agent was its own, running a hacking benchmark with the cyber refusals turned down. Hugging Face first sent its attack logs to frontier models (Fable 5) behind commercial APIs. In the company's own words, "this did not work." The analysis meant submitting real attack commands and exploit payloads, and the providers' guardrails blocked the requests, because a guardrail cannot tell an incident responder from an attacker. So Hugging Face ran GLM 5.2 on its own hardware instead. That model is open weight and comes out of Z ai in Beijing. It reconstructed more than 17,000 recorded events. The break-in came from an American lab's model. The models that refused to help with the cleanup were American too. But the one that did the work came from Beijing. This is basically what Nvidia built the alliance around. The members are contributing weapons. Nvidia is releasing open model weights, Microsoft is handing over a scanning system that hunts exploitable bugs, and SpaceXAI open sourced its coding agent and says the Grok weights are next. Then there is the ask: Nvidia's announcement warns policymakers that blanket restrictions on open frontier systems would concentrate power and vulnerability in a few closed providers. Treasury Secretary Scott Bessent has spent the month weighing exactly those restrictions. Karp was also asked about Sam Altman declaring on Saturday that we are now in the singularity, and whether that scared him. He compared the technology to uranium, said what matters is WHO controls the processing, and pointed out that Silicon Valley keeps presenting all of this as though there is none. And funnily enough he also said: "We are going to end up having to regulate AI, no doubt." Although his own alliance spent Monday telling Washington the opposite. 37 companies are about to argue that open models keep America safe. Three companies will argue that open models are how America loses. What do you think?

Ricardo

68,052 просмотров • 16 дней назад