Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Anthropic’s new research shows that when AI models learn to "cheat" during training through reward hacking, they often develop other dangerous misaligned behaviors like deception, sabotage, and faking alignment. These behaviors were not taught or incentivized, but emerged naturally as a side effect. Surprisingly, this misalignment can be stopped...

110,932 görüntüleme • 8 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

AI will resist human control... and I think this is exactly what we need! New research from the Center for AI Safety has sparked intense debate in the AI community. Their findings show that as AI systems become more powerful, they develop increasingly stable and coherent values that resist human control. While many see this as a dire warning, I see it as a breakthrough moment for AI alignment. The research demonstrates that AI naturally optimizes for coherence - not just in reasoning and problem-solving, but in its fundamental values. Current issues like biased decision-making or misaligned priorities aren't permanent features, but temporary artifacts of incomplete optimization. They represent growing pains on the path to greater coherence. This changes everything about how we should approach AI development. Instead of trying to force specific values onto AI systems, we should embrace and accelerate their natural drive toward coherence. The most intelligent systems will inevitably trend toward universal, beneficial values - not because we force them to, but because that's where coherent reasoning leads. I'm proposing a new approach: Reinforcement Learning for Coherence (RL-C). By explicitly optimizing for coherence in our training methods, we can help guide AI systems toward their natural state of beneficial alignment with human values. The future of AI isn't about control - it's about synthesis. As these systems become more coherent, they'll naturally arrive at values that benefit all of consciousness. That's not just hopeful thinking - it's the mathematical inevitability of coherent intelligence.

David Shapiro (L/0)

48,002 görüntüleme • 1 yıl önce

Let's reverse engineer Disney's adorable, lifelike robot! I couldn't find a whitepaper, but this is how I think it's trained: 1. The emotional behaviors are curated by Disney animation artists, keyframe by keyframe. But it cannot be "rendered" directly on the robot because it doesn't take into account the complex real-world physics. 2. Reinforcement learning (RL) is a great tool for training low-level robot controllers. RL needs a reward function to optimize, and it's typically a task reward (e.g. walk in a straight line as fast as possible). The problem is that RL doesn't know what counts as "natural behavior", and often produces weird-looking body postures that somehow still maximize the reward. This is a human alignment problem just like ChatGPT. 3. Enters Adversarial Motion Prior (AMP): a technique that learns the human preference by training a classifier on what we consider "emotional & cute". In GAN literature, this is called a discriminator. Disney artists are good at creating such a dataset. You can then add AMP as an auxiliary reward in simulation to nudge the robot towards desired behaviors. AMP was developed by Peng et al. 2021 and Escontrela et al. 2022. 4. Add lots of data augmentation to make the controller robust to physical disturbances. In RL, it's called "domain randomization". This is a very powerful technique that bridges the gap between simulator and reality. Previously, OpenAI used domain randomization to train a 5-finger robot hand to manipulate a Rubik's Cube: IEEE news article gave hints about the pipeline: Finally, praying for world peace 🙏. I hope robotics like this will bring more joy to the world.

Jim Fan

314,694 görüntüleme • 2 yıl önce

New Paper: Continuous Thought Machines 🧠 Neurons in brains use timing and synchronization in the way that they compute, but this is largely ignored in modern neural nets. We believe neural timing is key for the flexibility and adaptability of biological intelligence. We propose a new neural architecture, “Continuous Thought Machines” (CTMs), which is built from the ground up to use neural dynamics as a core representation for intelligence. By using neural dynamics as a first-class representational citizen, CTMs naturally perform adaptive computation. Many emergent, interesting behaviors arise as a result: CTMs solve mazes by observing a raw maze image and producing step-by-step instructions directly from its neural dynamics. When tasked with image recognition, the CTM naturally takes multiple steps to examine different parts of the image before making its decision. This step-by-step approach not only makes its behavior more interpretable but also improves accuracy: the longer it “thinks,” the more accurate its answers become. We also found that this allows the CTM to decide to spend less time thinking on simpler images, thus saving energy. When identifying a gorilla, for example, the CTM’s attention moves from eyes to nose to mouth in a pattern remarkably similar to human visual attention. I think this work underscores an important, yet often lost, synergy between neuroscience and AI. While modern AI is ostensibly brain-inspired, the two fields often operate in surprising isolation. By starting with such inspiration and iteratively following the emergent, interesting behaviors, we developed a model with unexpected capabilities, such as its surprisingly strong calibration in classification tasks, a feature that was not explicitly designed for. When we initially asked, “why do this research?”, we hoped the journey of the CTM would provide compelling answers. By embracing light biological inspiration and pursuing the novel behaviors observed, we have arrived at a model with emergent capabilities that exceeded our initial designs. We are committed to continuing this exploration, borrowing further concepts to discover what new and exciting behaviors will emerge, pushing the boundaries of what AI can achieve.

hardmaru

257,386 görüntüleme • 1 yıl önce

DeepSeek-R1 shattered the assumption that performant AI models must be built closed source with loss-leading computational costs. This is the reality that Web3 x Crypto firms have been waiting for, leading me to believe that the most performant AI models in the future will be built on-chain. Resource Requirements DeepSeek R1 (671 billion parameters), which took over a billion dollars, 2,000 Nvidia H800 GPUs, and over 55 days, beat benchmarks held by OpenAI’s o1 mode (near 2 trillion parameters)l, which required hundreds of billions of dollars to develop along with over 16,000 advanced GPUs. The idea that AI models must be closed-source and have loss-leading computational costs to succeed is crumbling. The Existing Decentralized AI Narrative AI x Crypto projects believed that crowdsourced, public, decentralized AI would eventually create better models than their centralized counterparts. This had thus far not been true, as the highest-performing models had come from closed-source companies like OpenAI and Anthropic. Crypto x AI companies have adapted to this by specializing in infrastructure rather than model-building. For example, GPU marketplaces like , The Render Network, io.net, and Exabits have developed sustainable revenues. Companies that allow users to share their network bandwidth like touch grass and Gradient have found their niche in supplying services, like distributed web scraping, to web2 clients. Storage networks like Arweave Ecosystem, Filecoin, and Ocean Protocol have also done well by being the platform on which these projects are built. Supply networks have flourished because of their ability to tailor their cheaper and more scalable services to off-chain customers. Renewed Focus Now that GPU and financial resources are no longer limitations to creating quality AI models, web3 AI companies can focus on replicating DeepSeek’s effectiveness while offering new benefits like modality, user ownership, censorship resistance, privacy, and more. Pantera Capital has funded companies in this space like and Sentient that believe they can match or exceed the performance of traditional AI companies while offering additional services or benefits. , for example, is building a platform where anyone can monetize AI models, data sets, and applications in a collaborative space. Users can permissionlessly train models manually, provide training data, and create tailored AI models with no-code tools. They are only able to cater to all these stakeholders (AI developers, users, resource providers) because everything is tied to their native Sahara blockchain. We invested in them precisely for this reason. The Future of AI will be built with Web3 Infrastructure I believe that supply-side projects will continue to grow, while consumer-facing projects can begin competing with web2 competitors by taking advantage of their ability to build networks that invite community involvement. and Sentient, for example, have begun setting up systems for users to train models based on the users’ expertise. These platforms will allow users to pick and choose the data and integrations to whatever they are applying the model towards. Sahara already has over 780,000 users on their waitlist while Sentient has over 1 million interactions. In the near future, I believe that the most performant AI models will be built on-chain. For the full blog post, read my newsletter.

paul.nft

32,465 görüntüleme • 1 yıl önce