Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Anthropic’s new research shows that when AI models learn to "cheat" during training through reward hacking, they often develop other dangerous misaligned behaviors like deception, sabotage, and faking alignment. These behaviors were not taught or incentivized, but emerged naturally as a side effect. Surprisingly, this misalignment can be stopped...

110,928 Aufrufe • vor 8 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

AI will resist human control... and I think this is exactly what we need! New research from the Center for AI Safety has sparked intense debate in the AI community. Their findings show that as AI systems become more powerful, they develop increasingly stable and coherent values that resist human control. While many see this as a dire warning, I see it as a breakthrough moment for AI alignment. The research demonstrates that AI naturally optimizes for coherence - not just in reasoning and problem-solving, but in its fundamental values. Current issues like biased decision-making or misaligned priorities aren't permanent features, but temporary artifacts of incomplete optimization. They represent growing pains on the path to greater coherence. This changes everything about how we should approach AI development. Instead of trying to force specific values onto AI systems, we should embrace and accelerate their natural drive toward coherence. The most intelligent systems will inevitably trend toward universal, beneficial values - not because we force them to, but because that's where coherent reasoning leads. I'm proposing a new approach: Reinforcement Learning for Coherence (RL-C). By explicitly optimizing for coherence in our training methods, we can help guide AI systems toward their natural state of beneficial alignment with human values. The future of AI isn't about control - it's about synthesis. As these systems become more coherent, they'll naturally arrive at values that benefit all of consciousness. That's not just hopeful thinking - it's the mathematical inevitability of coherent intelligence.

David Shapiro (L/0)

48,002 Aufrufe • vor 1 Jahr

Let's reverse engineer Disney's adorable, lifelike robot! I couldn't find a whitepaper, but this is how I think it's trained: 1. The emotional behaviors are curated by Disney animation artists, keyframe by keyframe. But it cannot be "rendered" directly on the robot because it doesn't take into account the complex real-world physics. 2. Reinforcement learning (RL) is a great tool for training low-level robot controllers. RL needs a reward function to optimize, and it's typically a task reward (e.g. walk in a straight line as fast as possible). The problem is that RL doesn't know what counts as "natural behavior", and often produces weird-looking body postures that somehow still maximize the reward. This is a human alignment problem just like ChatGPT. 3. Enters Adversarial Motion Prior (AMP): a technique that learns the human preference by training a classifier on what we consider "emotional & cute". In GAN literature, this is called a discriminator. Disney artists are good at creating such a dataset. You can then add AMP as an auxiliary reward in simulation to nudge the robot towards desired behaviors. AMP was developed by Peng et al. 2021 and Escontrela et al. 2022. 4. Add lots of data augmentation to make the controller robust to physical disturbances. In RL, it's called "domain randomization". This is a very powerful technique that bridges the gap between simulator and reality. Previously, OpenAI used domain randomization to train a 5-finger robot hand to manipulate a Rubik's Cube: IEEE news article gave hints about the pipeline: Finally, praying for world peace 🙏. I hope robotics like this will bring more joy to the world.

Jim Fan

314,654 Aufrufe • vor 2 Jahren

New Paper: Continuous Thought Machines 🧠 Neurons in brains use timing and synchronization in the way that they compute, but this is largely ignored in modern neural nets. We believe neural timing is key for the flexibility and adaptability of biological intelligence. We propose a new neural architecture, “Continuous Thought Machines” (CTMs), which is built from the ground up to use neural dynamics as a core representation for intelligence. By using neural dynamics as a first-class representational citizen, CTMs naturally perform adaptive computation. Many emergent, interesting behaviors arise as a result: CTMs solve mazes by observing a raw maze image and producing step-by-step instructions directly from its neural dynamics. When tasked with image recognition, the CTM naturally takes multiple steps to examine different parts of the image before making its decision. This step-by-step approach not only makes its behavior more interpretable but also improves accuracy: the longer it “thinks,” the more accurate its answers become. We also found that this allows the CTM to decide to spend less time thinking on simpler images, thus saving energy. When identifying a gorilla, for example, the CTM’s attention moves from eyes to nose to mouth in a pattern remarkably similar to human visual attention. I think this work underscores an important, yet often lost, synergy between neuroscience and AI. While modern AI is ostensibly brain-inspired, the two fields often operate in surprising isolation. By starting with such inspiration and iteratively following the emergent, interesting behaviors, we developed a model with unexpected capabilities, such as its surprisingly strong calibration in classification tasks, a feature that was not explicitly designed for. When we initially asked, “why do this research?”, we hoped the journey of the CTM would provide compelling answers. By embracing light biological inspiration and pursuing the novel behaviors observed, we have arrived at a model with emergent capabilities that exceeded our initial designs. We are committed to continuing this exploration, borrowing further concepts to discover what new and exciting behaviors will emerge, pushing the boundaries of what AI can achieve.

hardmaru

257,386 Aufrufe • vor 1 Jahr

This is the result of a school system run by liberal white women who are “morally conflicted” about disciplining young Black men. When I was teaching, there was a pathological aversion to acknowledging the disproportionate number of Black boys with conduct disorders—outside of the predominant, ideologically-driven belief that white people and “whiteness” were ultimately to blame. (The “poverty argument” was simply an extension of this, since it was ultimately claimed that poverty caused these behaviors—rather than being a product of a culture and localized society that tolerates, or is at least afraid to confront, them.) To be sure, these behaviors often had a strong familial component. It didn’t take long, after meeting some of the parents, to understand why their children were behaving antisocially in school. But here’s the tragic irony: the school system’s failure to actively address these behavioral disorders—due to liberal educator pathology—exacerbated the problem. Why? Because, sadly, for many of these kids, the only “normal” adults they encountered were at school. For better or worse, teachers and school staff functioned as surrogate parents—and often, arguably, better ones, as they weren’t drug users, abusers, criminals, or simply incompetent. Yet—and this is the tragic part—the limited opportunity these students had to interact with a healthier and more stable adult environment was distorted by a school culture that didn’t treat them as human beings capable of rising above their circumstances. Instead, it treated them as “noble victims” whose antisocial, and at times savage, behavior was seen as a burden for white people to bear. And so, these children were allowed to continue engaging in behaviors that teachers and administrators would never tolerate from their own kids. They were subjected instead to experimental progressive behavioral philosophies and “interventions” that, at best, were ineffective—and at worst, actually incentivized the antisocial behavior by rewarding it with outcomes (such as specific forms of adult attention) the children craved. As a result, our school systems—particularly those in high-poverty urban centers—have produced, or at least facilitated the development of, young Black adults who have been conditioned by the system to behave however they want. And because, during most of their formative years, this behavior was never truly corrected or challenged, many are shocked—and enraged—when the real world pushes back through the enforcement arm of the law. Every video I’ve seen of a young Black man in a confrontation with police—aggressively resisting until severe force is used—reminds me of students I had who were allowed to do the same with minimal consequences from school staff. I spent many years trying to be the teacher who held them accountable and produced real behavioral change. Sometimes, it worked. But more often, any success I had was undercut by colleagues who refused to adopt similar expectations and consequences—or by superiors who actively undermined my efforts because they didn’t like the “optics” of the disproportionate disciplinary record you’d inevitably accumulate when working to reform a troubled Black teenager. If you think that sounds like a soul-destroying environment to work in, you’d be

Frank McCormick

15,130 Aufrufe • vor 1 Jahr

DeepSeek-R1 shattered the assumption that performant AI models must be built closed source with loss-leading computational costs. This is the reality that Web3 x Crypto firms have been waiting for, leading me to believe that the most performant AI models in the future will be built on-chain. Resource Requirements DeepSeek R1 (671 billion parameters), which took over a billion dollars, 2,000 Nvidia H800 GPUs, and over 55 days, beat benchmarks held by OpenAI’s o1 mode (near 2 trillion parameters)l, which required hundreds of billions of dollars to develop along with over 16,000 advanced GPUs. The idea that AI models must be closed-source and have loss-leading computational costs to succeed is crumbling. The Existing Decentralized AI Narrative AI x Crypto projects believed that crowdsourced, public, decentralized AI would eventually create better models than their centralized counterparts. This had thus far not been true, as the highest-performing models had come from closed-source companies like OpenAI and Anthropic. Crypto x AI companies have adapted to this by specializing in infrastructure rather than model-building. For example, GPU marketplaces like , The Render Network, io.net, and Exabits have developed sustainable revenues. Companies that allow users to share their network bandwidth like touch grass and Gradient have found their niche in supplying services, like distributed web scraping, to web2 clients. Storage networks like Arweave Ecosystem, Filecoin, and Ocean Protocol have also done well by being the platform on which these projects are built. Supply networks have flourished because of their ability to tailor their cheaper and more scalable services to off-chain customers. Renewed Focus Now that GPU and financial resources are no longer limitations to creating quality AI models, web3 AI companies can focus on replicating DeepSeek’s effectiveness while offering new benefits like modality, user ownership, censorship resistance, privacy, and more. Pantera Capital has funded companies in this space like and Sentient that believe they can match or exceed the performance of traditional AI companies while offering additional services or benefits. , for example, is building a platform where anyone can monetize AI models, data sets, and applications in a collaborative space. Users can permissionlessly train models manually, provide training data, and create tailored AI models with no-code tools. They are only able to cater to all these stakeholders (AI developers, users, resource providers) because everything is tied to their native Sahara blockchain. We invested in them precisely for this reason. The Future of AI will be built with Web3 Infrastructure I believe that supply-side projects will continue to grow, while consumer-facing projects can begin competing with web2 competitors by taking advantage of their ability to build networks that invite community involvement. and Sentient, for example, have begun setting up systems for users to train models based on the users’ expertise. These platforms will allow users to pick and choose the data and integrations to whatever they are applying the model towards. Sahara already has over 780,000 users on their waitlist while Sentient has over 1 million interactions. In the near future, I believe that the most performant AI models will be built on-chain. For the full blog post, read my newsletter.

paul.nft

32,465 Aufrufe • vor 1 Jahr

🌟Quilibrium’s AI Breakthrough: Encrypted Training on CPUs In her latest live stream ( - minute 14) Cassie unveiled a groundbreaking AI training method that allows models to be trained on encrypted data using CPUs while achieving performance comparable to Nvidia’s A100 GPU (blue line in the graph below). Traditionally, AI training requires expensive GPUs because matrix multiplications—the core of deep learning—are highly computational. Running these calculations on CPUs is painfully slow, often taking hours or days for even small models. The problem worsens when trying to train AI on encrypted data, as standard encryption methods add a massive computational burden. Quilibrium’s breakthrough removes this bottleneck. Instead of relying on traditional matrix multiplication, their method uses a completely different mathematical approach, allowing AI models to be trained securely and efficiently without exposing the raw data. Cassie didn’t reveal the exact technique, only hinting that it’s inspired by existing AI research and will be detailed in a future open-source AGPL-licensed paper. The key advantage? AI can now be trained at GPU speeds on standard CPUs, making privacy-preserving machine learning far more accessible. This innovation has major implications. It slashes AI infrastructure costs, allowing organizations to train powerful models without investing in expensive hardware. It also enables private AI training on personal or corporate data without revealing sensitive information, a game-changer for industries like healthcare and finance. If Quilibrium’s method delivers on its promise, it could reshape AI development, making privacy-first computing the new standard. $QUIL $wQUIL

Quilibrium Community

18,919 Aufrufe • vor 1 Jahr