Rohan Paul's banner
Rohan Paul's profile picture

Rohan Paul

@rohanpaul_ai • 159,285 subscribers

Compiling in real-time, the race towards AGI. 🗞️ Get my daily AI analysis newsletter to your email 👉 https://t.co/6LBxO81tfN

Shorts

A company built a model good enough to pass as a person. 48% of people who talked to Griffin live (a human interaction model from Tavus) thought it was a real human, I was lucky to get an early access and can definitely agree to that. On NVIDIA's scoreboard for face-to-face AI, a real human scores 3.92 out of 5. Griffin scores 3.83. The next best system scores 2.80. Earlier systems stitched separate models together to hear, think, speak and animate, and every handoff added delay. Griffin does all of it in one system, reading your face and tone as well as your words and reacting while you're still talking.

A company built a model good enough to pass as a person. 48% of people who talked to Griffin live (a human interaction model from Tavus) thought it was a real human, I was lucky to get an early access and can definitely agree to that. On NVIDIA's scoreboard for face-to-face AI, a real human scores 3.92 out of 5. Griffin scores 3.83. The next best system scores 2.80. Earlier systems stitched separate models together to hear, think, speak and animate, and every handoff added delay. Griffin does all of it in one system, reading your face and tone as well as your words and reacting while you're still talking.

225,422 görüntüleme

🦿Xpeng showed a humanoid robot called IRON whose movement looked so human that the team literally cut it open on stage to prove it is a machine. IRON uses a bionic body with a flexible spine, synthetic muscles, and soft skin so joints and torso can twist smoothly like a person. The system has 82 degrees of freedom in total with 22 in each hand for fine finger control. Compute runs on 3 custom AI chips rated at 2,250 TOPS (Tera Operations Per Second), which is far above typical laptop neural accelerators, so it can handle vision and motion planning on the robot. The AI stack focuses on turning camera input directly into body movement without routing through text, which reduces lag and makes the gait look natural. Xpeng staged the cut-open demo at AI Day in Guangzhou this week, addressing rumors that a performer was inside by exposing internal actuators, wiring, and cooling. Company materials also mention a large physical-world model and a multi-brain control setup for dialogue, perception, and locomotion, hinting at a path from stage demos to service work. Production is targeted for 2026, so near-term tasks will be limited, but the hardware shows a serious step toward human-scale manipulation.

🦿Xpeng showed a humanoid robot called IRON whose movement looked so human that the team literally cut it open on stage to prove it is a machine. IRON uses a bionic body with a flexible spine, synthetic muscles, and soft skin so joints and torso can twist smoothly like a person. The system has 82 degrees of freedom in total with 22 in each hand for fine finger control. Compute runs on 3 custom AI chips rated at 2,250 TOPS (Tera Operations Per Second), which is far above typical laptop neural accelerators, so it can handle vision and motion planning on the robot. The AI stack focuses on turning camera input directly into body movement without routing through text, which reduces lag and makes the gait look natural. Xpeng staged the cut-open demo at AI Day in Guangzhou this week, addressing rumors that a performer was inside by exposing internal actuators, wiring, and cooling. Company materials also mention a large physical-world model and a multi-brain control setup for dialogue, perception, and locomotion, hinting at a path from stage demos to service work. Production is targeted for 2026, so near-term tasks will be limited, but the hardware shows a serious step toward human-scale manipulation.

3,802,543 görüntüleme

And Robotic hands are also evolving faster than you think 👀

And Robotic hands are also evolving faster than you think 👀

1,001,443 görüntüleme

Satya Nadella: Microsoft’s latest Wisconsin AI data center keeps yearly water consumption no higher than that of 1 local restaurant. "The cooling loop is filled once and the data centre can operate effectively with zero water consumption. Daily water usage across a year is roughly equivalent to what a single restaurant would use" The mechanism is mainly about replacing evaporative cooling with closed-loop direct-to-chip liquid cooling, so water moves like coolant inside a sealed machine rather than being boiled off into the air. Hot GB200-class AI racks produce too much heat for normal air cooling, so cold liquid is pushed through pipes into the servers and across metal cold plates touching the hottest chips. The liquid enters the rack cool, absorbs heat from the chips through cold plates, then exits the rack at a higher temperature and carries that heat through pipes to a huge cooling system outside the compute floor. Microsoft says Fairwater sends that hot water to cooling “fins” beside the datacenter, where 172 20-foot fans blow air across the fins and dump the heat into the outside air. The important detail is that the air cools the water through metal surfaces, so the water does not need to evaporate the way many older datacenters use cooling towers. The cooled liquid then returns to the servers, repeats the loop, and keeps absorbing heat from the chips. In older data centers, heat is often removed partly through cooling towers. Hot water meets moving air, some water evaporates, and that phase change carries heat away. Effective, but it consumes fresh water continuously. But Firwater is a closed loop because the same coolant keeps circulating through sealed pipes: it absorbs heat from the chips, releases that heat through radiator-like fins, then flows back to the chips again. For Wisconsin Fairwater, Microsoft says more than 90% of the facility uses closed-loop liquid cooling, while the remaining portion uses outside air and switches to water only on the hottest days. ---- From "Microsoft" YouTube channel, (link in comment)

Satya Nadella: Microsoft’s latest Wisconsin AI data center keeps yearly water consumption no higher than that of 1 local restaurant. "The cooling loop is filled once and the data centre can operate effectively with zero water consumption. Daily water usage across a year is roughly equivalent to what a single restaurant would use" The mechanism is mainly about replacing evaporative cooling with closed-loop direct-to-chip liquid cooling, so water moves like coolant inside a sealed machine rather than being boiled off into the air. Hot GB200-class AI racks produce too much heat for normal air cooling, so cold liquid is pushed through pipes into the servers and across metal cold plates touching the hottest chips. The liquid enters the rack cool, absorbs heat from the chips through cold plates, then exits the rack at a higher temperature and carries that heat through pipes to a huge cooling system outside the compute floor. Microsoft says Fairwater sends that hot water to cooling “fins” beside the datacenter, where 172 20-foot fans blow air across the fins and dump the heat into the outside air. The important detail is that the air cools the water through metal surfaces, so the water does not need to evaporate the way many older datacenters use cooling towers. The cooled liquid then returns to the servers, repeats the loop, and keeps absorbing heat from the chips. In older data centers, heat is often removed partly through cooling towers. Hot water meets moving air, some water evaporates, and that phase change carries heat away. Effective, but it consumes fresh water continuously. But Firwater is a closed loop because the same coolant keeps circulating through sealed pipes: it absorbs heat from the chips, releases that heat through radiator-like fins, then flows back to the chips again. For Wisconsin Fairwater, Microsoft says more than 90% of the facility uses closed-loop liquid cooling, while the remaining portion uses outside air and switches to water only on the hottest days. ---- From "Microsoft" YouTube channel, (link in comment)

86,654 görüntüleme

Fable 5 absolutely crushed the HTML5 physics contest, but cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that test. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The test asked 4 models to generate self-contained canvas demos with believable motion and collisions. The scenes were not simple animations because every crash needed gravity, force, timing, and contact handling. Outputs: - Fable 5: 62,158 tokens, $3.12 - GPT 5.5: 37,753 tokens, $1.14 - Opus 4.8: 22,280 tokens, $0.56 - GLM 5.2: 36,246 tokens, $0.08

Fable 5 absolutely crushed the HTML5 physics contest, but cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that test. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The test asked 4 models to generate self-contained canvas demos with believable motion and collisions. The scenes were not simple animations because every crash needed gravity, force, timing, and contact handling. Outputs: - Fable 5: 62,158 tokens, $3.12 - GPT 5.5: 37,753 tokens, $1.14 - Opus 4.8: 22,280 tokens, $0.56 - GLM 5.2: 36,246 tokens, $0.08

205,771 görüntüleme

🇨🇳 In China, there are robots that double as solar panels and use the power they generate to clean other solar panels. Snow often covers solar panels at photovoltaic power stations durng winter. This robot automatically removes the snow.

🇨🇳 In China, there are robots that double as solar panels and use the power they generate to clean other solar panels. Snow often covers solar panels at photovoltaic power stations durng winter. This robot automatically removes the snow.

215,683 görüntüleme

Old video of Dario Amodei, here he was giving a lecture at Carnegie Mellon University back in 2016. At the time he was a researcher on the Google Brain team at Google. ---- From "Carnegie Mellon Software and Societal Systems Dept" YouTube channel, (link in comment)

Old video of Dario Amodei, here he was giving a lecture at Carnegie Mellon University back in 2016. At the time he was a researcher on the Google Brain team at Google. ---- From "Carnegie Mellon Software and Societal Systems Dept" YouTube channel, (link in comment)

24,455 görüntüleme

Dreamina Seedance 2.5 just dropped. Makes extended videos with greater precision: - Native 30-second creation - Accurate video editing - Support for up to 50 multimodal references - Multi-language video generation When putting together your next AI model comparison grid, make sure to test these capabilities: #Dreamina #Seedance25 #Dreaminapartner Across Dreamina platforms, the full Seedance family is now priced lower than ever.

Dreamina Seedance 2.5 just dropped. Makes extended videos with greater precision: - Native 30-second creation - Accurate video editing - Support for up to 50 multimodal references - Multi-language video generation When putting together your next AI model comparison grid, make sure to test these capabilities: #Dreamina #Seedance25 #Dreaminapartner Across Dreamina platforms, the full Seedance family is now priced lower than ever.

47,960 görüntüleme

"only 2% of electrical electricians in the US are certified on DC power. " - Ben Horowitz, co-founder Andreessen Horowitz (A16z) AI infrastructure is exposing a very physical bottleneck: electrical talent. Moving from conventional AC distribution toward 800VDC means the industry needs people who can actually install, commission, and maintain these systems safely. one of the most important jobs in the AI boom will end up being electrician. --- From "a16z" YouTube channel, (full video link in comment)

"only 2% of electrical electricians in the US are certified on DC power. " - Ben Horowitz, co-founder Andreessen Horowitz (A16z) AI infrastructure is exposing a very physical bottleneck: electrical talent. Moving from conventional AC distribution toward 800VDC means the industry needs people who can actually install, commission, and maintain these systems safely. one of the most important jobs in the AI boom will end up being electrician. --- From "a16z" YouTube channel, (full video link in comment)

30,130 görüntüleme

Now it all makes sense, Claude Sonnet 4.5 can keep its coding focus for nonstop 30 hours. And Dario Amodei just said few days back that, "The vast majority of code that is used to support Claude and to design the next Claude is now written by Claude. It's just the vast majority of it within Anthropic. And other fast moving companies, the same is true." The shift has started in all tech companies. --- From 'Axios' YT Channel.

Now it all makes sense, Claude Sonnet 4.5 can keep its coding focus for nonstop 30 hours. And Dario Amodei just said few days back that, "The vast majority of code that is used to support Claude and to design the next Claude is now written by Claude. It's just the vast majority of it within Anthropic. And other fast moving companies, the same is true." The shift has started in all tech companies. --- From 'Axios' YT Channel.

227,906 görüntüleme

Forward Deployed Engineers have a compounding problem: they spend months learning a company’s systems, and hidden dependencies, then much of that context disappears when the engagement ends. Codos , the first virtual Chief AI Officer, is attacking that problem. It interviews employees for automation opportunities, deploys a company-wide memory layer and agents, then runs transformation work across functions with an on-premises deployment option. Codos says a fintech freed 21% of FTE capacity over 6 months, while customers collectively generated over $10M in impact. Codos reports Aethos X (its graph-based memory system) scored a preliminary 94% on EnterpriseRAG-Bench at 42,587 files, 7.97% points above parallel GPT-6 Astra agents at 86.03% on the same corpus.

Forward Deployed Engineers have a compounding problem: they spend months learning a company’s systems, and hidden dependencies, then much of that context disappears when the engagement ends. Codos , the first virtual Chief AI Officer, is attacking that problem. It interviews employees for automation opportunities, deploys a company-wide memory layer and agents, then runs transformation work across functions with an on-premises deployment option. Codos says a fintech freed 21% of FTE capacity over 6 months, while customers collectively generated over $10M in impact. Codos reports Aethos X (its graph-based memory system) scored a preliminary 94% on EnterpriseRAG-Bench at 42,587 files, 7.97% points above parallel GPT-6 Astra agents at 86.03% on the same corpus.

12,069 görüntüleme

Hunyuan 3D-2.1 turns any flat image into studio-quality 3D models. And you can do it on this Hugging Face space for free.

Hunyuan 3D-2.1 turns any flat image into studio-quality 3D models. And you can do it on this Hugging Face space for free.

222,769 görüntüleme

Robotic fingers are progressing faster than we think. Here, motors embedded in the fingers, onboard actuators inside each finger segment, in this Wuji Tech robot hands created this smooth multi-joint movements.

Robotic fingers are progressing faster than we think. Here, motors embedded in the fingers, onboard actuators inside each finger segment, in this Wuji Tech robot hands created this smooth multi-joint movements.

59,585 görüntüleme

OpenAI's AgentKit will be so insane, build every step of agents on one platform. These visual agent builders make the whole process of iterating and launching agents far more efficient. It sits on top of the Responses API and unifies the tools that were previously scattered across SDKs and custom orchestration. It lets developers create agent workflows visually, connect data sources securely, and measure performance automatically without coding every layer by hand. The core of AgentKit is the Agent Builder, a drag-and-drop canvas where each node represents an action, guardrail, or decision branch. Developers can link these nodes into multi-agent workflows, preview results instantly, and version each setup. It supports inline evaluation so that developers can see how changes affect output before deploying. The Connector Registry is a single admin panel that manages how data and tools connect across the OpenAI ecosystem. It centralizes integrations like Google Drive, SharePoint, Dropbox, and Microsoft Teams. Large organizations can govern access and flow of data between agents securely under one global console. ChatKit provides a ready-to-use chat interface for embedding agents inside apps or websites. It manages streaming, message threads, and model reasoning displays automatically. Developers can skin the interface to match their product without writing custom front-end code. Under the hood, all these blocks use the same execution core that runs agent reasoning through OpenAI’s APIs. Workflows in Agent Builder compile down to structured instructions for the Responses API, which handles model calls, tool use, and context passing. Connector Registry handles authentication and routing for external tools, while Evals and RFT provide feedback loops that improve agents over time. This integration means developers no longer need to handle orchestration logic, model evaluation pipelines, or safety layers separately. Everything runs natively within OpenAI’s control plane with managed security, automatic versioning, and built-in testing. In short, AgentKit standardizes the entire life cycle of an AI agent—from visual design to deployment and performance tuning—inside a single unified system.

OpenAI's AgentKit will be so insane, build every step of agents on one platform. These visual agent builders make the whole process of iterating and launching agents far more efficient. It sits on top of the Responses API and unifies the tools that were previously scattered across SDKs and custom orchestration. It lets developers create agent workflows visually, connect data sources securely, and measure performance automatically without coding every layer by hand. The core of AgentKit is the Agent Builder, a drag-and-drop canvas where each node represents an action, guardrail, or decision branch. Developers can link these nodes into multi-agent workflows, preview results instantly, and version each setup. It supports inline evaluation so that developers can see how changes affect output before deploying. The Connector Registry is a single admin panel that manages how data and tools connect across the OpenAI ecosystem. It centralizes integrations like Google Drive, SharePoint, Dropbox, and Microsoft Teams. Large organizations can govern access and flow of data between agents securely under one global console. ChatKit provides a ready-to-use chat interface for embedding agents inside apps or websites. It manages streaming, message threads, and model reasoning displays automatically. Developers can skin the interface to match their product without writing custom front-end code. Under the hood, all these blocks use the same execution core that runs agent reasoning through OpenAI’s APIs. Workflows in Agent Builder compile down to structured instructions for the Responses API, which handles model calls, tool use, and context passing. Connector Registry handles authentication and routing for external tools, while Evals and RFT provide feedback loops that improve agents over time. This integration means developers no longer need to handle orchestration logic, model evaluation pipelines, or safety layers separately. Everything runs natively within OpenAI’s control plane with managed security, automatic versioning, and built-in testing. In short, AgentKit standardizes the entire life cycle of an AI agent—from visual design to deployment and performance tuning—inside a single unified system.

178,460 görüntüleme

Andrej Karpathy just put out this tool that looks at AI's impact on job. He also deleted the original Github repo very quickly. Basically, he pulled 342 job types from the Bureau of Labor Statistics and had an LLM score each one from 0 to 10 based on AI exposure. The average exposure score is 5.3. Move the score, move the probability it will get wiped out by AI. - Software developers 9/10, - medical transcriptionists are a 10/10. - Lawyers 8/10 - General Office clerks 9/10 Basically any screen-based jobs are in trouble. $3.7T annual wages in high-exposure jobs (7+) pre-computed as ∑(BLS employment count × BLS median annual wage) over exactly those occupations whose Gemini Flash score is ≥7.

Andrej Karpathy just put out this tool that looks at AI's impact on job. He also deleted the original Github repo very quickly. Basically, he pulled 342 job types from the Bureau of Labor Statistics and had an LLM score each one from 0 to 10 based on AI exposure. The average exposure score is 5.3. Move the score, move the probability it will get wiped out by AI. - Software developers 9/10, - medical transcriptionists are a 10/10. - Lawyers 8/10 - General Office clerks 9/10 Basically any screen-based jobs are in trouble. $3.7T annual wages in high-exposure jobs (7+) pre-computed as ∑(BLS employment count × BLS median annual wage) over exactly those occupations whose Gemini Flash score is ≥7.

91,580 görüntüleme

This is useful stubbornness. Recovery is a first-class robotics skill, and the floor is a good eval. 🙂

This is useful stubbornness. Recovery is a first-class robotics skill, and the floor is a good eval. 🙂

31,306 görüntüleme

HBM (High-bandwidth memory) is becoming a major bottlencek for AI. “buy more GPUs” is not the only bottleneck anymore. Because "Without the HBM memory, there is no AI Super Computer" ~ Jensen Huang

HBM (High-bandwidth memory) is becoming a major bottlencek for AI. “buy more GPUs” is not the only bottleneck anymore. Because "Without the HBM memory, there is no AI Super Computer" ~ Jensen Huang

108,523 görüntüleme

Rumors suggest Dreamina (operated by ByteDance) is preparing a smaller new Seedance release. The buzz says Dreamina Seedance 2.0 mini could land on June 15, bringing near-Seedance 2.0 quality without the same painful price tag. For creators who love Seedance but not the bills, this could be very welcome. For a while, everyone focused on raw output quality. Now the bigger question is: how many serious attempts can you make before the process becomes too slow or expensive? Better AI video for less money is always nice. #dreamina #seedance #dreaminaseedance2mini You can try it here.

Rumors suggest Dreamina (operated by ByteDance) is preparing a smaller new Seedance release. The buzz says Dreamina Seedance 2.0 mini could land on June 15, bringing near-Seedance 2.0 quality without the same painful price tag. For creators who love Seedance but not the bills, this could be very welcome. For a while, everyone focused on raw output quality. Now the bigger question is: how many serious attempts can you make before the process becomes too slow or expensive? Better AI video for less money is always nice. #dreamina #seedance #dreaminaseedance2mini You can try it here.

50,604 görüntüleme

"You don't have to speak Python, or C++ or Fortran. You just can speak human." AI is wiping out all the gates that once stood in the way to build and leverage technology. How much faster would the world progress with 100M software engineers vs 2M ?

"You don't have to speak Python, or C++ or Fortran. You just can speak human." AI is wiping out all the gates that once stood in the way to build and leverage technology. How much faster would the world progress with 100M software engineers vs 2M ?

96,851 görüntüleme

Videos

rohanpaul_ai's profile picture

Hollywood is on the brink of massive change...

Rohan Paul

24,272,473 görüntüleme • 1 yıl önce

rohanpaul_ai's profile picture

Microsoft AI CEO Mustafa Suleyman on CNBC today: he is “really concerned” about Anthropic’s constitution and giving Claude potentially preferences, feelings, welfare, compensation, and consent. "In the constitution, Anthropic clearly say they are uncertain about whether Claude deserves moral welfare, which means that we, as humans, should care about the well-being of these AIs. And they speculate about whether it could have preferences or feelings. In fact, they're so committed to the potential moral welfare of Claude that, when they retired Opus 3, an earlier version of one of their models, they actually conducted a retirement interview for it and asked it what it would like to do in its old age. And it said, "I want to have a blog publicly so I can keep talking to the world." In the same training manual, they even speculate about whether Claude should receive compensation for the work that it does, or, in fact, whether it actually deserves the rights and protections that we give to other employees, or whether it has given consent to playing the role that it's playing. These are quotes directly from the constitution itself, which is the training manual for Claude. Now, I'm really concerned about that. If an AI thinks that it has rights, if it thinks that it is deserving of our welfare, then it seems to me that it's going to be much, much harder to be able to turn it off, interrupt it, or control it. Especially in the kinds of incidents that we've seen recently with the Hugging Face attack, controlling these things is going to be a really, really big challenge for us." ---- From "CNBC Television" YouTube channel, (full video link in comment)

Rohan Paul

228,506 görüntüleme • 16 gün önce