Google DeepMind is doing some crazy work. This is... so good and is going to help so many people. They just launched SL2T (Sign Language to Text), an AI model that translates sign language directly into text. Most spoken language dictation tools rely on audio to text, but sign languages have their own grammars and involve complex 3D movements. Instead of using raw camera feeds or physical gloves, SL2T tracks body landmark coordinates locally with MediaPipe Holistic and translates those spatial points into text. It’s powering sign to text dictation on Pixel 11 in Gboard and Live Transcribe, starting with ASL. Users can sign naturally to search, draft messages, or respond in live conversations instead of typing everything out. This is seriously awesome.show more

AshutoshShrivastava
27,080 görüntüleme • 22 gün önce
SL2T is our breakthrough sign language-to-text model powering new... features for Deaf and hard of hearing users on Android. Starting with American Sign Language-to-English on Pixel 11, people can sign directly into Gboard and Live Transcribe instead of typing.show more

Google DeepMind
388,482 görüntüleme • 22 gün önce
Rive text is live!! Animatable, variable, dynamic text for... runtime, with full support for any language. There are a MONSTER amount of new things to play around with! 🚀 Create Text objects, styles, runs, and modifiers. 🚀 Support for any language. LTR, RTL, special characters/features, ligatures, etc. 🚀 Animate variable fonts. 🚀 Google Fonts are provided by default, but you can upload custom fonts. 🚀 Runs are a way to segment a text field into different sections. You can target runs at runtime to change their content dynamically. See documentation here: 🚀 Apply different text styles to runs. 🚀 Convert text to paths (flatten). 🚀 Text modifiers are a powerful way to add layered animation to text, while maintaining its editability! 🚀 Use transform constraints to resize graphics to text objects. Coming soon: 🚧 Ability to export fonts out of band or via CDN. They are currently packaged in the .riv file.show more

Guido Rosso
25,950 görüntüleme • 3 yıl önce
PhD Students - How to detect AI text in... your writing? We often use ChatGPT for writing. However, this leads to AI-plagiarized text. This can be problematic in many scenarios. For example, if you use AI text in your papers. Your research paper can get desk rejected. 🍁How to detect if there is AI text in your writing? 1. Go to and log in. 2. Click on 𝐴𝐼 𝑑𝑒𝑡𝑒𝑐𝑡𝑜𝑟 from the left menu 3. Insert your text and click on 𝐴𝑛𝑎𝑙𝑦𝑧𝑒. 4. will generate AI detection report This report shows the following. → Percentage of AI generated text → Options for converting AI text into non-AI text 🍁How good is this AI detector? SciSpace conducted a benchmarking study. In this study, the detection capability was compared with other AI-detectors. SciSpace AI detector was tested with 4000 samples. It showed an accuracy of 96%. This means it can detect AI-generated text with 96% accuracy. The study showed that SciSpace AI detector has outclassed AI detectors like GPTZero, ZeroGPT, and Grammarly. 🔴Anything you'd like to add?show more

Faheem Ullah
13,102 görüntüleme • 10 ay önce
MiniMax H3 Instead of sharing the prompts for each... of these videos, I thought it would be more useful to share how I created that prompts. All of the videos were generated with text-to-video. First, find an image with the kind of scene, composition and mood you want to recreate. I used a few YouTube playlist thumbnails as references but Pinterest is also a great place to find inspiration. You can even use your own old or nostalgic photographs. Then upload the image to ChatGPT and ask it to describe the scene. The description it gives you can essentially become your text-to-video prompt. From there, you can generate completely new scenes with a similar composition, atmosphere and cinematic language. You can of course use the reference image directly with image-to-video or as a first frame. But if the original image isn't yours, I prefer using it only as visual inspiration and recreating the scene through text-to-video. This is the prompt I use with ChatGPT: "Describe the scene in this image in English, focusing primarily on what is happening, the characters, their actions and body language, the setting and the overall atmosphere. Also briefly describe the composition, framing, camera angle, approximate lens choice, lighting, color palette and cinematic aesthetic. Keep it concise and scene-focused rather than overly technical."show more

Kōda
53,987 görüntüleme • 19 gün önce
Create a 3D model from a single image, set... of images or a text prompt in < 1 minute 😮💨 This new AI paper called CAT3D shows us that it’ll keep getting easier to produce 3D models from 2D images — whether it’s a sparser real world 3D scan (a few photos instead of hundreds) or your favorite 2D image generator like Midjourney (just an image). How does this magic work? “This architecture is similar to video diffusion models, but with camera pose embeddings for each image instead of time embeddings. The generated views are passed into a robust 3D reconstruction pipeline to create the 3D representation (Zip-NeRF or 3DGS)”show more

Bilawal Sidhu
92,867 görüntüleme • 2 yıl önce
LayerAI Launches XRP AI Bot: A Simpler Way to... Transact on the XRP Ledger 🌐 Users can now perform essential actions such as checking balances, sending tokens, or executing swaps through natural language commands, without relying on traditional interfaces or dApps. It works like ChatGPT, but instead of generating text, it executes XRP transactions on your behalf. Example commands: “Convert 150 XRP to EUR” “Send 75 XRP to address [wallet address]” “Swap 50 XRP for USDT” To get started: 1. Connect your XAMAN wallet (one-time setup) 2. Start a conversation with the XRP AI Bot 3. Review, confirm, and sign the transaction The XRP AI Bot is now available to all users. Start here:show more

LayerAI | AI2Earn
123,814 görüntüleme • 1 yıl önce
We’re excited to introduce Text-to-LoRA: a Hypernetwork that generates... task-specific LLM adapters (LoRAs) based on a text description of the task. Catch our presentation at #ICML2025! Paper: Code: Biological systems are capable of rapid adaptation, given limited sensory cues. For example, our human visual system can quickly adapt and tune its light sensitivity to our surroundings. While modern LLMs exhibit a wide variety of capabilities and knowledge, they remain rigid when adding task-specific capabilities. Traditionally, customizing these models requires gathering large datasets and performing often expensive, time-consuming fine-tuning for specific applications. To bypass these limitations, Text-to-LoRA (T2L) meta-learns a “hypernetwork” that takes in a text description of a desired task, as a prompt, and generates a task-specific LoRA that performs well on the task. In our experiments, we show that T2L can encode hundreds of existing LoRA adapters. While the compression is lossy, T2L maintains the performance of task-specifically tuned LoRA adapters. We also show that T2L can even generalize to unseen tasks given a natural language description of the tasks. Importantly, Text-to-LoRA is parameter-efficient. It generates LoRAs in a single, inexpensive step, based solely on a simple text description of the task. This approach is a step towards dramatically lowering the technical and computational barriers, allowing non-technical users to specialize foundation models using plain language, rather than needing deep technical expertise or large compute resources.show more

Sakana AI
403,159 görüntüleme • 1 yıl önce
STEVE-1: A Generative Model for Text-to-Behavior in Minecraft paper... page: Constructing AI models that respond to text instructions is challenging, especially for sequential decision-making tasks. This work introduces an instruction-tuned Video Pretraining (VPT) model for Minecraft called STEVE-1, demonstrating that the unCLIP approach, utilized in DALL-E 2, is also effective for creating instruction-following sequential decision-making agents. STEVE-1 is trained in two steps: adapting the pretrained VPT model to follow commands in MineCLIP's latent space, then training a prior to predict latent codes from text. This allows us to finetune VPT through self-supervised behavioral cloning and hindsight relabeling, bypassing the need for costly human text annotations. By leveraging pretrained models like VPT and MineCLIP and employing best practices from text-conditioned image generation, STEVE-1 costs just $60 to train and can follow a wide range of short-horizon open-ended text and visual instructions in Minecraft. STEVE-1 sets a new bar for open-ended instruction following in Minecraft with low-level controls (mouse and keyboard) and raw pixel inputs, far outperforming previous baselines. We provide experimental evidence highlighting key factors for downstream performance, including pretraining, classifier-free guidance, and data scaling. All resources, including our model weights, training scripts, and evaluation tools are made available for further research.show more

AK
144,806 görüntüleme • 3 yıl önce
Fine-tune DeepSeek-OCR on your own language! (100% local) DeepSeek-OCR... is a 3B-parameter vision model that achieves 97% precision while using 10× fewer vision tokens than text-based LLMs. It handles tables, papers, and handwriting without killing your GPU or budget. Why it matters: Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow. DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents. The best part? You can easily fine-tune it for your specific use case on a single GPU. I used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate. ↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you're working with. I've shared the complete guide in the next tweet - all the code, notebooks, and environment setup ready to run with a single click. Everything is 100% open-source!show more

Akshay 🚀
126,213 görüntüleme • 9 ay önce
Google DeepMind is absolutely on fire 🔥 they have... just launched Gemini Robotics-ER 1.5 their first broadly available robotics AI model designed to act as the "high-level reasoning brain" for robots. This is Google's first Gemini Robotics model made available to all developers. - Available in preview through Google AI Studio and Gemini API. - First thinking model for robots interacting with the physical world. Handles complex commands and orchestrates sophisticated robotic behaviors. - Key capabilities: Advanced spatial reasoning, multi-step task planning, Google Search integration, precise 2D pointing, object reasoning, and video analysis. - Flexible thinking budget lets developers tune speed vs accuracy. Can think longer for complex tasks or respond quickly for reactive operations. - Enhanced safety filters refuse dangerous tasks and recognize physical limits.show more

AshutoshShrivastava
137,226 görüntüleme • 11 ay önce
We are getting absurdly close to the point where... “learning Blender” means learning how to direct an AI. A shot like this looks soft and playful on the surface, but under it is the usual 3D pain: modeling, layout, materials, lighting, atmosphere, animation, and endless tiny fixes until the frame stops looking dead. That is why Kimi K3 matters. With Blender MCP, you can describe a scene like a robotic goat walking through a dreamy field and let the model help build the environment, place the camera, shape the materials, script the motion, and iterate inside the real Blender project. The real shift is not text-to-image. It is text-to-workflow. Kimi K3 does not just give you a pretty output and disappear. It can help move the actual scene from rough setup to something that looks art-directed. Soon the hardest part of 3D will not be the software. It will be whether your imagination is good enough to deserve tools like this.show more

Rina
45,651 görüntüleme • 1 ay önce
The Hidden Language of Diffusion Models paper page: tackle... the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a pseudo-token that is a sparse weighted combination of tokens from the model's vocabulary, with the objective of reconstructing the images generated for the given concept. Applied over the state-of-the-art Stable Diffusion model, this decomposition reveals non-trivial and surprising structures in the representations of concepts. For example, we find that some concepts such as "a president" or "a composer" are dominated by specific instances (e.g., "Obama", "Biden") and their interpolations. Other concepts, such as "happiness" combine associated terms that can be concrete ("family", "laughter") or abstract ("friendship", "emotion"). In addition to peering into the inner workings of Stable Diffusion, our method also enables applications such as single-image decomposition to tokens, bias detection and mitigation, and semantic image manipulationshow more

AK
41,830 görüntüleme • 3 yıl önce
This week's ChatGPT feature drop - Aug 7: 1/... Rich formatting in our web composer – When you paste in emails or documents, our composer will retain the formatting; copying and pasting is so common, we should have done this a while ago! 2/ Updated model for paid users – GPT 5.6 Sol is more consistent across quick chats and deeper reasoning. You'll now get a slider that lets you choose how much thought ChatGPT puts into a response. The haptics (vibrations) on mobile slider are fun! 3/ Unlimited text messages - Free users will get GPT 5.6 Luna with unlimited text messages. More intelligence for all. Rolling out soon. 4/ Fast Android Camera – We've made it a lot faster on Android to tap "camera" in ChatGPT to take a new photo and ask a question. 5/ Voice x Files - You can now upload files and ask questions in our new ChatGPT voice experience powered by GPT-Live. Team demos were awesome this week. So much in the queue that the next few months are going to be good. Let us know what you're hoping for in the comments!show more

Adam Fry
360,478 görüntüleme • 27 gün önce
Online right-wingers can’t take it. Most people are going... through their lives generally ok, laughing and joking with those around them. To be an online right-winger means being consumed by hate, sitting around rubbing one’s hands together, dreaming of collapse or god sending the heathens to hell or whatever. They need “meaning” in the form of telling others what to do because they don’t believe in mental health so don’t understand that this is just a sign that they’ve got problems.show more

Richard Hanania
5,147,883 görüntüleme • 2 yıl önce
Hermes HUD mode just became my live stock analyst... I opened a chart and asked Hermes to analyze what it was seeing Instead of only explaining it in text, HUD drew directly over my screen: • Resistance • Support • Trend direction • Current price context So now Hermes can look at the same chart I’m looking at and annotate the analysis in real time. This is exactly what I wanted HUD mode to become. Not another chat window. An AI layer on top of whatever I’m already doing 👀show more

Luke The Dev
87,020 görüntüleme • 7 gün önce
Wait for it! Solid shift eastward - out to... sea - on many of our models and even the new NHC forecast. These are the Google AI ensembles. What’s happening here - partly - is #Humberto is now forecast to be a Cat 5 monster #Hurricane and is exerting an even bigger influence on #Imelda, helping to pull it away like a magnet. Also the upper low in SE is now forecast to be a bit weaker exerting less influence. This is not set in stone, and it may shift back, but it’s a good sign. It’s worth noting the Google AI has been pretty consistent for days with most members showing a stall, then turn east out to sea. Fingers crossed!show more

Jeff Berardelli
85,444 görüntüleme • 11 ay önce
I’ve been using the early version of Naval's new... app Airchat for a year. Last week it went viral. And with it the haters came out in full force. Here’s why I’m still bullish on Airchat and believe it will change social media forever: 💬 Talk with real people: Twitter is the closest app to Airchat right now. But it’s completely impersonal. I’ve already built stronger connections with people on Airchat than I have from years of using Twitter. 🎤 Transcription is awesome. Voice notes are long. You don’t know if they're relevant. Having a text version of the voice note to scan makes it easier to find relevant content. 🗣️ Transcription supports multiple languages. I can record in Hebrew, and you’ll read it in English. It works great! Try it! 👥 Groups: launched this week. This focuses discussions around specific topics. Can’t wait to see more of this in the weeks ahead. 🏆 Talk to people you couldn't in other settings. Part of this is because it's so early. But you can speak to people like Naval that you wouldn't have been able to speak to otherwise. And a response to the haters: “It’s just Clubhouse” It’s not. Saying it’s Clubhouse is like saying Zoom is Clubhouse. Or a Podcast is Clubhouse. Clubhouse is sync. Airchat is async. Clubhouse is 1:many. Airchat is 1:1. “It’ll die down. Every new social app dies” Maybe. Creating a new social app is hard. It might die. But what’s that matter to you. The relationships you build last whether or not the app becomes a unicorn. “Twitter will just copy it like they copied Clubhouse” They might eventually. But what’s that matter unless you’re an investor in Airchat? I predict apps will copy it. Discord channels will offer this feature, and it will be awesome. Turning text chats into real conversations. WhatsApp is long overdue transcription for voice notes. Ideas from Airchat will get added to many apps. “I can’t use Airchat in places I can’t talk” The intimacy of Airchat is a feature, not a bug. Certain settings are meant for talking, others aren’t. Airchat is about conversations with people. You can’t do Zoom meetings everywhere either. “Twitter conversations are also conversations” Yes, but they’re far less intimate. You’re talking to a wall of text, not a person. You don’t date over text, you date over voice. It’s more intimate. You’re talking to a real person. “The content is boring” It will improve over time. There will be more content. The algorithm will improve to show you the best content. Groups will allow you to find specific content you're interested in. “I prefer Twitter” Good. Use Twitter. I use it too. “The hype will die down” Probably. But what’s that matter to you? You’ll only use apps that have 100m DAU? Successful social apps have early hype that dies down, and then grow steadily after that. 🪭 As you can tell, I’m a fan 🪭 If you’re on the app feel free to say hi. My username is `eliezer`. And feel free to DM if you need an invite.show more

Elie Steinbock — oss/acc
62,767 görüntüleme • 2 yıl önce
🚨 A FAKE “HOLLYWOOD SIGN” JUST APPEARED OVER A... BUSY L.A. FREEWAY — AND DRIVERS CAN’T BELIEVE WHAT IT’S ADVERTISING Drivers in Los Angeles are doing double takes at highway speeds after a massive sign showed up above the freeway… designed to look almost identical to the real Hollywood sign. And this thing isn’t small. • Roughly 230 feet wide • About 30 feet tall • Giant white letters stretched across the hillside • Positioned right above heavy traffic But here’s the part that’s setting people off. It’s not promoting a movie. It’s not promoting Hollywood. It’s promoting Fiverr’s new AI video tools… where companies can generate ads, content, and campaigns without hiring real creatives. So now you’ve got: • A fake Hollywood sign • Sitting over one of the busiest roads in the city • Advertising AI replacing the very industry Hollywood was built on It’s just similar enough to the real thing to make drivers look twice… exactly where they shouldn’t be. So is this just a clever ad… or are they quietly telling you what happens to Hollywood next?show more

HustleBitch
18,121 görüntüleme • 5 ay önce