Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

THIS $599 MAC MINI CAN DRIVE THREE DISPLAYS NATIVELY - THIS SETUP TOOK "MORE SCREEN SPACE" SO FAR IT ENDED WITH BINOCULARS. the M4 Mac mini shrinks a full desktop into a 5 x 5-inch box. the base model launched with: 10-core CPU 10-core GPU 16GB unified memory support...

16,290 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

A wrist force sensor fires at 100Hz. The policy only ever sees it at 30Hz, downsampled to land on the same control step as the camera and the joint state. That's not a bug, it's the whole point, and it sits inside a bigger pattern in VLA research this year. Every major release has been Markovian at its core, mapping the current frame straight to the next action. The fix everyone reaches for is more vision: more history frames, longer image context. FM-VLA makes a clean case that the fix is the wrong channel for a whole class of tasks. Press a button three times and stop. A camera watching that has almost nothing to work with, the scene barely changes between press one and press three. Force doesn't have that ambiguity problem. Each press is a sharp, distinct spike in the wrench signal, whether or not the camera noticed anything at all. So FM-VLA doesn't add more frames. It compresses the wrench history into eight tokens with a VAE, pretrained purely on reconstructing force signals, frozen before it ever touches the policy, then hands those tokens to the action expert alongside a short window of joint state. That's the entire memory system. Averaged across three contact-rich tasks, FM-VLA hits 83.3 percent success against 33.3 percent for the strongest vision-memory baseline on the button-counting task specifically, where the ambiguity problem is worst, 72.2 percent for FM-VLA there. Strip out the short-state window and force-only performance drops well below the combined system, so force alone isn't the answer either. The two channels are doing different jobs. The field has defaulted to one memory channel for every kind of temporal problem. This is a clean data point that the channel should match the ambiguity you're actually trying to resolve, not just get bigger. Source: Paper: Credit to the teams at Tsinghua University, Microsoft Research, and Fudan University. #Robotics #PhysicalAI #RobotLearning

Stephen James

11,658 Aufrufe • vor 2 Monaten

Seedance 2 is so good at character and scene consistency across cuts. I tested it with a short film on something I learned last week about cats. The key is detailed prompts - take your story to an LLM to nail the structure, then to Krea to make the video. Take my prompt 👇 Style: high-quality 3D animation, expressive character acting, cozy lighting, cinematic composition Timestamped structure: 0:00 - 0:02 Open on a cozy desk setup in a home office. A black cat, Finn, is near the desk, alert and curious. The woman with brown hair and blue eyes sits at her computer, pulling up a YouTube-style bird video. The screen shows colorful birds hopping on branches. Finn immediately locks in on the movement. 0:03 - 0:05 The woman smiles knowingly and adjusts the monitor so Finn can see better. She lightly gestures toward the screen. Finn jumps onto the chair or desk edge and sits facing the monitor, ears forward, body still, completely mesmerized by the birds. 0:05 - 0:07 The woman gets up from the desk and walks away casually, leaving Finn watching. Finn remains seated, staring at the bird video with total concentration. Maybe a subtle tail flick or tiny head movement tracks the birds on screen. 0:07 - 0:10 Cut to later in the day, in the exact same office. The woman is now back at the desk in a Zoom meeting, still with the single monitor on the same desk. The computer screen now shows a video call grid or work presentation instead of the birds. She is mid-conversation, professional and focused, seated upright, saying "and now we'll move to the next slide" 0:10 - 0:12 Finn suddenly jumps up onto the desk from below, entering frame with urgency. He looks at the screen, confused and dissatisfied that the birds are gone. He starts pawing at the screen and the woman is shocked and embarrassed

Justine Moore

27,449 Aufrufe • vor 5 Monaten