Loading video...

Video Failed to Load

Go Home

Are we done with object detection? What about tiny objects beyond 200 meters? 🔎 Telescope 🔭 addresses long-range perception by explicitly tackling extreme scale imbalance ⚖️ in images. It hinges on a learnable hyperbolic foveation transform from a low-resolution image, magnifying distant regions 🔍 while compressing nearby ones -...

188,500 views • 2 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Wonderland: Navigating 3D Scenes from a Single Image Contributions: • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.

MrNeRF

52,801 views • 1 year ago

🚨 MASSIVE ASTEROID ALERT 🚨 😱 5 ASTEROIDS TO STRIKE EARTH ON JANUARY 4! ⚠️On January 4, something unusual appeared on NASA’s tracking systems. Not one. Not two. But five separate space objects entered Earth’s monitored region of space — all on the same day. 🛰️ WHAT WAS DETECTED NASA’s Near-Earth Object monitoring network identified five asteroids moving at extreme speeds, each following its own calculated path around our planet. • Classified as Near-Earth Asteroids (NEAs) • Traveling at tens of thousands of km/h • Detected days in advance • Tracked continuously as they approached and passed. Their distances varied, but all were close enough to activate automatic observation protocols. 🔭 WHY THIS STOOD OUT Asteroids pass Earth often — but multiple objects on the same date always draw attention. Every trajectory was recalculated. Every data point rechecked. Ground-based telescopes and automated systems stayed locked in. No public countdown. No dramatic warning. Just silent monitoring. 🌑 WHAT THE SYSTEMS SAW • Stable orbits • No sudden course changes • No fragmentation • No interaction with Earth’s atmosphere. One by one, the objects passed Earth’s vicinity and moved back into deep space. 🌍 THE RESULT No impact. No damage. No visible sign in the sky. To most people, January 4 felt completely normal. But above our heads, space traffic moved quietly — and was watched closely. 🛰️ THE BIGGER PICTURE Earth travels through a solar system filled with ancient debris left over from planet formation. Most objects remain distant. Some pass close. A few demand attention. This was one of those moments. The universe didn’t slow down. Earth didn’t notice. NASA kept watching. 🌌 Five asteroids came and went. The planet remained untouched. And space moved on. #Asteroids #NASA #SpaceUpdate #NearEarthObjects #CosmicWatch #Astronomy #Universe

3I/ATLAS updates

23,554 views • 6 months ago

Everyone's sleeping on image-to-3D AI models. They can make your app look incredibly unique, with just a little effort. Here's how. This is my calorie tracker, built in a week with nothing but prompting. Just Claude Code + a couple APIs. The visuals are all AI-generated. I'll be sharing the full workflow + all the crazy technical stuff Claude and I did to make this work, so nobody has to struggle through it like me. Deep dive coming soon! Till then, this is the high-level idea: 1. Get a clean image of the food (or whatever your asset is) - In my app, the user describes foods via text, or attaches images (or both) - If text, an LLM extracts the food description and formats it into a specific prompt I tuned for this design, and we generate an image using Z-Image Turbo through fal - If image, we do the same thing but with FLUX.2 [dev] to edit the user image into our reference design - Originally, both used Google Nano Banana, but switching to open models cut costs and latency a ton 2. Gaussian splatting (2D image → 3D model) - I tried various 2D-to-3D options on fal and ended up with TripoSplat as my preferred balance of speed, cost, latency; this turns an image into a 3D model that looks super high quality (link below) - The app displays the 2D image while our backend generates the 3D splat - We "groom" the splat to reduce size and load time by culling low-opacity/scale points 3. Render efficiently on device Originally, it looked great but ran at 10 FPS. Getting to 120 FPS was a crazy journey. TL;DR: - SwiftUI had to go; it forced us to render each asset in independent MTKViews, which wasn't workable - Instead, we composite every dish into one full-bleed CAMetalLayer using MetalSplatter (link below) - We had to make some optimizations within MetalSplatter's code too, to reduce the overhead of sorting points per render Then I added some finishing touches like the subtle rotation and parallax as they move around. I think it turned out pretty cool :) Overall, this took some effort, but we still got it done in less than a day. Hopefully your agent can follow in the footsteps of mine and do it much faster. Keep an eye out for the bigger writeup, which'll give your agent everything it needs. If you have any questions, drop em below!

Anshu

19,931 views • 29 days ago

Nanobanana Pro Higgsfield AI 🧩 なるほどこれは便利!! いろんなショットを一発で出して、そこから選ぶ感じ 見事にアップスケールしてくれる プロンプトはリプ欄の元投稿を使わせていただきました ””” Analyze the entire composition of the input image. Identify ALL key subjects present (whether it's a single person, a group/couple, a vehicle, or a specific object) and their spatial relationship/interaction. Generate a cohesive 3x3 grid "Cinematic Contact Sheet" featuring 9 distinct camera shots of exactly these subjects in the same environment. You must adapt the standard cinematic shot types to fit the content (e.g., if a group, keep the group together; if an object, frame the whole object): **Row 1 (Establishing Context):** 1. **Extreme Long Shot (ELS):** The subject(s) are seen small within the vast environment. 2. **Long Shot (LS):** The complete subject(s) or group is visible from top to bottom (head to toe / wheels to roof). 3. **Medium Long Shot (American/3-4):** Framed from knees up (for people) or a 3/4 view (for objects). **Row 2 (The Core Coverage):** 4. **Medium Shot (MS):** Framed from the waist up (or the central core of the object). Focus on interaction/action. 5. **Medium Close-Up (MCU):** Framed from chest up. Intimate framing of the main subject(s). 6. **Close-Up (CU):** Tight framing on the face(s) or the "front" of the object. **Row 3 (Details & Angles):** 7. **Extreme Close-Up (ECU):** Macro detail focusing intensely on a key feature (eyes, hands, logo, texture). 8. **Low Angle Shot (Worm's Eye):** Looking up at the subject(s) from the ground (imposing/heroic). 9. **High Angle Shot (Bird's Eye):** Looking down on the subject(s) from above. Ensure strict consistency: The same people/objects, same clothes, and same lighting across all 9 panels. The depth of field should shift realistically (bokeh in close-ups). A professional 3x3 cinematic storyboard grid containing 9 panels. The grid showcases the specific subjects/scene from the input image in a comprehensive range of focal lengths. **Top Row:** Wide environmental shot, Full view, 3/4 cut. **Middle Row:** Waist-up view, Chest-up view, Face/Front close-up. **Bottom Row:** Macro detail, Low Angle, High Angle. All frames feature photorealistic textures, consistent cinematic color grading, and correct framing for the specific number of subjects or objects analyzed. """

yachimat - AI Short Anime

115,284 views • 7 months ago

🚨 CHINESE SCIENTISTS JUST INVENTED 3D PRINTING THAT CREATES OBJECTS IN 0.6 SECONDS USING ONLY LIGHT. Researchers at Tsinghua University have developed a new method called DISH (Digital Incoherent Synthesis of Holographic light fields) that can print complex millimeter-scale objects almost instantly. Instead of slowly building layer by layer, the system fires thousands of precisely patterned light images from multiple angles into a still vat of liquid resin. Where the light overlaps, the resin instantly hardens into a solid 3D object. The entire process takes just 0.6 seconds. Why this matters: • It’s currently the fastest volumetric 3D printing method ever demonstrated • Achieves extremely fine detail features thinner than a human hair • The resin stays completely still, so there’s no vibration or distortion • It can work with watery (low-viscosity) resins, making it suitable for biological applications • The team has already printed complex structures like blood vessel-like tubes and even a tiny bust of a historical figure The deeper implication: Traditional 3D printing has always been limited by speed and the need to move either the print head or the resin. This approach removes both constraints by using light itself as the sculptor. Because it can print directly into still liquid (and potentially onto living tissue), it opens new possibilities in bioprinting, medical devices, and rapid manufacturing. If the technology can be scaled beyond millimeter sizes, it could fundamentally change how we think about making physical objects turning “print” from a slow process into something closer to instantaneous fabrication. We’re moving from “layer by layer” to “all at once.” How do you think instant volumetric 3D printing like this could change medicine, manufacturing, or everyday life if it becomes widely available? Follow for more frontier manufacturing and materials science breakthroughs.

TheNewPhysics

347,458 views • 1 month ago

Dear Tarun Chitra 1. We are the original creators of DeSci back in 2016. What DeSci has become today is largely unrelated with its original model of producing rigorous peer-reviewed scientific studies published in reputable medical journals. The model we introduced. We are tirelessly fighting against pseudoscience, and we are showing the world that people can understand the difference between legit science and pseudoscience with the success of $INNBCV. Yes, meritocracy is possible in crypto. Even against all odds. 2. We are the only project in the entire crypto space that ever funded, performed, and published highly innovative HIV cure research ( We are the project that produced the first peer-reviewed study on blockchain-based biomedical data storage in the world’s most reputable scientific network, Springer Nature ( $INNBCV is not for the privileged few; it is for the many. We resisted all the pressure from those who wanted us to provide big allocations to VIPs of other DAOs “because it is good for the marketing” and put our users first, ensuring a fair launch, a launch for the people, and they turned $70k into $2,000,000. $INNBCV shows that you can have a sustainable model, provided you are backed by actual science. And thanks to the amazing guys at daos.fun baoskee and Solana community. Behind our project there is the sweat and blood of years of work to produce publications in the most reputable medical journals. Just to put things into perspective, it took us 3 years to publish our latest work in Springer Nature. 3. Unlike many other projects, we had no ICO/VCs, meaning we had to prove ourselves every single day because we are only supported by our community. If we deliver products, we survive; it is either publish or perish for us, and that’s why we have such a close connection to our community. $INNBCV is a struggler, $INNBCV is a survivor, $INNBCV is not for the privilege of the few but for the people. Our community makes it possible by supporting us. You guys are the real heroes.

InnovativeBioresearch🇮🇹

10,867 views • 1 year ago

Tried this viral prompt idea on BudgetPixel AI using GPT Image 2 + Seedance 2.0 Prompt remove the arrows immediately while starting the video. The camera generates footage in a first-person, ultra-high-speed perspective, faithfully following the exact path of the red line marked on the reference image. Cinematic presentation. Low-angle ground-level shot racing across the lush green meadow filled with wildflowers and tall grass, following the exact curving path shown in the reference image. Pass smoothly right beside the large fluffy long-haired cat sitting alert on the left foreground, its fur detailed and catching warm sunlight. A second cat lies playfully on its back in the grass nearby. Scattered throughout the vibrant field are numerous hens and chickens grazing and moving naturally in the same positions as the reference photo. A rustic wooden farmhouse sits nestled in the midground among the greenery. The camera then continues the smooth ascent following the curving path upward along the gentle stream area, soaring through the dense trees toward the towering rocky mountain peaks. It dramatically circles the central mountain before pulling back into a breathtaking bird’s-eye panoramic view of the entire valley, showcasing the full scale of the river, forests, fields, farmhouse, and animals below. One continuous fluid cinematic shot with no cuts. Photorealistic, ultra-detailed textures on the cats’ fur, chicken feathers, vegetation, water reflections, and rocks. Warm golden hour lighting with soft volumetric god rays, rich atmospheric depth, vibrant natural colors, National Geographic level realism, masterpiece --ar 16:9 --stylize 25 --v 6

Aaliya

14,611 views • 1 month ago

🚨 ANOTHER BRAND NEW VIDEO DIRECT FROM IRAN: The call to be armed has evolved. This is no longer about small-scale tactics or asking for pistols to target low-level regime forces; it is the deployment of a next-level strategy. A highly sophisticated, organized force inside Iran is preparing for a decisive action to finish the matter. These individuals are sending videos directly to me to break through the silence and project a completely new reality on the ground. This latest footage reveals a group operating east of Tehran near Pardis and Jajrud, confidently displaying their exact GPS location before setting up a fleet of FPV and surveillance drones. Their message is absolute: the era of disorganized vigilantes playing around with guns is over. They are operating with modern tools and calculated precision. On the audio track, a mechanically disguised voice delivers a clear warning directly to the state: "This time you are in the center of our target. We are not waiting for the help of others, and we are waiting for you where you do not expect it." The message ends with the declaration "Long live Iran, Javid Shah" (Long Live the King). While they are maintaining strict operational security by keeping their exact targets and timelines secret, they promise the results will be seen soon. For the millions they represent, peaceful protest against live ammunition has proven impossible. They aren't asking for foreign boots on the ground. They are demonstrating that if Crown Prince Reza Pahlavi has a force at his command, that force is intelligent, advanced, and ready. And let’s be clear: the only viable path to effectively arming and coordinating this movement is through direct alignment with the Prince and his team. They risked their lives to deliver this proof straight from the ground. More people need to amplify their call. Do not let the regime bury their bravery.

Armin Navabi

37,271 views • 16 days ago

‼️DECLASSIFIED: New 2024 Infrared UAP Footage Released — Military Still Can’t Identify It 👀 As part of the Presidential Unsealing and Reporting System (PURS), the Department of War just cleared for release a stunning new video. This isn't a leaked cellphone clip—it is one minute and 39 seconds of high-definition infrared footage submitted by the U.S. Indo-Pacific Command. The 2024 Incident: This footage was captured by a U.S. military platform in 2024, proving that these incursions are not just "historical" relics but an ongoing reality of our current airspace. The "Unresolved" Status: Unlike previous clips that the Pentagon tried to dismiss as "parallax" or "tricks of the eye," this Indo-Pacific footage remains an unresolved case. The government is officially admitting they cannot determine the nature of the observed phenomena. Secretary of War Pete Hegseth stated that this release is part of a mandate to find, review, and declassify these files expeditiously. The days of "justified speculation" are being replaced by raw data. This footage follows a series of other chilling reports from 2024, including government contractors witnessing a "large metallic cylinder" the size of a commercial airplane that vanished in mid-air. We are seeing objects with morphological features and behaviors that defy the state of the art. The Reality: The "War for the Mind" includes keeping us in the dark about what is actually in our skies. But with the unsealing of these files, the "scales" are finally being removed. 🏛️📉 If these objects can't be identified by the most advanced military sensors on the planet, then who is flying them? The Department of War has even invited private-sector analysis to help resolve these cases. That means the "Noticing" community is now officially part of the search for the truth. RT to spread the declassified proof. The era of secrecy is over. 👇

Project Constitution

64,986 views • 2 months ago

A moment suspended between Saudi Arabia's football passion and coffee tradition. GPT Image 2 + Seedance 2.0 on BudgetPixel AI prompt A highly cinematic, photorealistic single-shot sequence that preserves the exact original location, environment, architecture, objects, lighting conditions, camera perspective, and subject position from the source video. Do not replace, redesign, relocate, or alter the setting in any way. The person remains in the exact spot where they were filmed, maintaining their original pose, facial expression, body position, and interaction with the environment. The subject is wearing the official Saudi Arabia national football team uniform throughout the entire sequence: authentic green Saudi Arabia jersey with white details, official team crest, matching football shorts, athletic socks, and football boots. The uniform must appear naturally integrated into the original scene with realistic fabric folds, stitching, texture, shadows, reflections, and movement-free realism. Every background element, object, texture, structure, shadow, reflection, and environmental detail must remain identical to the original footage. The effect transforms the captured moment into a frozen-time cinematic sequence while keeping the real-world location completely unchanged. The subject is captured at the exact moment they pour a beverage from a transparent cup. Time has completely stopped. The liquid erupts from the cup in a dramatic suspended splash, forming elongated ribbons, twisting streams, intricate arcs, and hundreds of individual droplets frozen midair. Every droplet, splash fragment, and liquid strand appears perfectly suspended in space, creating the impression of a sculptural masterpiece made of liquid. The liquid spilling from the cup must be identical to the liquid inside the cup, with perfectly matching color, texture, thickness, reflections, transparency, and material properties. The beverage can be any type or color, but it must remain visually consistent throughout the scene with extreme realism. The subject remains absolutely motionless, frozen in the precise instant of action. Their posture, facial expression, fingertips, hair strands, jersey fabric folds, shorts texture, socks, football boots, accessories, and every micro-detail are perfectly preserved. Tiny condensation droplets on the cup, reflections on the surface, and subtle imperfections remain locked in place as if the entire world has been paused between two frames of time. The surrounding environment is equally frozen. Every object visible in the original footage remains completely static. Nothing moves. No wind, no shifting light, no falling droplets, no environmental motion. The entire world exists in a state of perfect suspension. The only moving element is the camera. The camera performs a slow, smooth cinematic arc movement around the subject, beginning from the original camera viewpoint and gradually orbiting to one side while maintaining focus on the frozen action. As the camera travels through three-dimensional space, it reveals changing perspectives of the suspended liquid sculpture, the Saudi Arabia football uniform, the subject, and the original environment. Strong spatial parallax is visible throughout the movement. Foreground droplets, liquid strands, the subject, nearby objects, and distant background elements shift relative to one another, creating a powerful sense of depth and dimensionality. The scene feels like moving through a perfectly preserved moment in time. Natural lighting remains consistent and unchanged throughout the shot. Shadows stay fixed, reflections remain stable, and materials such as glass, metal, stone, wood, fabric, football jersey fabric, embroidered team crest, and liquid exhibit highly detailed photorealistic textures. Captured with a premium wide-angle cinema lens, the scene emphasizes depth, scale, and immersive three-dimensional realism. The Saudi Arabia football uniform appears crisp, premium, and authentically detailed, with realistic fabric texture and professional sportswear quality. Core visual concept: The entire world is frozen in time exactly as it appeared in the original footage, like a hyper-detailed sculpture, while a subject wearing the Saudi Arabia national football team uniform pours a beverage that explodes into a suspended liquid masterpiece. The camera freely moves through the frozen moment, revealing dramatic parallax, depth, and cinematic realism from multiple angles. Style: Hyper-realistic, cinematic, ultra-detailed, 3D stop-motion illusion, frozen-time photography, volumetric depth, realistic lighting, film-quality rendering, smooth camera orbit, strong parallax, premium commercial sports production, museum-like suspended motion sculpture, exact environment preservation, original location consistency, photorealistic liquid simulation, luxury football advertisement aesthetic, FIFA World Cup promotional quality, 8K photorealism.

Sharon Riley

42,983 views • 1 month ago

Yesterday at Brown University ICERM's workshop on “Agentic Scientific Computing and Scientific Machine Learning” I spoke about “Adaptive Swarms Across Scales”, making the case for scientific AI as systems that can create representations, stress them, fracture them, and enlarge the category in which future representations live. The category here is a composable and breakable working universe of science: data, hypotheses, simulations, measurements, tools, failures, figures, papers, provenance, and the transformations that connect them. Discovery happens when those transformations become executable, inspectable, composable, and capable of changing the world model they operate within. Atomistic modeling gives one category - states, forces, trajectories, observables, boundary conditions, conservation laws. Neural surrogates learn fast morphisms inside or between such categories. But discovery is higher-order: it changes which objects and morphisms are available in the first place: what variables exist, what operations are allowed, what evidence counts, what scale is active, what invariant is being preserved, and what kind of explanation the system is even capable of forming. This is scientific method as adaptive architecture: compression, stress, fracture, recomposition. Fracture matters here because it makes the logic physical: a non-commuting diagram realized in matter. The imposed load, material hierarchy, defect field, and assumed continuum description no longer map cleanly into the observed outcome. The crack is the obstruction and it identifies where the old morphism failed and where a new representation must be introduced. The physical crack and the categorical obstruction are the same event viewed in different substrates. ScienceClaw × Infinite is a machine for constructing and transforming a category of scientific artifacts. Each artifact is typed. Each operation has lineage. Each failed branch remains in the category as reusable structure. The “paper” is no longer the terminal object of science; it is one projection of a larger compositional trace, and it can be generated at any time for consumption by a human or an AI. With that the unit of scientific labor is changing. For most of the twentieth century the unit was the result (a measurement, a theorem, a synthesized molecule). It is now becoming the algorithm that produces results, and after that, the substrate of discovery itself. The static PDF is the wrong terminal object for this regime, and the role of the scientist with it. We now design algorithms that build algorithms, and eventually substrates in which such algorithms compose themselves. At that point, the scientist is no longer outside the discovery system. The scientist becomes one of the representations the system can transform. In that sense, the systems will eventually do science to us, and that is the structural consequence of the principle they are built on.

Markus J. Buehler

10,095 views • 2 months ago