Do 3D reconstruction transformers really need a billion parameters,... or are most of those layers just doing the same thing over and over? Introducing Déjà View: a single transformer block, looped K times, that matches or beats models 8–10× its size with lower compute. 🧵show more

Tobias Fischer
112,894 次观看 • 4 个月前
Wonderland: Navigating 3D Scenes from a Single Image Contributions:... • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.show more

MrNeRF
52,849 次观看 • 1 年前
Create a 3D model from a single image, set... of images or a text prompt in < 1 minute 😮💨 This new AI paper called CAT3D shows us that it’ll keep getting easier to produce 3D models from 2D images — whether it’s a sparser real world 3D scan (a few photos instead of hundreds) or your favorite 2D image generator like Midjourney (just an image). How does this magic work? “This architecture is similar to video diffusion models, but with camera pose embeddings for each image instead of time embeddings. The generated views are passed into a robust 3D reconstruction pipeline to create the 3D representation (Zip-NeRF or 3DGS)”show more

Bilawal Sidhu
92,919 次观看 • 2 年前
Show-o One Single Transformer to Unify Multimodal Understanding and... Generation discuss: We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model.show more

AK
124,085 次观看 • 2 年前
I love that g1 Soundwave is genuinely just terribly... clumsy. Like he cannot walk or run in a straight line for the life of him; and even as he became Soundblaster in headmasters, he still tripped over himself and flipped over a rock. What’s wrong with him #Transformersshow more

Moth
34,116 次观看 • 18 天前
🚨 Wait… I think I just discovered something HUGE... about AI-generated 3D models. I found Hyper3D’s new Agentic Mode that can BUILD, EDIT & ANIMATE 3D models!! And there’s one thing that really surprised me. Unlike regular 3D generation, Agentic Mode can automatically filter and organize more informative input. You can give it different views, individual parts, or even a single reference image of separate components… And it can figure out how they fit together and generate a complete 3D model — with surprisingly accurate structure and details. That’s something I haven’t seen regular 3D generation handle this well. Check out this 3D windmill animation 👇 Hyper3D Agentic Mode lets you control 3D models with PARAMETERS !! Instead of rebuilding the sofa from scratch, you can change things like: → Parametric Control → Animation Pressets → Enhance Materials So you can generate the model once... and keep editing its geometry without starting from scratchshow more

SARAH
90,925 次观看 • 7 天前
The US "diet industry" makes over $70 billion a... year... Yet Americans are fatter than ever. Why? They sell you useless fad diets and supplements that don't work. Here are the only 8 diet rules you need for fat loss (bookmark this): 🧵 1. Eat 150g+ of protein a dayshow more

Chace Chambers
156,837 次观看 • 9 个月前
and he’s right! love me a man of science... 🤓 🍒 : with how i look at it.. the core and the erector spinae muscles.. those that support your lower back.. i think working on these areas consistently is good. even if you don't go as far as doing pilates or things like that.. just the basics! because eventually, if you lose strength on those muscles over time, it puts strain on your spinal discs. so you really have to work on your lower back muscles!show more

NLNL 🍒🔥
21,488 次观看 • 1 个月前
Conducting road safety audits, scheduling maintenance, or planning logistics... routes used to require combing through imagery or, worse, driving miles of roads yourself. Not anymore. New AI-powered data layers in Google Earth are changing that. Pulling from billions of Google Street View images, these layers now spot infrastructure assets like stop signs, speed limit signs, and more to map infrastructure for you. By combining these assets into a single project and diving into Street View you can locate and validate assets on Google Earth in seconds. These layers are now available for Professional & Professional Advanced customers on web and Android, with coverage expanding over the coming weeks.show more

Google Earth
21,347 次观看 • 3 个月前
A VIEW OF FOUR STORMS A moving satellite image... released by the US National Oceanic & Atmospheric Administration (NOAA) reveals the view from above last Nov. 11 over 8 hours of 4 typhoons that hit or are set to hit the Philippines this week: Marce, Nika, #OfelPH, and #PepitoPH. (Courtesy: CSU/CIRA & NOAA/NESDIS) | via Anjo Bagaoisanshow more

ABS-CBN News
1,139,623 次观看 • 1 年前
This is... not a remotely accurate description of what... the 2023 Al executive order did? Undersecretary Emil Michael: "If you remember the Biden executive order on Al, which was this crazy executive order that limited the amount of compute any model company could do and was essentially grandfathering in a small number of ai companies that they were gonna designate as the winners, and everyone else was out" Its not true that the EO limited the compute that AI companies could do. What it did do was require companies who were training models above a certain very high compute threshold (10^26 FLOP or 10^23 FLOP for models trained primarily on biological sequence data) to notify the government and share what testing and red teaming they were doing for certain national security risks. People are free to dislike the Biden AI EO! But it seems good to factually describe what the policy said.show more

Nathan Calvin
58,745 次观看 • 7 个月前
There is no such thing as "lower abs". Your... rectus abdominis is one continuous muscle. You cannot contract the bottom half on its own, and no exercise on earth isolates a region that does not anatomically exist. So all the things sold to you as "lower ab" builders are actually just glorified hip flexor exercises: - Leg raises - Reverse crunches - Flutter kicks - Scissor kicks - Mountain climbers - The dreaded "lower ab burnout circuit" None of them is carving a separate lower section. They train the same single muscle every other ab exercise trains, just with more hip flexor and more wishful thinking. The reason you cannot see your lower abs is not a missing exercise. It is the layer of fat sitting over them, which collects there last and leaves there last. So the actual programme is dull and two-part: - Get into a calorie deficit and strip the fat that is hiding them - Build the abs themselves with weighted crunches, 4-6 reps, adding load over time, like any other muscle That's it. There is no lower-ab secret. There is just a visible-ab one, and it lives in the kitchen and under a loaded cable.show more

Sama Hoole
27,193 次观看 • 3 个月前
On Wednesday, Brandon Johnson was asked whether he is... running for re-election. He said: “Why are you mad at me for doing what the people of Chicago elected me to do? I’ve kept every single promise. I have not lied to the people of Chicago.” Did Mayor Johnson just lie when he said that? Do you feel he has lied to the people of Chicago? Do you feel he had a mandate for his progressive agenda when Chicago voter turnout was about 35%, and he was elected with roughly 18% of the vote—or just over 50% of those who voted?show more

Reporter William J. Kelly #thatreporter
76,738 次观看 • 7 个月前
Solana, BNB Chain And Arbitrum Clash Over Robinhood Chain... Fee Model Robinhood (Robinhood Crypto) Chain’s revenue model has sparked a wider debate about blockchain economics. Solana co-founder Anatoly Yakovenko previously criticized Robinhood’s revenue arrangement with Arbitrum. The network lets Robinhood retain most gas revenue while sharing 10% of its protocol net revenue under its Arbitrum agreement. Its applications generated $2.66 million in revenue over 24 hours on August 31. Offchain Labs CEO Steven Goldfeder defended the arrangement, saying Robinhood keeps most gas revenue. Nina Rong, Executive Director of Growth at BNB Chain said lower gas fees are no longer the industry’s top priority. She argued that blockchains need sustainable revenue models to keep funding technology and development.show more

BSCN
27,518 次观看 • 1 个月前
Transformer by hand ✍️ ~ 6 steps walkthrough below... Open the hood of a transformer and the parts list is overwhelming: embeddings, positional encoding, attention weighting, self-attention, cross-attention, multi-head attention, layer norm, skip connections, softmax, linear, Nx, shifted right, query, key, value, masking. Which of those actually make the car run? Two of them. Attention weighting and the feed-forward network. Everything else is an enhancement to make it run faster and longer, which is how we got from a car to a truck, and to the word "large" in large language model. So I drew and calculated those two parts entirely by hand. Goal: push five features through one transformer block, filling in every cell yourself. 1. Given Five positions of input features, arriving from the previous block. 2. Attention matrix Let us feed all five features to a query-key module (QK) and read back an attention weight matrix, A. The details of that module are a post of their own. 3. Attention weighting We multiply the input features by A to get the attention weighted features, Z. Still five positions. The effect is to combine features *across positions*, horizontally: X1 becomes X1 + X2, X2 becomes X2 + X3, and so on. 4. First layer Let us feed all five weighted features into the first layer of the FFN. Multiply by the weights and biases. This time the combining happens *across feature dimensions*, vertically, and each feature grows from 3 numbers to 4. Note that every position goes through the same weight matrix. That is what "position-wise" means. 5. ReLU We cross out the negatives. They become zeros. 6. Second layer Let us bring it back down: 4 dimensions to 3. The output feeds the next block, which has a completely separate set of parameters, and the whole thing runs again. You have just calculated a transformer block by hand. ✍️ The takeaway: the two parts are doing two different jobs, and neither one alone is enough. Attention mixes *across positions*, so a feature can see its neighbours. The FFN mixes *across feature dimensions*, so each position can think about itself. Horizontal, then vertical. Then that pattern repeats N times, each block with its own separate set of weights. That is the Nx from the list up top, and that is what makes the transformer run. 💾 Save this post! #AIbyHand #Transformers #DeepLearningshow more

Tom Yeh
26,211 次观看 • 2 个月前
Once she stepped out of that car towards him... with a gun. That is not self-defense. She wanted a fight, and she killed an unarmed man at Walmart. All she had to do was stay in her car, or leave. He should of done the same. But you can't just kill someone over a parking spot and call it self defense.show more

The SCIF
95,330 次观看 • 3 个月前
Girls were outraging over that Rs 370 biryani joke.... But you'll see thousands of such reels. I just have one question - Why can't these beggar girls earn their own money? Do you want a boyfriend, love, or sugar daddy? Everything starts and ends with money for them. And you'll never see them doing anything for the man, zero effort or gifts. Modern relationships are a humiliation ritual for men.show more

︎ ︎venom
54,334 次观看 • 1 个月前