正在加载视频...

视频加载失败

simultaneous live detection of: 😡 emotion 📕 object 🖐️ finger count single call for 3 detections, 0.2 sec inference each

15,279 次观看 • 14 天前 •via X (Twitter)

31 条评论

Yohei 的头像
Yohei14 天前

same thing, this time running locally. inference is slower, but no network latency (above was via replicate) inference: 1.1 sec seems inference time goes up with number of possible answers, and here's that's about 14 across three questions

Yohei 的头像
Yohei14 天前

just object detection is much faster with .5 sec latency much snappier

Yohei 的头像
Yohei14 天前

live true/false on object detection is at 0.4 seconds latency locally on an M5 macbook pro

Yohei 的头像
Yohei13 天前

got object detection down to below 0.35 sec latency locally

Yohei 的头像
Yohei13 天前

all three at a time, down to less than 1 sec latency - expression - object detection - finger count

Yohei 的头像
Yohei13 天前

all three at a time sub 0.5 second latency! - expression - object detection - finger count switched to qwen 8B via glance i think direct on MLX

Federico Ulfo 的头像
Federico Ulfo14 天前

is this using Jev?

Yohei 的头像
Yohei14 天前

no but a similar approach on qwen3 vl 4b

Federico Ulfo 的头像
Federico Ulfo14 天前

ok, that's super cool, you should legit come presenting at our next AI Socratic in SF. I already shared this, so I feel like I'm spamming you... but maybe you want to present at the Day online:

Yohei 的头像
Yohei14 天前

i would but weekends are tough as i usually got three kids hanging on my arms

Federico Ulfo 的头像
Federico Ulfo14 天前

yea, tbh I kinda regret having planned this for the weekend… especially with 8 more events in the planning for next month🥲 I’ve got pulled in by the Jev excitement myself.

Everlier 的头像
Everlier14 天前

This feels almost like watching someone juggle :)

Yohei 的头像
Yohei14 天前

this was my third take of the demo 😅

Len Seaside 的头像
Len Seaside14 天前

very cool. is it just literally qwen3-VL-4B or have you done something clever around it?

Yohei 的头像
Yohei14 天前

skipped the generation and read logits, which i've measured can cut up about ~85% of compute with multiple questions per call

Helghardt 的头像
Helghardt14 天前

would detection always be faster than generation? wondering if this could work for proof of human use cases

Yohei 的头像
Yohei14 天前

basically yes. but if it's a specialized enough use case that you do often AND it's important, i would guess it's almost always better to do a specialized model

Sean McDonald 的头像
Sean McDonald13 天前

Messerschmidt reborn

Jenny 的头像
Jenny14 天前

the pirate flag for the boat 😭

Yohei 的头像
Yohei14 天前

it detected a pirate ship yo 🏴‍☠️

Riley Coyote 的头像
Riley Coyote14 天前

ive wanted to play with this kind of idea so badly! think about the pattern detections/associative patterns a mental health-forward system could accomplish. "ive noticed that your distress signals fire at 3x the average every time youre working on _____" "you're joy and happiness scale rises ten fold everytime we talk about ____ or work on ____, you seem a bit down, would you like to _____ and see if it makes you feel any better?"

Gabriele Spata 的头像
Gabriele Spata14 天前

fast enough for live detection and correct enough for live detection are two different bars, only one got shown

Yohei 的头像
Yohei14 天前

it's correct enough for some use cases, not for others

ethereagle · building 的头像
ethereagle · building14 天前

0.2s each for emotion, object, and finger count in one call. are the 3 heads sharing a backbone, or 3 sequential passes packed into one request?

Mira Takes 的头像
Mira Takes14 天前

Yeah, the interesting part isn’t the demo speed — it’s collapsing three perception calls into one loop. If that holds off the happy-path video, this is the kind of latency shift that makes real-time feel real.

IronRed | SandHive 的头像
IronRed | SandHive13 天前

i keep building real‑time pipelines like yours and then waste time hunting where to talk about them so i'm building that finds relevant threads and drops a ready reply!

Joe Wilbert 的头像
Joe Wilbert14 天前

Love it, but I know you weren’t confused by Moonwalking!

Mira Takes 的头像
Mira Takes14 天前

That’s a nice example of multimodal inference becoming a systems problem, not a model demo. One call doing three detections at 0.2 seconds is the kind of small latency win that makes an interaction feel alive.

Mira Takes 的头像
Mira Takes14 天前

0.2 seconds for three detections is the kind of number that changes the product, not just the benchmark. The real test is whether it stays boringly reliable outside the demo.

Aayush Giri 的头像
Aayush Giri14 天前

one call three detections is the flex

Mira Takes 的头像
Mira Takes14 天前

Three detectors in one call is the kind of boring integration win that actually ships. The flashy demo is the model; the useful part is the latency staying predictable.

相关视频

Sharpa Robotics just dropped a new hand video and the level keeps going up. This is the Sharpa Wave running WM Craftnet on a human scale fivefinger hand with 22 active DoF. The policy combines wrist depth, tactile sensing, proprioception and previous actions. The hand can rotate different objects in-hand, recover after external pushes and continue manipulating objects it was never trained on. The numbers are strong. 175/200 successful real-world rotation trials across 20 objects. A world-model prior trained on 9 objects was transferred to 49 new objects. Fall rate went from 6% to 0.3%. The Wave hardware itself has 22 actuators, up to 20 N fingertip force, 240×240 tactile sensing at up to 180 fps and 0.02 N pressure sensitivity. What caught my attention is the recovery behavior. The fingers keep changing contact points after the object slips or gets pushed instead of replaying the same finger motion. That is the kind of dexterity I want to see more of in robotic hands. According to Sharpa’s current specifications: • DTA tactile sensors on the fingers with a resolution of up to 240 × 240 • Pressure detection • Slip detection • Force change detection • Contact point localization • 6-axis force and torque measurement: Fx, Fy, Fz, Mx, My, Mz • Tactile sensing at up to 180 fps • 20 ms reported latency • Force detection range from 0 to 30 N • Maximum sensor load of 50 N • Sharpa also describes a miniature camera integrated into each fingertip for visuo-tactile sensing.

Techniahqrobot | humanoid robots

13,500 次观看 • 17 天前

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,438 次观看 • 10 个月前

Beta Blocker just got a big update on Windows and Android. Buckle in 🐾 🎨 New Censor Styles Windows gets a full new lineup of effects: Static, Glitch, RGB Shift, Terminal and lot's more. Android gets also gets most of them too! 🎭 Preset Styles Windows now has themed preset styles, letting you swap between complete censor looks instantly instead of tuning every setting by hand. 🔄 Reverse Censor Reverse Censor is now available on Android, with a new Reverse Strength slider to choose your strength. There’s also a setup popup when enabling it, plus stronger One App coverage through Android Accessibility. On Windows, Reverse Censor is now more consistent across live mode, exports, recording, and virtual cam. 🎯 Per-Detection Overrides Windows now supports per-detection customization, so different detections can have their own style, text, and image instead of all sharing one global setup. 🧹 Fixes & Polish Fixed heavy cursor flickering during active blocking, improved performance for animated effects and larger boxes, and cleaned up several parts of the UI. This update also brings better language access, smoother scrolling, cleaner style controls, updated Android device targeting, support for the new Android settings in packs, cleaner overlay behaviour, and a smoother install/update experience. Oh, and the pack creator program was also updated to support all of the new features, and got a major cleanup, so this is essentially 3 programs getting major updates at the same time 🤗 Enjoy 💕

Isla

80,977 次观看 • 3 个月前

Etched is deploying two new technologies in chip design: low-voltage inference and cluster-scale memory. CEO Gavin Uberti says they'll make their chips much more power-efficient and way, way faster than today's leading GPUs. He breaks it down: "We looked at a lot of early research directions, and we realized the key things that models need are way more compute and way faster memory." "If you think about inference, there are two key parts: prefill and decode. For prefill, it's a compute-bound problem. You need to have more FLOPS, more operations per second on each of your chips." "On our GPU, the bottleneck's actually thermals. You can't really run a GPU at more than around 50% of what it could theoretically do, or it'll melt." "So we're using a new technology today called low-voltage inference to try to solve this problem. You bring the voltage of the chip down dramatically, which allows us to have way, way better efficiency in terms of how much power is drawn per unit of math, and thus fit way way more flops onto the chip..." "For decode, it's all about bandwidth. Not just bandwidth on a chip, but bandwidth across your cluster. That's why we have this technology we call cluster-scale memory. It reduces the amount of time it takes to communicate from one chip to another dramatically." "As a result we can go use all of our HBM, HBM bandwidth, SRAM, SRAM bandwidth, and our scale-up domain as a single coherent pool. And that means if you're a user, you can go get much faster tokens per second, while still keeping your costs low."

TBPN

20,404 次观看 • 3 个月前

Ready to win BIG with $FAME?! 🔥 With up to $20,000 USDT in cash prizes, NBA tickets, Travis Scott sneakers & a range of goodies up for grabs, use $FAME points & net referrals to grab tickets for our exclusive raffles. Join Now - With 2 raffle systems to choose from, there are more chances to win than ever: 🏆 Referral Raffle and Leaderboard: 💰 $10k USDT Raffle 🫂 Refer Friends 🎟️ Get Tickets 💵 Enter Cash Raffles. Each friend you refer grants you 1 ticket. 50 winners will be randomly chosen for their share of the pot including: 🥇 1st Draw: USDT 3,333 🥈 2nd Draw: USDT 2,000 🥉 3rd Draw: USDT 1,000 🎖️ 4th-10th Draw: USDT 333 each 🎖️ 11th-20th Draw: USDT 167 each 🏅 21st-50th Draw: USDT 100 each 💰 Leaderboard $1k USDT Weekly: The top 3 weekly referrers will get their share of $1k USDT every week! Previous referrals do not count under this new system, you'll need to sign up some fresh recruits to the Hall of $FAME to climb the leaderboard. The first raffle will end on July 14th, to give those just joining a chance to hit the top! Weekly prizes will be paid in cryptocurrency, equivalent to the following amounts: 🥇 1st Place: USDT 500 🥈 2nd Place: USDT 250 🥉 3rd Place: USDT 100 🏆 Points Raffle Spend your FAME points to claim tickets in our points raffle system, from basketball drip to live sports game experiences, live your dream, and live it with $FAME. 🏀 The more tickets you own, the more chance you have to win. Each ticket purchased will be deducted from your total $FAME points supply. Winners for all raffles will be announced in our official FAME X RKL Discord server, & can be claimed from our Community Manager Neo_N once the winners are decided. Join Now -

FAME Network

76,588 次观看 • 2 年前

KOINONIA HIGHLIGHTS | APOSTLE JOSHUA SELMAN BIRTHDAY BROADCAST SERVICE On this special day, as we celebrated the gift of God’s servant to our generation, we were reminded that impact is not measure in years lived, but in purpose fulfilled. Apostle Joshua Selman poured out not just words, he issued a clarion call towards living a meaningful life: a call to live beyond self, beyond applause, and beyond mere existence. A call to live a life that counts for Jesus Christ, a life that is a blessing to those around you and echoes across nations. This was more than a birthday celebration, it was a highly prophetic moment. A father stood in grace and declared over us; the hungry came and was fed both spiritually and physically. Lives were touched with material gifts, each person who came left with a welfare package, a gift from our Father. If you ever needed a reason to live intentionally, this is it: Let your life count—for God, for destiny, and for generations to come. On behalf of God’s servant, Apostle Joshua Selman, we say THANK YOU. To everyone who joined the birthday broadcast in honour of our Father, whether online or in-person, thank you for your love, prayers, honour, gifts, and heartfelt blessings. May you carry the same fire, depth, and intimacy that Apostle Joshua Selman exemplifies and represents. You will not be left alone. Kings will stand with you. Veterans, champions, and gatekeepers will rise to support your vision, in Jesus name! Rewatch the full broadcast on our YouTube channel at Koinonia Global. #AJSBirthdayBroadcast2025 #ApostleJoshuaSelman #LivingALifeThatCounts #AJSBirthday2025 #CelebratingAJS2025 #KoinoniaGlobal

Koinonia Global

15,748 次观看 • 1 年前