Загрузка видео...

Не удалось загрузить видео

На главную

Gemini 4 argon vs GPT 6.1 sol both models were tested with same prompt but results came out really different > sol was tested by me in devin cloud at default reasoning > gemini was tested by lentils earlier this month on pre release checkpoint argon isn’t publicly available...

14,442 просмотров • 2 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

before pumpfun livestream feature is updated, before my account goes big and before I got some money, before everything, there was this token the beginning was so small at that time(over 2 years ago) rug was still rampant and I thought like 'why wouldn't they grow their project bigger rather than rugging at 10k?' sadly I became one of them now but regareless of it, I launched it just for 100% fun with buying 1 sol and turned on camera at TG group just for fun too someone said, "yo bro you should keep doing this this gonna be hella huge" so I did it I still remember the 2 guys who carried the whole project with max shilling and leading community members: Noble(this guy was pretty mean to me lol but still he was a goat) and Cassius(actual goat) it was a pure joy I did a stream 24/7 even while I was sleeping at TG and people were having fun in here(one girl took off her shirt when I was sleeping and I fucking missed it 💀) and my big bro Tyrelle Anderson-Brown came into my coin and helped me with 2 sol. I still remember this thankful money. with this I ate a nice dinner with my gf I still don't know the reason(maybe money laundaring?) but it went 5M and at this day when I woke up and checked my trojan, THIS WAS THE FIRST TIME I HIT 6 FIGS IN MY LIFE and I didn't sell a penny because community was more important than money at that time here's a list what community did for me - bought me a new iPhone - bought me a new MacBook(these two were for a better stream - my PC and phone was trash) - funded me almost every equipment for stream - formed a team with 10-11 members and kept supporting me they even put my sleeping video at Timesquare, NY here's a video so how can I sell this lmao but sadly the coin goes up, the coin goes down too and this happened to my coin too it was sooo tough days but I kept doing my best and it ended from up 130k to making 3k only this project was like if someone asks me "what did you do in this year? can you answer to this question with confidence?", I will answer this coin with 100% sure and I just turned on livestream with this coin just for fun too was very good days

letterbomb 🟪🔶🟦⟠

14,212 просмотров • 5 месяцев назад

AI has had exactly two scaling axes that worked so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inference

Sasha Malysheva

15,053 просмотров • 1 месяц назад

Astra (GPT-6) is here!!! I've had early access and tested it like crazy with things like games, code, writing, browser control, presentations and general knowledge work. This is the best model I've ever used. Period. (Incredible demos below in this thread ⬇️) Here's my take on Astra: > It's insanely capable. This feels like a massive improvement, not just an incremental change. This is especially true with zero-shot prompts. > It's all about knowledge work. Slide creation, analysis, writing, and browser control. And oh my...it's so good at browser control. GPT-5.6 was already fantastic at doing things in the browser, Astra is another level and significantly faster. > We're closer than ever (arrived?) at prompt-to-playable game. And I don't just mean only playable, these are actually fun games. I bet if someone with a great eye for games used Astra, they could create a viral game within 1-2 weeks. > Astra is better at writing but not perfect. It removes much of the "AI Smell" we're all familiar with but some stink still survived. > It has a tendency to use the same design colors and look/feel as GPT-5.6 (forrest green anyone?) but it is more steerable in design than previous models. > It's highly steerable in general. A little nudge goes a long way. When I first started using Astra, almost every task I gave it would go for ~30 minutes. I wanted it to keep working. Adding more specifics to a prompt helped greatly with it's ability to work for a long time. > Astra's 3D understanding is unmatched. 3D asset creation was consistent and easy and its spacial awareness while building complex 3D worlds blew me away. I'm still getting familiar with Astra but this will now be my go-to model for any difficult work I have. Check out the demos below: 👇

Matthew Berman

1,907,120 просмотров • 29 дней назад

zb1 is trending on theqoo HOT category, where over 200+ comments are praising zb1’s vocals, dance, everything, especially their cover of “sherlock (clue + note)” by shinee “damn the determination is insane, even with just five members the stage still feels full” “honestly, this was the most memorable stage for me at kcon today lol. the nerve to cover “sherlock” live is insane, they did so well. i didn’t realize zb1 were this talented” “sherlock is such a difficult song, so honestly i’m shocked a 5th gen group can pull this off. they’re seriously good” “the determination was crazy. this song is really hard and keeps repeating high notes… and the dancing was perfectly in sync too” “just deciding to cover sherlock already shows legendary determination lol but they seriously did so well” “since this is their first stage after reorganizing into five members, i love that you can feel the determination right from the song choice. they did well” “i guess because the current members are all main vocal or strong vocal members, they’re insanely good” “sherlock is insanely high but hearing them do it live so powerfully feels so satisfying. zb1 are insanely good” “it’s been a while since i’ve seen a stage this full of determination, my dopamine is going crazy” “i usually don’t watch other idols’ stages but the determination was insane so i got totally focused. they’re seriously good” “wait this sounds fully live though? sung hanbin is surprising, i thought he was only good at dancing but he sings well too??”

☁️ (IA)

15,880 просмотров • 4 месяцев назад

A lesson for every Polymarket bot developer: I built a strategy that looked perfect on paper. Backtested it. Looked like a winner. Almost went live. Then i actually measured real costs. Strategy was dead before the first trade. And this is what bot building on Polymarket actually looks like. Here is what happened (and what you MUST know): Backtested mean reversion on crypto dips. SOL came back at +44% return and 70.5% win rate. Beautiful clean curve. Looked ready to ship. Then i measured real round-trip costs on SOL flash dips. Backtest assumed 0.45% in fees and slippage. Reality was 1.44%. Strategy stops working at 0.70%. Starting again. But the lesson was worth more than any profit the strategy could have made. Here is what i actually learned: Taking a dip with a market order means you eat the spread the dip just created. The volatility making your signal is the same volatility destroying your fill. You see the opportunity. You enter. You already lost. But resting a limit order below market and letting the dip come to you? You collect the maker rebate instead. Same thesis. Completely opposite execution. One bleeds money, one prints it. That one realization changed how i think about bot strategy entirely. 180 strategies tested to get there. 179 dead. That is not failure. That is how you find the 5 that actually work. Building a bot on Polymarket is not about finding a magic strategy. It is about eliminating every wrong answer until only the right one is left.

Oracle Boar

14,379 просмотров • 5 месяцев назад