Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Battle of the new coding models: Cursor vs Cognition They both make big claims being near the frontier - how well do they pass the Golden Gate Bridge test? I've included 5x non-cherry picked generations from each, followed by examples from pre-release test of Gemini 3 Pro, GPT-5 and...

23,583 görüntüleme • 10 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Opus 4.7 - 400k vs 1m context - is there a difference? I've heard Theo - t3.gg talk about the fact that it is unlikely that Anthropic would have offered up a model with 1m context at the same cost, if it wasn't a different (i.e. cheaper to serve) model. I did a test where I toggled the 1m default model on & off in Claude Code (otherwise default settings, xHigh reasoning) and compared the outputs with 3x generations - same prompts etc. My observations: - Models feel DIFFERENT - often when you ask a model for the same generation, you get a somewhat different answer, but it feels & smells the same. Here 400k and 1m are very different every time - 400k model seems better - not that 1m is trash and 400k is amazing, but there are definitely issues with the level of ambition and accuracy that 1m model seems to have Examples of 1m failing: - Voxel Rome: the colosseum is nowhere near as impressive - Golden Gate: cars go sideways, waves not very high, bridge goes into land; though the structure of the bridge is a bit better - Stonehenge: structure is more 'wrong', lighting, shadows & textures are more flat and not as rich This isn't a conclusive evidence of course, but at least to me the two models do not behave the same way. Anecdotally as well when building 1m felt like it was doing more weird validation (e.g. going around in circles) and 400k was more straightforward. These sorts of things are harder to capture in tests, but you'd notice in Claude Code. You can review the hosted generations, see the code & prompts in the links below

Peter Gostev (SF: 22-26 June)

29,203 görüntüleme • 4 ay önce

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 görüntüleme • 2 ay önce

Baseten Head of AI Model Training Charlie O'Neill says the future is many specialized LLMs dedicated to specific tasks, with bigger labs deployed on the frontiers of areas like science and math: "People are thinking about intelligence capabilities in the wrong way. People are thinking about intelligence relativistically. They say, 'OK, the open-source gap is like 6 months behind closed-source, and GLM 5.3 is as good as Opus 4.8,' or whatever." "The best way to think about what models can do for you, and for the world, is in an absolute sense." "So for any given task that you want to do with an LLM, there's some intelligence threshold where below that you can't do the task, and above that you have very diminishing returns to more intelligence on the task." "So when you think about it that way, the game of LLMs over the last 5 years has been, 'OK, we have these things we want to do with them. Closed source hits it first... but open-source can eventually do that task. And then for many reasons, once you have the base level of intelligence required to do it, you probably do want to swap to open-source." "It's not really about the [frontier lab] God model being better. Like, if I'm filing a tax return, there is a limit to how much intelligence I need to do that particular thing." "So I think the world is going to look like — frontier closed-source labs are going to continue to push the frontier. You do want to use the most intelligent model. You have very inelastic demand for intelligence when you're doing frontier science or frontier math." "But for a lot of the economically valuable things, it looks a lot like, 'I'm a Cursor, or I'm one of these big companies who are realizing I can't just be a wrapper anymore. I've been through the life cycle of building a product that people love. And I should be using that information to make my model better at the things that I care about, and not at anything else.'"

TBPN

46,566 görüntüleme • 1 gün önce

How good is GPT-4-Vision at extracting text from images? I wanted to find the limit - but I found weirdness instead Most surprising: GPT-4V performance varies depending on the *structure* of text it sees Let me explain A set of images with progressively more text was presented to GPT-4-Vision. GPT-4V was asked what text it saw in the image. The response from the model was compared against the image’s original text and scored for similarity. The model was tested with 4 types of text: essay, random words, random tokens, and random characters. Findings: * Performance degrades - Yes, the models are good at basic OCR, but as you get more text and words then performance drops (this is expected) * Type of context matters - You should expect different recall on your texts based on your context types * Hallucination Errors - I thought that the model would make errors of omission (it wouldn’t return all the words). But instead the model mostly made hallucination errors - it replaced words with made up words. * Evals Matter - This test in isolation doesn’t mean that your data will have the same results, but it should motivate you to create eval tests for your data and anticipate errors which are hard to spot Notes: * Next step would be to add additional image types like tables or PDFs * GPT-4V would routinely get stuck in repeat-token-loops when trying to extract random tokens * GPT-4V would refuse to answer most random character images

Greg Kamradt

49,111 görüntüleme • 2 yıl önce

He's right, I do have it in my notes because he did say it! 😆 On multiple instances, but here's him saying it on Trader Round Up: Timestamp 1:00:00 “They have to have the last two Mondays and the last two Fridays always, because they are the beginning and the end of the week.” ---- Timestamp 59:53 "In fact, every Monday, you should have two Mondays and two Fridays. In my opinion, if you’re going to ask me, like I teach my sons, they have to have the last two Mondays and the last two Fridays always because they’re at the beginning and the end of the week. So that’s what Cameron made money with today. He used information that is gleaned from that range, where we’re at in price action, and using the higher timeframe expectancy that we saw, that rejection when it went up to take that short term high." ---- Timestamp 1:00:46 "But every one of you listening, if you were my kid, if you were sitting with me, you would be hearing me tell you: Do you have two Mondays worth? Do you have both Fridays? Present today’s Friday and then last Friday’s Opening Range high, low and close, and the gradient levels on it? And then where are we in relation to that? Are we significantly above it, significantly below it, are we in close proximity to it? Because Monday’s Opening Range gap and Friday’s Opening Range gap are like a big huge draw on liquidity. It'll act like a big magnet. Not all the time, it’s not a panacea, but if we’ve traveled a lot one way, those Opening Ranges tend to be a factor on the close of the week and the beginning of the week. So Monday’s trading and Friday’s trading, using that framework that includes two weeks worth of Monday and Friday, have them in there. Even if you don’t want to carry Tuesday, Wednesday, and Thursday, the week prior to last week at least have that previous Monday in there, because that way you’ll gonna have a full range of opportunity that the algorithm will refer back to."

Pipmunch

187,151 görüntüleme • 1 ay önce