Загрузка видео...

Не удалось загрузить видео

На главную

GPT-6 Astra vs DeepSeek V4.1 Flash is an interesting tradeoff. Astra: $10 input / $50 output per 1M tokens. DeepSeek V4.1 Flash: as low as $0.15 input / $0.60 output off-peak. That’s a massive price gap -- while the visual results here are surprisingly competitive. For production workloads, the...

13,794 просмотров • 8 дней назад •via X (Twitter)

Комментарии: 15

Фото профиля Nic Vandewetering
Nic Vandewetering8 дней назад

Deeeeeepseek is king. This is where I go when I run out of Astra usage

Фото профиля Matt Parrott
Matt Parrott7 дней назад

Comparing the price of a token here is kinda awkward. What's the price of the equivalent final results?

Фото профиля Kunal
Kunal8 дней назад

Give me prompts for these games man

Фото профиля Big Kid Nation
Big Kid Nation8 дней назад

Do a comparison for mulptle prompts "If I use multiple prompts and march however much GPT spent do I get equal or greater results with DeepSeek"

Фото профиля Pritam Chauhan
Pritam Chauhan8 дней назад

Check this as well

Фото профиля Lim EunCheon
Lim EunCheon8 дней назад

Used DeepSeek V4 Flash (previous version) for several months. No subscription API option meant high-volume use often cost more than frontier models with sub plans.

Фото профиля Alejandro Maestre | AI
Alejandro Maestre | AI8 дней назад

Veinte veces más barato por output y resultados competitivos? DeepSeek está haciendo trampa o Astra está cobrando lo que no vale.

Фото профиля Magic Turd Fish
Magic Turd Fish8 дней назад

Show me a reasonable 20 dollar a month deepseek sub with gpt-image 2.5 quality image model and texture mapping understanding.

Фото профиля Wurt
Wurt7 дней назад

This doesn’t mean anything if you don’t give the final token usage/costs of each for the final results

Фото профиля Sam Ivere
Sam Ivere8 дней назад

Ultra 6 beats deepseek in so many ways

Фото профиля Edgex
Edgex8 дней назад

True that nails consistency and detail, definitely a step up from the flash model.

Фото профиля Hisserel
Hisserel7 дней назад

The problem is that DeepSeek only wants to suppress GPT, so you shouldn't vote for DeepSeek. So far, its desire to overtake GPT far outweighs its desire to provide you with a good product. You should give GPT a chance, not DeepSeek

Фото профиля DRAMYYDS (薄肌男孩👦)
DRAMYYDS (薄肌男孩👦)8 дней назад

v4.1 is a junk model. It's literally worse than Qwen3.8-Flash-Next and more expensive. please stop the fuss. @deepseek_ai has lost it's research taste. such an ugly model

Фото профиля Arron2006
Arron20068 дней назад

You are a senior full-stack game developer and 3D technical artist. Create a complete, fully playable, single-file browser-based first-person shooter game using Three.js (or React Three Fiber + drei + cannon-es). The game must run immediately in the browser with no external assets except procedurally generated or simple geometric models. Game Title: Harbor Skirmish Subtitle: A Little Island. A Big Skirmish. Setting: A small, colorful, sleepy harbor island town called Seabreeze. Low-poly, vibrant pastel color palette (mint, coral, yellow, teal, lavender), cartoonish proportions, clean cel-shaded look, bright sunny lighting with soft shadows. Core Gameplay: - First-person shooter with WASD movement, mouse look, space to jump, Shift to sprint. - Wave-based survival: Waves of small robot-like rabbits invade the town. - Weapons (switch with 1/2/3 or scroll): 1. Ink Rifle (semi-auto, medium range, 24 ammo) 2. Scattergun (shotgun, close range) 3. Paper Blade (melee) - Grappling hook (hold E to attach to poles/buildings and swing). - HUD must include: Health bar, current weapon + ammo, score, current wave number, remaining rabbits, crosshair, "Hold E to Grapple" prompt when near grapple points. - Start screen with title "HARBOR SKIRMISH!", "WELCOME TO SEABREEZE", "A sleepy little island. A very unruly rabbit invasion.", and a "LET'S PLAY" button. - Simple wave system: Wave 1 starts with "RABBITS ON THE LOOSE!" announcement. - Rabbits have health, can be headshot for bonus points, drop score on death. - Player has 100 HP, regenerates slowly or via pickups if possible. World Requirements: - Small circular island with sandy beaches, colorful low-poly buildings (cafes, shops, lighthouse, clock tower), streets, market square, rooftops. - Destructible or interactive elements optional but preferred (signs, boxes). - Clear visual distinction between playable area and ocean. - Bright, cheerful, slightly chaotic atmosphere. Technical Requirements: - Smooth 60fps on mid-range hardware. - Collision detection, gravity, jumping. - Basic AI for rabbits (chase player, wander, attack). - Particle effects for shots, hits, and deaths. - Sound effects optional but recommended (gunshots, rabbit noises, UI clicks). - Responsive design, works on desktop. Output a complete, self-contained HTML file with all CSS and JavaScript inline. Make it look polished and actually fun to play, not just a tech demo.

Фото профиля im6
im68 дней назад

Fake

Похожие видео

China is making Dario Amodei's AI slowdown proposal worthless. The thing doing it is a 510 GB file that anyone can download for free. Two days before that essay went out, DeepSeek shipped a model called V4.1 Flash. The weights went straight onto Hugging Face under an MIT license, which allows commercial use, modification and redistribution with no royalties owed to anybody. On its own model card, V4.1 Flash beats OpenAI's GPT-5.6 Sol and Anthropic's Claude Opus 5 on four of the five hardest agentic benchmarks: - DeepSWE v1.1: 74.2 against Sol's 73.0 - AutomationBench: 54.8 against 45.8 - Agent's Last Exam: 31.8 against 26.7 - CyberGym: 88.1 against 84.5 CyberGym is the cybersecurity one. So the exact capability Amodei is asking the industry to pace is now a PUBLIC DOWNLOAD LINK. Now look at the price... OpenAI listed GPT-6 Astra on September 3 at $10 per million input tokens and $50 per million output. DeepSeek listed V4.1 Flash seven days later at 15 cents and 60 cents off-peak. Cached input runs at a third of a cent. And the reasoning benchmarks tell the same story over a longer window. Scoring 87.5% on ARC-AGI-1 cost roughly $4,560 per task in December 2024. 20 months later the same score cost about 30 cents. DeepSeek's earlier Flash build scores higher than that, 89.0%, for 2 cents. On OpenDesign's arena on September 9, DeepSeek landed 1.5 points behind OpenAI's newest model at $0.023 per task against $1.61. 1.5 points. At 70 times the price. That's the American "lead“, measured this month. And that number is what breaks the entire proposal. The essay caps the permitted slowdown at the size of the lead, because slowing by more than that lets Chinese projects pull ahead. So the slowdown Anthropic, OpenAI, Google and xAI are allowed to take is 1.5 points wide. V4.1 Flash carries 552 billion parameters but switches on only 8 billion of them for each word it reads, which is why it runs for almost nothing. It trained on 45 trillion tokens. It's sitting on Hugging Face right now. A file on a hard drive signs nothing. Nobody can UNDOWNLOAD it. And the distillation crackdown in Dario‘s plan also does nothing here either, because nothing was distilled. DeepSeek published the weights outright. There's no theft to prosecute and no copy to trace. They gave it away on purpose. DeepSeek's Pro model already hit 80.6% on SWE-bench Verified back in April, at roughly a thirty-fourth of the price of the American flagships. This has been happening all year. The frontier still wins the hardest work. On Terminal-Bench 4.0, OpenAI reports 57.9 for its newest model against the 31.2 DeepSeek reports for this one. Those four benchmark wins come from DeepSeek's own card, and nobody has replicated all of them under one shared protocol. If you look at it that way, the American labs still hold the dangerous end of the curve, and pacing that end is worth doing. But a speed limit only binds companies that can be sued in an American court. Everyone else just downloads. What exactly is Dario Amodei's slowdown protecting you from?

Ricardo

400,709 просмотров • 6 дней назад

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,272 просмотров • 26 дней назад

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 просмотров • 22 дней назад

China just made Silicon Valley's entire AI industry look like a scam. The US government spent 3 years trying to stop China from building competitive AI. But this backfired HORRIBLY. Here's what happened: Yesterday, a Chinese startup called DeepSeek released a new AI model called V4. It matches the performance of OpenAI and Anthropic's best models. At 1/7th the price. And for the first time ever, it was built on Chinese chips. NOT American ones. That last part is the one that terrifies the west. For context: Since 2022, the US has banned the export of advanced AI chips to China. The entire strategy was built on the assumption that if China can't access Nvidia's best hardware, they can't build frontier AI. But DeepSeek just proved that assumption wrong. Their V4 model was trained and runs on Huawei's Ascend chips. Huawei spent months working directly with DeepSeek to make sure V4 runs across their entire line of AI processors. Jensen Huang even predicted this on a recent podcast: "The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation." That day was yesterday. And the numbers are crazy: DeepSeek V4 costs $3.48 per million output tokens. OpenAI's latest model GPT-5.5 costs $30. Anthropic's Claude charges $25. Same ballpark performance. 7x cheaper. Uber's CTO just admitted they burned through their ENTIRE 2026 AI budget in 4 months using Anthropic's tools. If Uber had used DeepSeek instead, that same budget would have lasted 7 YEARS. 4 months vs 7 years. Same work getting done. But the pricing isn't even the big thing here. The real story is what DeepSeek did with their technical report: They published the benchmarks where they LOSE. Every AI company cherry-picks the tests where their model wins. DeepSeek ran the full comparison against GPT-5.4 and Google's Gemini, found they trail frontier models by 3 to 6 months, and printed it anyway. They literally don't care because the price gap makes the performance gap irrelevant for 90% of use cases. So the US export controls didn't slow China down. They ACCELERATED China's independence. Because Chinese developers were FORCED to train models with limited resources, they had to figure out how to make AI radically more efficient. That constraint became their competitive advantage. Every generation of DeepSeek has gotten dramatically cheaper to train. V4 continues the trend. Meanwhile US companies are going the OPPOSITE direction: OpenAI's GPT-5.5 Pro costs $180 per million output tokens. That's 51x more expensive than DeepSeek V4 for comparable work. The Commerce Secretary confirmed this week that ZERO Nvidia advanced chip shipments have actually gone through to China despite being approved in January. So China built frontier AI anyway. Without American chips. At a fraction of the cost. And the market response tells you everything: Chinese chipmaker SMIC surged 10%. Huahong Semiconductor jumped 15%. DeepSeek's Chinese AI competitors Zhipu AI and MiniMax dropped 9% because V4 is destroying them too. DeepSeek is making Silicon Valley's pricing model look like a scam. US tech companies spent $650 billion on AI infrastructure this year. DeepSeek just showed the world you can match their output for pennies. The export controls were supposed to be America's ace card. Instead they taught China how to win without American chips, at American prices nobody can compete with. Jensen Huang was right. This is a horrible outcome. But it's the outcome America built for itself.

Ricardo

281,190 просмотров • 4 месяцев назад

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,997 просмотров • 24 дней назад

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 просмотров • 1 месяц назад

Switch Transformer by hand ✍️ ~ 13 steps walkthrough below The Switch Transformer, by Fedus, Zoph, and Shazeer in 2022, is one of the papers that made sparse Mixture of Experts practical at scale. Today, frontier models use MoE to pack enormous parameter counts while activating only a small slice per token: GPT-4, Claude, DeepSeek-V3, and Kimi all follow this pattern. If you want to understand how those models can be huge to store yet still cheap to run, this paper is a good place to start. How does it work? Goal: run five input features through attention, route each one to a single best expert, and read the output off the page. = 1. Given = Input features X1-X5 arrive from the previous block. = 2. Attention matrix = Feed all five features to a query-key attention module to get an attention weight matrix A. = 3. Pooling = Multiply the input features by A to get attention-weighted features Z1-Z5. The effect is to combine features across positions. = 4. Visualize pooling = Z4 is X4 + X5 because the fourth column of A is [0,0,0,1,1]. = 5. Gate values = Multiply the weighted features by the switch matrix. Each gate value says how well expert A, B, or C can probably handle the feature. = 6. Top expert = Pick the row with the highest gate value. Sparse means only the top expert is selected, not all of them. = 7. Routing = Route each Z to its best expert. Every expert has a fixed capacity of 2, so one feature may overflow. = 8. Expert A, linear = Apply the linear layer to the features routed to Expert A. The effect is to combine features across feature dimensions. = 9. Expert A, aggregate = Send the combined feature to the corresponding output column. = 10. Expert B, linear = Apply the linear layer, as in step 8. = 11. Expert B, aggregate = Send the result to the corresponding output column, as in step 9. = 12. Expert C, linear = Apply the linear layer, as in step 8. = 13. Expert C, aggregate = Send the result to the output column. Since one feature exceeded Expert C's capacity, it passes through as-is. Takeaway: a Switch Transformer keeps attention unchanged, then replaces the dense feed-forward network with a sparse set of experts. Most of the parameters sit in the experts, but only a small fraction are used for any one input. That is how GPT-4, Claude, DeepSeek-V3, and Kimi can be enormous to store and still cheap to run. 💾 Save this post!

Tom Yeh

37,124 просмотров • 1 месяц назад