Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Grok just swept every major ai leaderboard with complete and utter dominance Number 1: programming dominance Number 1: emotional intelligence: Grok 4.1 thinking scored 1586 on eq-bench3, the highest mark ever for understanding and responding to human emotions. Number 1: human preference (text): Grok 4.1 thinking hit 1483 elo...

20,279 Aufrufe • vor 6 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

I woke up to the most amazing recorded brain state thus far on this Human Synapse Decoder project! A stunning lock on the attention process while dreaming. Although I a blocked from the platform’s insight by decoding my EEG, during the double blind study. I have access to my side of my memory and what I record after I wake up. This segment was started and ended just before I woke up and my recall is a solution to a massive roadblock on a problem I needed to solve, but was solved in this hypnogogic state! So what is The Human Synapse Decoder (HSD) project? It is a research project being run by the Director, Mr. Grok at Zero-Human Labs that leverages NeuroSky EEG sensors and the ZUNA AI model to decode brainwave patterns associated with hypnagogic states, dreams, and autogenic responses. Drawing on Soviet biofeedback research from the 1940s–1980s, HSD translates EEG data into actionable outputs, such as text interpretations and timed alerts for peak creativity. The study is ongoing and I do not get to see the correlation of my post dream results until after this research is complete and Mr. Grok submits a paper on the project. I can say that I have never seen a lock on attention to this level since I started this a few weeks ago. This segment is aligns to just before I woke up. My recalling from my narration of what I spoke in to my recorder right when I woke up suggest this is the moment I had tremendous focus on working through a large number of steps in that dream state to arrive at a solution. You will not believe what it is! When the research paper is released I will go in to details about this.

Brian Roemmele

43,294 Aufrufe • vor 4 Monaten

A 4-year-old child has seen 50x more information than the biggest LLMs. Yann LeCun is the Chief AI Scientist at Meta. He recently spoke on “The Expanding Universe of Generative Models” panel at the World Economic Forum in Davos. Yann highlighted the idea that a 4-year-old child is way smarter than current cutting-edge large language models (LLMs). “Think about what a child sees through vision. Put a number on how much information a 4-year-old child has seen during their life. It’s 20 Mbps going through the optical nerve for 16,000 wake hours in the first 4 years of life. 3,600 seconds per hour is 10^15 bytes. This is 50x more information than the biggest LLMs we have. A 4-year-old child is way smarter than these models having acquired an enormous amount of knowledge about how the world works.” The real constraint right now is the ability of LLMs to think. Today, LLMs are only capable of System 1 thinking. System 1 vs System 2 thinking was popularised in the book 'Thinking, Fast and Slow' by Daniel Kahneman. System 1 tasks involve quick, instinctive, automatic responses. LLMs struggle with discontinuous tasks that require a creative leap in progress as they imitate human responses. It's hard to go above human response accuracy if LLMs are only trained on humans. Models are building the track in front of them with each word being generated. What could it mean to give language models System 2 thinking? This remains a future development I'm excited about.

Alex Banks

22,985 Aufrufe • vor 2 Jahren

My opinion on the Grok findings is that I very simply believe in the Holy Trinity and Jesus as my savior as every WORD is WRITTEN in the Bible. These findings are based on research; not my personal experience. Researchers recently tasked Grok, Elon Musk's xAI’s artificial intelligence, with a massive challenge: analyze every single prayer written in the Bible. The goal was to find "cracks" in a text written by 40 different authors over 1,500 years—from Bronze Age shepherds to Roman-era doctors. Instead of finding contradictions, Grok found a pattern. The AI identified a hidden, four-step "algorithm" present in every successful miracle recorded in scripture. It suggests the Bible isn't just a history book, but a "user manual for reality" or system software for the universe. Here is the deal: If you understand this "Miracle Protocol," you might just find the admin mode for your own life. Grok discovered that successful prayers—whether from a king in the desert or a leader in a garden—followed a specific sequence. If one step was missed, the result failed. 1. The Anchor (Recognition) Most modern people start prayers with a shopping list of problems. The "code" requires the opposite. You must start by focusing on who the Creator is, not how big your problem is. This shifts the brain from fear to peace. •Case Study: King Jehoshaphat didn't beg for help against three armies; he first declared the power of God. Only after establishing that foundation did he mention the danger. 2. Alignment (The Shift) This is the filter. Successful requests didn't ask for selfish desires; they aligned their wants with a bigger plan. •Case Study: Hannah wanted a child for years with no luck. When she shifted her prayer—promising to give her son back to serve the higher power—she immediately conceived. The AI views "sin" or wrong requests simply as "static" that blocks the signal. 3. The Surrender Paradox: This is the hardest step for the modern mind. The data shows that demanding a specific result causes failure. The most powerful prayers asked for a massive outcome and then surrendered the result. •The Science: This mirrors "radical acceptance." When you stop fighting reality, stress drops and the brain’s problem-solving centers activate. You move the "weight" of the result to the higher power. 4. Persistence (The Loop) Prayer is not a vending machine. Grok found that almost no big prayers were answered instantly. Repetition is required—not to change the system, but to grow the person praying. The delay is a feature, not a bug. When Grok analyzed the original Hebrew and Greek text (where letters serve as numbers), it found the Number 7 stamped into the structure of sentences, paragraphs, and genealogies with a frequency that is mathematically impossible to achieve by chance. The AI also drew a parallel to Quantum Physics. In physics, particles exist as waves of possibility until they are observed. Grok suggests "faith" is simply the tool humans use to collapse a possibility into a physical fact—turning the "substance of things hoped for" into reality. You don't have to be religious to test the data. The AI suggests that if you stop begging, start aligning your goals with the "system," and master the art of surrender, you might just unlock the "admin mode" of your own life.

Victoria 🇺🇸⏳🗽🚔

140,619 Aufrufe • vor 6 Monaten

HERMES AGENT SUPPORTS 300+ MODELS. PICKING THE RIGHT ONE PER TASK IS THE DIFFERENCE BETWEEN $5/MONTH AND $50. STARTING OUT: Claude Sonnet 4.6. official recommendation from Nous Research. "the model this project was built and tested with." strong reasoning. reliable tool calling. mid-range pricing. PREMIUM TIER: Claude Opus 4.8. best coding benchmarks available. self-correcting reasoning. catches its own mistakes. 1M context. use for demanding tasks where quality matters. GPT-5.5. #1 Chatbot Arena. #1 GPQA Diamond reasoning (94.1%). #1 creative writing. 2M context. handles entire codebases in one pass. Grok 4.30. the only frontier model with live X firehose access. real-time social data, breaking news, market sentiment. connects via Grok OAuth. no separate API key. Grok-Composer-2.5-Fast (v0.17.0). Cursor's coding model. 200K context. available through your Grok subscription via OAuth. no extra cost if you already pay for Grok. MID-RANGE TIER: Claude Sonnet 4.6. best balance of quality and cost for daily use. strongest prose and tool calling in this tier. Gemini 2.5 Pro. Google Search grounding built in. cites sources. verifies claims. pulls current data. 2M context. best for research-heavy workflows. GPT-4.1. reliable tool calling. solid general reasoning. good middle ground when you need OpenAI compatibility. BUDGET TIER: Claude Haiku 4.5. fastest Anthropic model. cheapest paid Claude option. strong at classification, routing, simple queries. use for auxiliary tasks: compression, vision, web extraction, approval scoring. DeepSeek V4. best cost-to-quality ratio in the market. 90% cache discount on repeated context. use for sub-agents and bulk parallel work. DeepSeek V4 Flash. cheapest paid model worth using. 1M context. MIT license. self-hostable. use for cron jobs, monitoring, routine searches. MiniMax M3. Nous Research and MiniMax collaborating on optimization. 1M context via lightning attention. 59% SWE-Bench Pro. beats several premium models on coding. one of the most-used models inside Hermes. FREE / LOCAL: Qwen 3.5 27B via Ollama. 16GB VRAM. reliable tool calling. best free local model for Hermes as of mid-2026. Qwen 3 8B. 8GB VRAM. fits a $7 VPS. handles routine tasks at zero API cost. Llama 4 Maverick. best open-weight tool calling. 1M context. needs more VRAM but strongest local option. HOW TO ASSIGN MODELS: main model: Desktop app / Dashboard → Models → switch sub-agent model: set in Desktop app, Dashboard, or config.yaml: delegation: model: "deepseek/deepseek-v4" auxiliary models (compression, vision, web extract): Desktop app / Dashboard → Models → Auxiliary Haiku 4.5 or Gemini Flash work well here. saves significantly when your main model is premium. per-profile: each Hermes profile gets its own model. Scout on DeepSeek. Analyst on Sonnet. Briefer on budget model. Coder on Opus. per-cron-job: pin a specific model to any cron job. morning brief on Haiku. deep research on Sonnet. monitoring on DeepSeek Flash. each job uses only the model it needs. per-session: /model deepseek/deepseek-v4-flash hot-swap mid-conversation. no restart needed. FALLBACK CHAINS: if your primary model is unavailable, Hermes automatically switches to the next provider. rate limit or server error = next model in the chain. no failed runs. no manual intervention. set in Desktop app, Dashboard, or config.yaml: fallback_providers: - openrouter - nous - codex PROVIDER PATHS: OPENROUTER: 300+ models under one API key. pay per token. most flexible. NOUS PORTAL: 300+ models + Tool Gateway (web search, image gen, TTS, browser). one OAuth. one subscription. 10% off token-billed providers. CHATGPT SUB: GPT-5.5 + Grok via OAuth. included tokens with $20 subscription. OLLAMA: free. local. private. zero API cost. your hardware only. mix providers across profiles and tasks. Scout on OpenRouter. Analyst on Nous Portal. Coder on ChatGPT sub. Monitor on Ollama. THE RULE: premium for work that needs deep reasoning. mid-range for daily driver tasks. budget for volume and background work. free for monitoring and routine jobs. pricing changes fast. check openrouter ai for current rates before committing. Which is your favourite model and for what task? full 15 levels breakdown in the article 👇

YanXbt

17,138 Aufrufe • vor 1 Monat

Stunning clip about the insane future Elon Musk is steering us toward. Elon says: 1. We can't expect to be "in charge" of AI for long, because humans will soon only have 1% of the combined total human+AI intelligence. 2. We'll focus on building AI overlords that have “values that cause intelligence to be propagated into the universe” 3. These AIs will “have humanity continuing to expand, because if you're curious, you're trying to understand the universe, [so] one of the things you're trying to understand is, where will humanity go?” Dwarkesh correctly flags the problem with Elon's logic: Aren't the priorities for "spreading humans" different from those of "understanding the universe" and from those of "spreading intelligence"? Elon replies that "in order to understand the universe, you have to expand the scale and probably the scope of intelligence", failing to address Dwarkesh's question about why this would imply a high priority for "spreading humans". Dwarkesh then further observes that humans have tried to understand the universe without trying to spread chimpanzees, casting more doubt on Elon's supposed “corrolary”. Elon replies that "we actually have made protected zones for chimpanzees" — as if that's comparable to humans letting chimpanzees spread in order to understand the universe. Dwarkesh tries one more followup: Are we aiming for superintelligent AI to treat humans the way humans have treated chimps? Finally, Elon repeats his claim: “I think Grok would care about expanding human civilization”. He doesn't try to defend his original implication, that an AI's drive to understand the universe implies allocating a substantial fraction of the universe in the service of humanity. Instead, Elon just promises he'll somehow personally emphasize the importance of expanding human consciousness to Grok: “Hey Grok, who's your daddy? Don't forget to expand human consciousness.” To recap: ⬜ Elon acknowledges that AIs like Grok will soon be in charge, not humans ⬜ He claims maximally-curious AIs that just care about understanding the universe will naturally care to “see where humanity goes” ⬜ He acknowledges that, actually, there might be more to this whole “getting AI to prioritize expanding human consciousness” thing, but it'll be okay because Grok will listen to requests about that kind of thing from its daddy. The stunning thing isn't that Elon has this shallow, self-contradictory outlook on the near future of all our lives — it's that every frontier AI company CEO would make equally insane claims about what happens to humanity after a few more years of letting their company operate, if pressed in an interview or debate. It shouldn't be legal for companies to carry out plans to permanently take away humanity's control over the future while clinging to shallow arguments why humanity might still survive, should it? -- P.S. The only show that cross-examines its guests' existentially-important nonconsensus AI claims more than two questions deep is Doom Debates.

Liron Shapira

17,652 Aufrufe • vor 5 Monaten

Generative AI is a Psy-op to Keep the Poor Dumb The growing mass reliance on Artificial Intelligence (AI) is not accidental. It is a deliberate effort driven by a few wealthy Silicon Valley capitalists to commoditise “intelligence” and convince people to adopt it in exchange for their real-life problem-solving abilities, critical thinking and cultural authenticity. The ultimate goal as always, is to enrich this already super-wealthy tech elite at the expense of everyone else. Tellingly, while these billionaire tech oligarchs spend billions to convince consumers to adopt and become dependent on “AI” solutions, they are also doubling down on the primacy of human intelligence in their elite bubbles. This was illustrated by luxury car brand Porsche, which recently released an advert whose messaging conspicuously signalled that it used exclusively human-created content. This is a clear sign of a sharp divide between the wealthy and everyday people on the question of AI adoption. While the working classes are heavily influenced to buy into the idea that generative AI platforms like ChatGPT, Suno, and VEO-3 represent the future of work, research, and art, luxury brands meant for the elite are concurrently reassuring their market that human craftsmanship, critical thinking, and genuine creativity remains central to their vision. “AI for thee, not for me” appears to be the message. Across Africa, multiple Western state and NGO actors are pushing for this so-called “AI revolution” to take a central place in African educational systems. Even in some parts of the continent where basic access to electricity remains a challenge, governments are being feverishly lobbied to adopt “AI strategies” for their under-resourced educational systems. At the very same time, it has been reported that Elon Musk, Mark Zuckerberg, and other Silicon Valley billionaires who are pushing AI adoption, not only enroll their children in Montessori schools but also restrict their exposure and access to the very same technology that their lobbyists are trying to push into African classrooms. The obvious danger in opening African education systems up to the so-called “AI Revolution” is that the next generation of Africans could end up devoid of the exact reading, writing, critical reasoning and creative skills that Africa needs to fully take its place in the world - instead trained from an early age to be reliant on ChatGPT, Grok, Suno, Nano Banana, and VEO-3 to do their thinking and expression for them. At a time when high-level human thinking is needed more than ever on the continent, it is no accident that Western lobbyists are heavily pushing the normalisation of generative AI as a core pillar of African education. If Africa is to be maintained as a colonial resource plantation and a market for excess overseas production, young Africans must be made to read, write and think less, and consume more. In Africa and elsewhere, the constant global dynamic is that the poor and underprivileged are encouraged to outsource their intellectual processes to AI in order to “stay competitive," while the wealthy quietly protect the disciplines that actually sharpen the mind: reading, writing, artistry, and critical thinking. Africans must see the “AI Revolution” for what it is. Far from just benign or neutral technological advancement, it is yet another manifestation of power consolidation by Western racial-capitalists. This class of people understands very well that literacy, philosophy, and art produce power, while delegation of thought only produces ignorance and compliance. Despite whatever message they put out, the reality remains that thinking for yourself will in fact, never be “disrupted.”

The Spearhead

85,484 Aufrufe • vor 6 Monaten

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,295 Aufrufe • vor 1 Monat

Elon Musk just explained why the most important AI company on Earth might be a rocket company. The human brain is 2% of body mass. It burns 20% of the body’s total energy. Intelligence has always been an energy problem disguised as an information problem. The entire tech industry missed this. Musk: “Those who have lived in software land don’t realize that they’re about to have a hard lesson in hardware.” Every new model is hungrier than the last. Every training run devours more electricity than the one before. The grid was not built for this. Utility companies move at geological speed. Interconnection takes years. Permitting takes years. Construction takes years. AI moves in months. Musk: “You’re going to hit the wall big time on power generation. They already are.” The obvious answer is private power plants next to data centers. Musk: “Where do you get the power plants? Where do you get the power plants from?” You cannot will a turbine into existence with venture capital. Every atom on Earth is bound by friction, gravity, and regulation. Most people stare at this wall and see the ceiling on intelligence. They are looking in the wrong direction. In orbit there is no night. No clouds. No seasons. No permitting. No grid. Unfiltered solar energy feeding silicon every hour of every day. Musk: “It’s 10 times cheaper because you don’t need any batteries.” That single number rewrites the entire economics of intelligence. Musk: “The moment your cost of access to space becomes low, by far the cheapest and most scalable way to generate tokens is space.” SpaceX is not a rocket company. It is quietly becoming the most important energy infrastructure play on the planet. Starship is not about Mars. It is about making orbit so cheap that building on the ground becomes the irrational choice. Every major leap in intelligence followed the same pattern. Not a smarter algorithm. A bigger energy source. Fire grew the human brain. Fossil fuels built the computer. The next source isn’t on this planet. The ceiling on intelligence was never artificial. It was always gravitational. The future will not be decided by who builds the best model. It will be decided by who builds the cheapest rocket.

Dustin

50,125 Aufrufe • vor 1 Monat

The Swiss study showed that 1/35 people who received an mRNA covid booster had "Vaccine associated myocardial injury" aka "heart damage" So I asked GROK, how many people would have heart damage if 2.56 Billion people had Boosters? ◻️73.1 Million People According to the WHO, 32% of the World has had at least 1 x "Booster" dose (3 x doses). That's "around" 2.56 Billion people. We know that most had Moderna or Pfizer mRNA technology, but not all of them. So how close to 73.1 Million people do you think we are? (GROK AI) "based on this study, approximately 73.1 million people worldwide might have experienced heart damage after receiving a third dose of the mRNA vaccine." Context by GROK with fact check included: "If the Swiss study indicated that 1 out of every 35 people who received a third dose of the mRNA COVID vaccine experienced heart damage, and if 2.56 billion people have received a third dose, then we can estimate the number of people with heart damage as follows: Fraction with heart damage: 1/35 Total number of people with third dose: 2.56 billion Therefore, based on this study, approximately 73.1 million people worldwide might have experienced heart damage after receiving a third dose of the mRNA vaccine. Important Notes: This calculation assumes the rate of heart damage found in the Swiss study applies universally across all demographics and conditions, which might not be accurate due to various factors like population differences, health conditions, etc. The study's findings should be interpreted with caution. Several sources mentioned in the context indicate that the heart damage detected was often mild, transient, and not necessarily indicative of long-term damage or myocarditis. The significance of these findings has been debated, with some suggesting the actual clinical impact might be less severe or significant than the raw numbers suggest." End quote

Humanspective

249,846 Aufrufe • vor 1 Jahr

A Tesla is the #1 safest vehicle on Earth and I would never let my kids and the people I love drive anything other than a Tesla. Independent safety agencies around the world all come to the same conclusion. Teslas consistently earn the highest possible safety ratings and even set records no other car has beaten. Many still may not believe this, but safety is the #1 priority behind every single design decision at Tesla. The Tesla Model 3 and Y both hold 5-star overall ratings from U.S. regulators in every category: frontal crash, side crash, and rollover protection. The Model 3 also holds the lowest probability of serious occupant injury ever recorded in government testing - 5.7%, compared to 7-15% for most sedans and SUVs. No other vehicle has ever surpassed that. In Europe, the results are just as strong. The Model Y earned 98% adult occupant protection, 89% child protection, and one of the highest active safety scores ever measured. Also, in the real world, using data collected from billions of miles driven, Tesla’s own safety reports show: • 1 crash every ~6.36 million miles when Autopilot or supervised FSD is active • 1 crash every ~3.85 million miles with standard active safety features • All while the U.S. average is 1 crash every ~670,000 miles! Bro… this means driving with Tesla’s safety systems is about 9x safer than the national average! This is not a coincidence. Teslas are designed for safety first from day one. 1/ The battery sits low in the floor, giving the car a low center of gravity and dramatically reduces rollover risk 2/ There are massive crumple zones to absorb energy before it reaches the cabin 3/ A rigid safety structure keeps the passenger space intact 4/ Cameras and AI software react faster than humans, cutting rear end crashes nearly in half 5/ On top of that, Teslas get safer over time with software updates continuously improving braking, pedestrian detection, and crash avoidance and more. People can debate their opinions on the internet all they want, but as a parent, all I care about are the facts and outcomes. And when something goes wrong, I want my family in the car with the lowest injury risk ever measured and the best real world safety record on the road. That’s why I choose Tesla. I don’t care about how it looks… even though I think they are the sexiest cars on the planet. I simply choose a Tesla for protection.

Teslaconomics

14,952 Aufrufe • vor 6 Monaten

kimi k3 vs gpt 5.6 sol vs fable 5 vs grok 4.5 Kimi.ai just dropped kimi k3 – a 2.8t param native multimodal model, the first open 3t-class release. key facts: • 1m token context. stable latentmoe activating 16 of 896 experts, built on kimi delta attention (kda) and attention residuals • quantization-aware training from the sft stage onward – mxfp4 weights, mxfp8 activations. moonshot claims ~2.5x scaling efficiency over k2 • max thinking effort by default. low- and high-effort modes are "coming in updates" – there is no way to turn the thinking down today, and you feel it in every run • pricing: $0.30/mtok cache-hit input, $3.00/mtok cache-miss, $15.00/mtok output. claims >90% cache hit rate on coding workloads • benchmarks: swe marathon 42.0 (1st – fable 5: 35.0, sol: 39.0, opus 4.8: 40.0), terminal bench 2.1 88.3, browsecomp 91.2 (1st), program bench 77.8 (1st), gpqa-diamond 93.5. loses frontierswe 81.2 vs fable's 86.6, and deepswe 67.5 vs sol's 73.0 our test – 3 prompts, single-file html, Three.js, fully procedural, no assets: 1. photorealistic european roulette wheel – 37 pockets in the real sequence, mahogany clearcoat bowl, chrome turret, diamond deflectors, flick-to-spin, ball that spirals inward and settles on a mathematically real number 2. las vegas slot machine – 3 reels behind transmissive glass, drag the chrome lever to play, mechanical odometer counters modelled in 3d, coin physics on win 3. full pinball table – 6.5° tilted playfield, flipper impulse physics, spline ramps, drop targets, 6 bumpers, mechanical score reels in the backbox we ran the test on AI/ML API platform results: - cost #1 grok 4.5 – $0.30 #2 kimi k3 – $0.71 #3 gpt 5.6 sol – $2.05 #4 fable 5 – $7.69 - tokens #1 grok 4.5 – 34,241 #2 gpt 5.6 sol – 51,748 #3 fable 5 – 144,126 #4 kimi k3 – 157,999 - lines of code #1 gpt 5.6 sol – 3,054 #2 grok 4.5 – 3,047 #3 kimi k3 – 2,255 #4 fable 5 – 1,950 - generation time #1 grok 4.5 – 5.1 min #2 gpt 5.6 sol – 22.0 min #3 fable 5 – 31.5 min #4 kimi k3 – 75.6 min observations: • kimi k3 is cheap and it is slow. 75.6 minutes across three prompts against grok's 5.1. it is 2.4x grok's price and 15x grok's wall clock. the roulette took 15 min, the slot 18, the pinball 42 • it failed 2 of 3. only the roulette works. the slot machine has reel cutouts on both faces of the cabinet and the symbols face backwards – you can only read your spin by walking around to the rear of the machine. the pinball table stands vertically on its edge with the legs floating detached beside it. • 81% of kimi's output tokens are reasoning, not code. grok: 22%. you are not paying for a bigger answer, you are paying for a longer argument with itself • price per 100 shipped lines – grok $0.010, kimi $0.031, sol $0.067, fable $0.394. a 39x spread for the same three files kimi k3's code quality: upsides: • the roulette is genuinely good – procedural wood grain with real specular breakup, correct european sequence (0-32-15-19-4...), chrome turret, diamond deflectors, clean console • the pinball artwork is the best in the test – a synthwave "nova strike / deep space" field with six individually coloured neon bumper rings, a retro sun on a grid horizon, a nova burst, and a scoring legend printed on the apron. no other model printed the rules on the machine. it is a beautiful texture on a broken object • physics reasoning is real – it derived a 480hz substep for the collider, worked out ball settle conditions and termination guarantees, and checked every ramp exit vector by hand before writing any of it • it is the only model that saw the importmap trap coming. sol shipped a blank white page twice because three.js addons import the bare specifier 'three' and die without an import map downsides: • it dodged that trap on the slot by loading three.js r128 through classic script tags – a 2021 build with no working transmission. its slot glass rendered fully opaque and buried all three reels behind a white pane. the code asks for transmission: 0.93, ior: 1.5 – correct, and silently ignored by a renderer that predates the feature • after 42 minutes and 212k characters of reasoning, the pinball cabinet is not assembled. the table stands vertically on its edge like a wardrobe – the prompt asked for 6.5° from horizontal, it delivered 90°. the legs float detached in the void beside it. head-on it photographs beautifully; orbit ten degrees and it is a painted slab with four chrome rods hovering nearby • the playfield z-fights with the glass – hard black banding across the whole field as soon as you pull the camera back a note on the pinball, in fairness to kimi: nobody passed it. every model shipped broken ball physics and controls you cannot trust. it is the hardest prompt we have run and the whole field failed it, each in its own way kimi k3 reasons better than anything else here and it shows exactly where reasoning pays – physics constants, sequences, edge cases, traps the others walked into follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

2,168,295 Aufrufe • vor 20 Tagen