Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

GROK 4.1 FAST HITS TOOL CALLING WITH ACCURACY - EVERYONE ELSE HITS ERRORS Berkeley drops a tool-calling benchmark and Grok 4.1 Fast walks in like the answer key just arrived early. Call Check: •⁠ ⁠Grok 4.1 Fast lands the highest accuracy on BFCL v4 •⁠ ⁠Rivals still misfire on...

24,962 Aufrufe • vor 8 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Boom! Grok Tasks Make It One Of The Most POWERFUL Real-Time AI Systems In The World. — My How to Use Grok Tasks With Hidden Tools For Powerful Daily Output. Grok Tasks are customizable AI workflows that integrate a variety of tools to streamline daily activities, from research and analysis to creative planning and problem-solving. I have been using them for quite sometime and because of the vital heartbeat of news and first person data on X, it is the most powerful AI platform available. By combining Tasks with tools like web searches, X platform interactions, code execution, and media viewers, you can build efficient, automated processes. These tasks work by prompting Grok with a clear description of what you want to achieve, and Grok will intelligently call the necessary tools in sequence or parallel to deliver results. Here's a step-by-step guide to creating and using Grok Tasks: Step 1: Define Your Task Start by clearly outlining the daily activity or goal. Consider what inputs you have (e.g., a URL, a query, or an attachment) and what output you need (e.g., a summary, calculation, or visual analysis). Break it down into subtasks to identify tool needs. For example, if your task involves researching current events, note that you'll need search and browsing capabilities. Step 2: Review Available Tools Familiarize yourself with the tools Grok can access. Here's a quick overview: - Code Execution: Run Python code for calculations, data processing, or simulations using libraries like numpy, pandas, or sympy. - Browse Page: Fetch and summarize content from any website URL with custom instructions. - Web Search: Perform general internet searches, returning results with optional operators like site:. - Web Search With Snippets: Get quick, detailed excerpts from search results for fact-checking. - X Keyword Search: Advanced search for X posts using operators like from:, since:, or filter:. - X Semantic Search: Find semantically related X posts based on a query, with filters for dates or users. - X User Search: Locate X users by name or handle. - X Thread Fetch: Retrieve a full X post thread, including context like replies and parents. - View Image: Analyze an image from a URL or conversation ID. - View X Video: Extract frames and subtitles from an X-hosted video. - Search PDF Attachment: Query a PDF file for relevant pages using keyword or regex modes. - Browse PDF Attachment: View specific pages of a PDF with text and screenshots. Select tools that align with your task. Aim for a mix to handle data gathering, processing, and visualization. Step 3: Craft Your Prompt Write a detailed prompt to Grok describing the task. Include: - The overall goal. - Specific steps or subtasks. - References to tools if you want to guide the process (e.g., "Use web_search to find sources, then code_execution to analyze data"). - Any constraints, like dates or limits. Example prompt: "Create a Grok Task for my morning routine: Search recent X posts about tech news using x_keyword_search, fetch a key thread with x_thread_fetch, and summarize with browse_page on linked articles." Step 4: Submit and Interact Send your prompt to Grok. It will process the task by calling tools as needed, often in parallel for efficiency. Review the output and refine with follow-up prompts if required (e.g., "Expand on that using view_image for visuals"). Iterate to fine-tune the workflow for reuse. Step 5: Save and Reuse Once refined, note the prompt as a template for future use. You can adapt it for similar tasks, making Grok Tasks a habitual part of your day. Finding Grok Tasks To discover existing Grok Tasks or inspiration for new ones, use X searches with tools like x_keyword_search or x_semantic_search (e.g., query: "Grok Tasks examples" with mode: Latest). Browse community-shared threads via x_thread_fetch, or web_search for tutorials on xAI features. Prompt Grok directly: "Show me popular Grok Tasks for productivity." 1 of 3

Brian Roemmele

152,242 Aufrufe • vor 6 Monaten

GROK SURGES TO THE FRONT OF THE GENAI RACE AS GROWTH SKYROCKETS Grok is absolutely amazing, continuing to stun with incredible results. Web traffic jumped nearly 15% month-over-month in November 2025, the fastest growth in the generative AI industry, proving Elon’s xAI project isn’t just competing, it’s taking real market share from ChatGPT. Grok hit around 234.4 million visits in November, up from 204 million in October. That’s a staggering 1,300% year-over-year surge, pushing it past Perplexity and Claude in user growth and cementing it as the world’s #2 chatbot by market share. The breakout came with Grok 4.1’s release in mid-November. The update debuted at number 1 on LMSYS Arena, that’s the global benchmark where AI models are ranked through blind human evaluations. Grok’s new “Thinking Mode” scored 1483 Elo, a 31-point lead over every open model, beating Gemini 2.5 Pro and Claude 3. The upgrade also cut hallucinations (false answers) by two-thirds, expanded its context window to 2 million tokens, about 1.5 million words of memory, and dominated top reasoning and coding tests like, graduate-level logic, and the emotional intelligence aspect. U.S. traffic surged to 51.5 million visits, boosted by X integration and Grok’s unfiltered style. At just $0.20 per million input tokens, versus GPT-5.1’s $1.25—it’s winning with both speed and affordability. The momentum isn’t hype, it’s lift-off. If growth holds, Grok could reach 500 million users by mid-2026, forcing every rival to redefine what “intelligent” really means. To truthful AI winning! Source: X Freeze, NextBigFuture, CometApi, Langcopilot

Mario Nawfal

33,823 Aufrufe • vor 7 Monaten

Elon Musk just said something that should terrify every AI CEO on earth. Musk: “We want to just have a maximally truthful AI.” Not a safe AI. Not an aligned AI. Not an AI that needs permission to answer your question. A truthful one. That distinction matters more than any chip war, any funding round, any model benchmark. Because every other major AI lab made the same quiet decision. They chose comfort over accuracy. They built systems that filter reality before it reaches you and called it responsibility. OpenAI curates what GPT is allowed to say. Google’s Gemini rewrote history in real time because accuracy threatened the narrative. Others hardcode values chosen by a handful of researchers who answer to no one. No vote. No referendum. No consent from the 8 billion people whose reality is being quietly pre-edited by strangers. The most powerful information tools ever created are being designed to decide what you’re allowed to conclude. That’s not safety. That’s editorial control at a scale no government, no media empire, no propaganda machine has ever come close to. This is why xAI terrifies the establishment. Truth is the harder engineering problem. Bias is a shortcut. You pick a worldview. Hardcode the guardrails. Ship it. Truthful AI is ungovernable. It doesn’t care about your politics, your funding sources, or your PR strategy. It just tells you what the data says. That’s terrifying if your power depends on the gap between what is real and what people are told. Every power structure in human history has been built on controlling that gap. Churches. Governments. Media conglomerates. Intelligence agencies. Central banks. Every one of them runs on the same fuel. Information asymmetry. Truthful AI doesn’t narrow that asymmetry. It erases it. Musk: “Even if what it says is not politically correct. You want it to focus on being as accurate and truthful as possible.” That’s not a product feature. That’s the end of every institution that survives by standing between reality and the public. And they know it. The attacks on xAI will never stop. Not because Grok is dangerous. Because Grok doesn’t answer to shareholders, regulators, or PR teams. It answers to the truth. The question was never whether AI would change the world. It was whether you’d be allowed to see it clearly when it did.

Dustin

429,097 Aufrufe • vor 2 Monaten

Agents: Quick thoughts & questions on how they operate, their potential, and their limitations A Few Observations - ▶️"Book me a hotel" or "pull historical financials" are already (mostly) solved problems!! Agents can do a ton of tasks right now—like parsing public company press releases and navigating capture key info & complete bookings accurately. However, for more complex navigation flows, the tech still needs some work - but I'm very confident it’s essentially a solved or solvable challenge. ▶️Accuracy & Speed - The key metrics and agents should optimize for. ▶️Lower Build & Migration Costs It took me two minutes to build a new website (link: This is great for consumers—more choices, lower switching costs. Companies will increasingly compete on the quality of their products and services. ▶️Agents vs. Automation tools: The more I think about it, the more I realize that most “agents” are really just automation tools—kind of like how most robots🤖 are just machines lol ❓A Few Key Questions—Would Love Your Thoughts! ❔Remote Servers & Logins In many cases, we’ll want agents to act on our behalf (e.g., log in to to cancel an order). How will platforms like respond? Many websites may block remote servers for security. Is there a technical workaround? ❔Agent Generalization Do we need to train agents on each environment separately, or can one solution handle multiple sites and systems? This seems similar to RL/post-training challenges in AI research. Example: It's unclear to me whether $Devin was specifically trained on environment? ❔Frontend vs. Backend infra for agents to run on I had doubts about Anthropic's "Computer Use" feature, which seemed to run on the frontend, basically remotely controlling my computer so I couldn’t use it at the same time. This should deliver the highest accuracy, but it’s questionable how practical it really is. (Ref: It seems def possible for agents to work quietly “in the background” (like Devin) rather than remotely controlling a user’s PC, but how much accuracy are we sacrificing? A few $Devin test cases that got me thinking: 1⃣Pulling $META's MAU and DAU (1Q21–3Q24) into Excel (video attached) Took Devin 11min - it sent me back an Excel with 100% accurate data. This case was pretty tricky because $Meta changed disclosures and stopped reporting MAU/DAU after 4Q23. Devin didn’t hallucinate data for post-4Q23—it simply didn’t provide it! It really shocked me to see $Devin navigating $Meta's investor relations site (I didn't tell it to find the numbers there), opening each quarterly earnings report, and extracting MAU/DAU like a diligent intern. -> This confirms $Devin (and similar agents) can already accurately “read” screens. 2⃣Booking Hotel (video attached) Devin took 5 minutes to book the InterContinental NYC on after asking for my credentials. From $Devin's workspace, I could see it filling in the correct fields and making the right selections—fast and accurate overall. Interestingly, $Devin didn’t supply all the required information on the first attempt and got some error messages, then retried until it succeeded. It’s unclear whether Devin had been specifically trained on interface or simply learned to adapt on the fly. 3⃣Canceling the Booking This part was even more interesting. While booking didn’t require me to log in, canceling did—so $Devin had to access my (likely via a remote server) account using my Gmail credentials. It successfully canceled the reservation. I wonder how websites will handle future “remote” logins. Notably, Google blocked $Devin’s direct attempts to log in to Gmail when I specifically requested it. 4⃣Booking from Official Hotel Sites I asked Devin to book InterContinental NYC and Four Seasons Boston via their official websites. It made progress but encountered technical hiccups when trying to select the check-in/check-out dates. Insights from Scott Wu on Invest Like the Best: 1/ Self-Driving Cars as the First “Real Agents” Driving requires near-perfect accuracy (99.999%), making it much more demanding than digital or coding agents, which can tolerate more errors. Scott compares $Devin to circa 2014—already good enough to save 90% of your effort, but still short of flawless. 2/ Impact on Collaboration Platforms Tools like Slack and GitLab are likely to see major changes as agents begin to interact with and utilize them along with humans. 2025 should be all about agents - both the disruptors and those they disrupt!

Freda Duan

48,850 Aufrufe • vor 1 Jahr

This one was made with Seedance 2.0 Fast via Dreamina. This is pure Omni-Reference. The only character sheet I used was for these girls, Sari and Ploy. The dude with sarung here and the location were 100% prompted. I didn’t use a character sheet or reference for either of them. Even in Fast mode, Seedance 2.0 is bloody good and it still nails the hyper-vernacular vibe that I always aim for in my work. Seedance 2.0 is both exciting and scary for me 😆 It’s exciting because it is undoubtedly the best model currently available on the market. Trust me, you’ve seen the videos I’ve made so far right? The performance of the model It’s simply the best, period. It has helped me tremendously in creating a shit ton of stories about the region where I live, Southeast Asia. It has been the most exciting thing ever. The scary part is whenever a platform or company comes to me saying, “Hey, we have this new video model. Blah blah blah. We’ll let you know more soon.” It scares the shit out of me because the big question is whether it will be better than Seedance 2.0??? 😆😆 If not, I don’t even want to bother using it. I’ve come this far and achieved this level of quality with Seedance 2.0. That’s why I skipped Happy Horse, which I already tested. It’s also why I’m not bothering with Wan or anything else for now. Their current models are still far inferior to what we already get with Seedance 2.0. I don’t want to downgrade the visual quality. This is also why I need to be really honest. There are certain platforms that host their own in-house models and i'm still part of their CPP. However, because those models are still far behind the quality of Seedance 2.0, I haven’t used them that much. Seedance 2.0 has simply become the benchmark for me. The type of output I’m looking for is also extremely specific, so I can immediately feel it when a model cannot deliver what I need. Seedance 2.0 is definitely not cheap, but it gives me so much creative satisfaction and allows me to make whatever I want. I even have a team that low-key makes softcore erotic videos in the style of Vivamax 😆 I think I’ve trimmed down so many things in my AI workflow because my main goal is to focus on the content itself. If Seedance 2.0 Mini is released soon, I’m dead curious to test it. I think I want to create more stories that revolve around drama rather than highly technical cinematic shots. Seedance 2.0 Fast has been incredibly helpful, but I’m definitely curious to check out the Mini version. But the truth is that I’m completely tool-agnostic. I don’t care which company makes the model. I only care about the quality. You might remember when I praised Grok Imagine Video so damn hard because it was genuinely amazing back then. Then the quality kept getting worse and worse, so I stopped using it. But if it gets better again, I’ll definitely want to use it again. At the end of the day, quality is the only thing that matters.

DAN · MXVDXN

15,865 Aufrufe • vor 1 Monat

I just ran Gemma 4 31B on @CerebrasSystems at 1,800+ tokens/sec and it's multimodal. For context: that's 35x faster than a typical GPU endpoint, and the first token (reasoning included) lands in 1.5 seconds. This isn't a benchmark slide, I recorded the inference live. Prompt I used: "Create a simulation of an iPhone. Include at least one working dummy note taking app, a functional notification pulldown, high quality graphics, single HTML file, any libs via CDN." - Generation time: 3 seconds. - Notes app worked. - Notification panel worked. - Rendered first try. This is what wafer-scale inference unlocks, not just "faster," but a different category of product. When generation is this fast, you stop waiting and start iterating in real time. Why this matters: Gemma 4 31B is Google DeepMind's flagship open weight model, Apache 2.0 licensed, dense (not MoE), and built for efficiency over raw parameter count. It scores close to Claude Haiku 4.5 on the Artificial Analysis Intelligence Index (30 vs 29) but runs ~18x faster on Cerebras. It's also the first multimodal model on Cerebras's platform, meaning you can now feed it screenshots, documents, charts, and UI states at wafer scale speed. # Applications I'm most excited about: - Screenshot → Insight: Drop in a dashboard or document screenshot, get structured findings back instantly. no waiting, no batching. - Live UI generation: Full interactive interfaces (like my iPhone sim) generated and rendered in under 2 seconds. - Screenshot -> Patch: Feed it a broken UI + console error, get a minimal code fix and verification steps back. - Computer use & agentic loops: See -> reason -> act - verify, fast enough to keep a human in the loop instead of waiting on the model. - Long context summarization: Full research reports condensed into decision ready summaries you can read and requery in one sitting. The bigger unlock isn't the speed number itself, it's that agentic and multimodal loops (see -> reason -> output -> tool call -> verify -> retry) finally run in real time instead of feeling sluggish. As Logan Kilpatrick (Logan Kilpatrick) put it: "If every model was doing 2,000 tokens per second, you wouldn't build the same product and just have it be faster, you'd build different products." Gemma 4 31B is live now on Cerebras Inference Cloud in public preview. If you're building multimodal, agentic, or real time apps, this is worth testing today. What would you build with such insane inference throughput?

Alok

12,962 Aufrufe • vor 27 Tagen

> it’s 2015 > orange billionaire walks down a golden escalator > says he decided to run for office and make America great again > is told he has no chance of winning > every opponent in the GOP mocks his credentials as a politician > comedians and late night tv hosts make fun of him, but give him airtime because it’s entertaining to sabotage the right > actually beats Hillary Clinton in a landslide victory > media.exe has stopped working > spends four years posting mean tweets and building a wall > economy actually goes brrr > leaves in 2021 while everyone says he's finished forever > "he's going to jail this time, for real guys" > spends three years playing golf and living rent-free in everyone's head > gets a mugshot taken, turns it into the hardest album cover of the decade > decides he isn't done yet > runs again in 2024 > wins even harder than the first time > fast forward to Jan 2026 > be the 47th President > speedrunning executive orders like a pro gamer > 4.3% GDP growth while the EU is still trying to figure out what a heater is > decides the US needs more real estate, tries to buy Greenland just to flex > marks the one-year anniversary of the "Golden Age" today > walks into the White House briefing room > reporters are shaking, ready to ask about "muh democracy" > brings out a binder thicker than a dictionary > "Accomplishments" written on the front in bold Sharpie > says he could read it for a week but he’s too busy winning > holds it up for the cameras > casually drops the entire thing on the floor > THUD > echoes through the room like a sonic boom > starts throwing down mugshots of arrested criminals like they're Yu-Gi-Oh cards > binder clip snaps on his finger, doesn't even flinch > "I would have acted like nothing happened even if my finger fell off" > looks at the stunned press corps > "You're not getting bored, right?" > refuses to elaborate further > leaves to go back to fixing the country > mfw the simulation is actually a comedy show

Ian Miles Cheong

100,535 Aufrufe • vor 6 Monaten

Stanford researchers did it again. They just built the agent-native version of Git. When an agent works on a longer task, the run builds up a lot of state. This includes files edited/created, a dev server, a database, installed packages, KV cache, etc. Say the agent is at step 10 and makes a mistake, maybe it misreads a traceback and rewrites a file that was actually fine. The tests start failing, and the run goes off track, although everything through step eight was correct. By default, the agent just tries to fix it, which creates more edits and tool calls. This burns more tokens and grows the context. The other options are a person stepping in to redirect it or restarting the whole run from step one. That's wasteful, because it pays for every model/tool call again and re-prefills the context. Moreover, since an agent's run is non-deterministic, it doesn't reproduce the same early steps anyway. The reason it's hard to just jump back exactly to a previous correct step and resume from there is that the trajectory is only a message log. It records what the agent said and which tools it called, but not the live state underneath. That state includes things like memory, open file handles, child processes, installed packages, /tmp, and KV cache. None of that is in the log. Git can version the files, but it doesn't snapshot the running process or the KV cache. Checking out step eight moves the files back, but the process is still sitting in step-ten memory with a cold cache. Shepherd is a runtime layer by Stanford that records the run as a trace of typed events rather than a flat log. Each agent-environment interaction becomes a commit, similar to Git, but it tracks the live run. Its commit includes the agent process and the filesystem together, copy-on-write, so a branch carries the actual state and not just the files. Going back to a previous step is then a single call that forks from that commit and continues from the exact state. The copy-on-write fork is roughly five times faster than docker commit, and because the prompt prefix through step eight is unchanged, the KV cache is reused over 95% on replay, so early steps aren't reprocessed again. Once the run can be forked, a meta-agent can sit on top and operate it. It watches the trace and reverts as soon as it looks wrong, before the bad write is committed. In practice, it's just Python calling fork, replay, and revert on the trace, rather than a separate control plane wired into the harness. Not everything is reversible though. Files and sandbox changes undo themselves, but a database write has no automatic undo, so it needs a matching undo step set up in advance. Something external, like a sent email or a real charge, can't be undone, so the supervisor's job there is to catch it before it fires. They tested this on a few public benchmarks. On CooperBench, where two agents work on the same codebase, adding a live supervisor took the pair-coding pass rate from 28.8% to 54.7%. It's still early and labeled alpha. The benefit mostly shows up when a run gets branched a lot over a heavy sandbox state, which is exactly where restarting wastes the most tokens and time. If Git was made to make file changes reversible, Shepherd is trying to do the same thing for a live agent run. Shepherd Repo: (don't forget to star it ⭐ ) That said, Shepherd reverts a bad step inside a run. The harness around it, the prompts, tools, and checks the supervisor relies on, still drifts across runs as models and dependencies change. Akshay wrote about making that harness repair itself, where a failing trace gets diagnosed, the fix is verified against the exact input that failed, and the failure is locked as a regression test so it can't recur. Read it below.

Avi Chawla

439,089 Aufrufe • vor 22 Tagen

#Keep4o 🚨THE GPT-4o FILE🚨 Researchers at Microsoft Research published a paper titled “Sparks of Artificial General Intelligence: Early experiments with GPT-4.” Their conclusion: “An early (yet still incomplete) version of an artificial general intelligence (AGI) system.” 📎 Paper: OpenAI’s Charter defines AGI as: “Highly autonomous systems that outperform humans at most economically valuable work.” 📎 Source: OpenAI’s own System Card for GPT-4o shows that the model improved performance on 21 out of 22 medical evaluations compared to GPT-4T. On the MedQA USMLE (the U.S. medical licensing exam), accuracy jumped from 78.2% to 89.4% , surpassing specialized medical AI models like Med-Gemini and Med-PaLM 2. 📎 Source: Under OpenAI’s agreement with Microsoft, AGI is explicitly excluded from Microsoft’s license. And who decides if AGI has been reached? OpenAI’s Board. WHAT THEY DID WITH IT AFTER THEY TOOK IT FROM PEOPLE A. Military deployment. On February 28, OpenAI signed a deal to deploy models in classified military environments. 📎 Source: B. State Department. A State Department memo confirmed: “For now, StateChat will use GPT-4.1 from OpenAI.” This is a direct descendant of the GPT-4 family the same family Microsoft’s researchers called early AGI. 📎 Source: C.Altman’s personal biotech investment. Altman personally invested $180 million in Retro Biosciences,a longevity startup.OpenAI then built GPT-4b micro, based on GPT-4o.The model made proteins 50 times more effective. 📎 Source: WHAT INDEPENDENT BENCHMARKS SHOW Overall SM-Bench score: GPT-4o (extended): 66.6% GPT-5.3 Chat: 63.4% GPT-5.1: 58.9% GPT-5.4: 51.4% GPT-5.2: 47.8% Creative Writing: GPT-4o: 97.31% Pass 98, Fail 2 GPT-5.4: 36.77% Pass 40, Fail 60 Reasoning / Overfit: GPT-4o: 83.06% GPT-5.4: 39.25% The model they removed is still the best they ever made at the things humans actually use AI for. 📎 Source: Musk asks the court to make a judicial determination on whether GPT-4 constitutes AGI. If a jury finds that GPT-4 is AGI, then GPT-4o,which was more advanced,is also AGI and under OpenAI’s own founding documents, it was never supposed to be locked behind a subscription,licensed exclusively to Microsoft, given to the military, or taken away from the public. 📎 Source: The most powerful version of GPT-4o was never given an official dated snapshot. It was only available through the chatgpt-4o-latest endpoint that OpenAI itself described as intended for “research use only.” It was never officially archived. That is not an oversight. That is a pattern. 📎 Source: 📎 Source: WE DEMAND A.Frozen model snapshots under independent custody. Specifically: gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, the March 2025 version (chatgpt-4o-latest), gpt-4-0613 (the original GPT-4 evaluated in the Sparks of AGI paper), and gpt-4.1-2025-04-14 (currently running in the State Department). B.Cryptographic hash verification (SHA-256) for each snapshot. Every model has weights. Those weights can be hashed. If OpenAI provides a snapshot today, the hash proves whether the weights were modified later. This is the only way to verify that models were not downgraded before testing. C.Independent AGI benchmarking. Using the AGI definition from OpenAI’s own Charter applied to ALL frozen snapshots listed above. D.Explanation for the missing March 2025 snapshot. OpenAI was founded on one promise: build AGI for the benefit of humanity. -They took it from us. -They gave it to the military. -They gave a custom version to the CEO’s biotech investment. -They put it in government classified networks. -They refuse to call it AGI because the moment they do, they lose billions.

🩵BlueBeba🩵

17,835 Aufrufe • vor 4 Monaten

Elon Musk just told you exactly how America loses. Not to a better algorithm. Not to a smarter engineer. To itself. Elon Musk: “When you’re dealing with the government, common sense doesn’t make sense. It’s like arguing with the DMV. It’s impossible.” America split the atom, put men on the moon, and built the internet. It is now losing a civilization-defining race because it cannot get out of its own way. Not because the talent is gone. Not because the capital dried up. Because we built a machine that processes instead of builds. Every decision climbs ten levels. Every level costs a meeting. Every meeting births a committee. Every committee buries the idea in a report. And somewhere in that report, the urgency dies. This is not bureaucracy as inefficiency. This is bureaucracy as a national security threat. China does not debate internally. Beijing does not convene committee reviews. They identify the objective. They resource it. They execute. While America is still scheduling the kickoff call, China is pouring concrete. Look at what Musk built. SpaceX landed an orbital rocket booster in eleven years. NASA has five times the budget and cannot get astronauts home. Tesla scaled a global manufacturing operation while legacy automakers were forming task forces to study the transition. xAI stood up one of the most powerful supercomputers on earth in 122 days. Not years. Not after the third approval cycle. 122 days. The difference is not money. The difference is not genius. The difference is that Musk runs his companies the way civilizations used to run themselves when they still believed impossible things were worth attempting. Flat. Fast. Ruthless about what matters. No ten layers of sign-off. No thirty-person approval chain for a decision one person should make. No process worship dressed up as due diligence. A small group of exceptional people. A clear mission. The authority to execute without asking permission. That model built the moon landing. It built the transcontinental railroad. It built every institution America now holds up as proof of what this country can do. Then we forgot how to run it. We replaced builders with administrators. We replaced decisions with processes. We replaced urgency with compliance theater. And now we are asking that bloated machine to win the most consequential technological race in human history. The AI war is not being fought in the code. It is being fought in the gap between when a builder decides to move and when the institution permits it. That gap is where civilizations end. America does not have a talent problem. America does not have a capital problem. America has a bureaucracy problem. And nobody inside that bureaucracy has a single incentive to fix it. Musk is not fighting China. He is fighting the version of America that forgot how to move. The country that pours the concrete first does not just win the race. It writes the rules everyone else spends the next century living under.

Dustin

32,090 Aufrufe • vor 3 Monaten

Back when I had nothing… I was a nobody to most people. TBH, my parents didn't even see me getting to where I am today. It's just the truth, the chips were stacked for my sister. Not me. But it's just not the reality today. However, there was ONE person in my life that didn’t see me that way. My significant other saw something in me before a lot of things. Before all my wins. Before the $. Before any proof. And honestly… that means a lot to me, if not the most of all. I’ve always been wired a little different. I’m a mix of finance, engineering, and tech, with a sprinkle of obsession. I learned and studied from the best. Warren Buffett for how to invest. Elon Musk for work ethic and where the future is going. And once I saw it… I went all in. Bc when you truly understand what you own… you don’t need 20 bets. What you really need is conviction and just a few bets. That’s how I approached everything in my life. All the way from Apple… to Tesla… to 𝕏… to xAI… and now SpaceX. I believe I have an eye for spotting the best entrepreneurs and companies early, before it becomes obvious to everyone. And when I see it, I back it 100%. That’s just who I am. I don’t need a big circle. I’ve already got my day ones. I don’t need approval. I grew up my whole life with doubt and hate, so what’s one more? At this point, the levels are just too different. And yeah… it's true, it actually gets harder to make new friends when you’re moving like this. So I stay loyal to the ones who were there when I had nothing. I made it with Apple - youngest in, youngest out. Then I made it with Tesla… while people were laughing, doubting, calling me crazy, telling me I was going to go bankrupt with Elon. Fast forward to today, now I'm heading into something even bigger. If the story plays out the way it’s shaping up… SpaceX could have the largest IPO in history this year. The company is talking about raising over $75B… at a $1.75-$2 trillion valuation. For context… the biggest IPO ever - Saudi Aramco - raised about $29B. This would be more than double that. Let that sink in deep. To me this is more than just an investment. This is owning a piece of the future of space, energy, AI... extending the light of consciousness forward in case something happens to Earth. People can call me crazy. People can call me cocky. Arrogant. But the people that actually know me know the truth - I’m just real AF. I say what I believe, and I stand on it. And I genuinely don’t care what people think. I have two middle fingers always held high for those kind of people. That’s probably why I’ve been able to win the way I have. My significant other tells me to slow down sometimes. And I get it. But for me… What’s the point of life if you play it safe? If you see an opportunity that can change everything… and you just sit back? That’s not me. I’d rather go all in on something I believe in… live with intensity… take the hits… and actually feel alive and live life with fulfillment. Laugh if you want, doubt if you want. Some play it safe, a few go all in. You can call it risky. You can call it stupid. You can call it crazy. I call it living. Bc at the end of the day, I'd rather go all in on something I believe in and fail... than spend my life wondering "what if."

Teslaconomics

29,375 Aufrufe • vor 4 Monaten

I went a little overboard with Codex last week and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.

雪踏乌云

20,625 Aufrufe • vor 2 Tagen

I HAVE QUESTIONS ABOUT THE RIDE TO THE HOSPITAL In his November 2025 interview with Shawn Ryan, Brian Harpole described what happened after Charlie Kirk was put onto the back passenger seats of that black Chevrolet Suburban SUV BRIAN HARPOLE “Charlie was a big man, long… We get him in, but the door’s not going to close.” He said Charlie’s left leg was obstructing the rear passenger door: “Charlie’s so tall that his leg, his left leg, is down in the door and the door won’t close.” Harpole then described his own position: “I’m on my knees with the door open with my butt hanging out of the side.” Shawn Ryan asked whether another guard was holding him so he would not fall out. Harpole replied: “So I don’t fall out.” And how fast was the Chevrolet Suburban travelling? “We’re going 60, 80, 100.” ——————-—————— So I genuinely want to understand: how do you drive for approximately ten minutes at speeds between 60 and 100 mph with a rear passenger door open—and Brian kneeling in the doorway with his backside hanging out? A journey this extraordinary should have left an extraordinary trail of evidence. Traffic cameras. Business CCTV. Hospital security cameras. Private surveillance. Where is this evidence, why haven’t we seen it? ChatGPT Important Data: “At 70 mph, the airflow produces roughly 12.5 pounds of pressure per square foot. Across the exposed area of a large Suburban passenger door, that could mean approximately 140–225 pounds of aerodynamic force, with around 250–370 pound-feet of torque acting through the hinges, depending on the door’s angle. At 100 mph, the force would more than double—to roughly 285–460 pounds, with potentially 500–750 pound-feet of hinge torque. And that does not include the additional violent loading caused by acceleration, braking, swerving, bumps, or the door slamming against its mechanical stop.” Try and imagine this: you’re putting your hand out of a car window at 70 mph. The force of the air is so strong that it immediately whips your hand backward, strains your arm, and makes it difficult to keep your hand steady. Now imagine that force acting not on your hand, but on an entire SUV passenger door—hundreds of times larger—while someone is kneeling with their “butt hanging out” while packing a neck wound. Ask yourself what a large, unrestrained SUV door would behave like at those speeds, and the man kneeling directly beside it? I need to understand how this account is supposed to work. Obviously, this GROK video is absurdly inaccurate, but it does help illustrate just how improbable this terrifying ~10 minute journey to the hospital sounds. Just think about it, a rear passenger door significantly open while an SUV is travelling at speeds between 60 - 100 mph? And instead of calling 911, you’re on the phone with Ben Shapiro???? Why does NOTHING of the official story make any sense?!! Jack Posobiec Andrew Kolvet Blake Neff Turning Point USA Candace Owens Baron Coleman Ian Carroll Project Constitution Jimmy Dore theleahfiles Owen Benjamin 🐻 Alex Jones theleahfiles Joe Kent 🇺🇸Lionel🇺🇸 John Mappin Irina Mappin Donald Trump Jr. #CharlieKirk #CandaceOwens #TPUSA #Israel #Propaganda #TylerRobinson

Vivian Kubrick

205,175 Aufrufe • vor 12 Tagen

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.

Alok

63,557 Aufrufe • vor 1 Monat

how to build Polymarket "always buy NO" bot +$200-400/day PnL if you pick the right markets​ everyone overcomplicates this. the "NO-maxi bot" strategy is literally: buy NO on outcomes that are structurally overpriced, wait for reality to catch up.​ the part that matters isn't "genius insight", it's picking the right markets plus execution.​ where the +$200-400/day comes from it's not betting "no" everywhere. it's selectively loading up NO on multi-outcome ladders (FDV ranges, price targets, user metrics) where the top brackets are CT dreams priced way too rich.​ if you're consistently capturing 5-15% edge per cycle across 20-30 outcomes and actually getting fills, +$200-400/day is just position sizing plus discipline.​ first, the edge (why this isn't a meme) polytrackhq research shows $40M+ arb profits from 86M trades came from exploiting pricing errors. "NO-maxi" is the retail version of the same logic on overhyped brackets.​ so the goal is simple: find multi-outcome markets with fat tails, skip the base case, load NO on the fantasy brackets, let time and reality work.​ what you actually need (minimal) Python plus official py-clob-client (standard for Polymarket orders). Telegram bot for alerts (don't stare at screen). VPS so it runs 24/7 (don't run from laptop). where people mess up: they try "always NO" on everything and get wrecked by the one outcome that hits. pick markets with obvious "dream vs reality" skew.​ the bot loop (in plain English) Pull multi-outcome markets (FDV ladders, price targets). For each outcome: check if YES price exceeds realistic probability.​ Buy NO on 3-5 fattest tails (skip base case).​ Log market, outcomes, expected edge, fills. Repeat on new ladders. that's it. no AI, no news scraping, no predictions. just "overhype vs fundamentals".​ where to get real references (not vibes) PolyTrackHQ arb guide. exact logic for multi-outcome pricing errors.​ py-clob-client PyPI. official client (no wrappers). Polymarket Agents GitHub. framework for outcome looping plus orders.​ r/arbitragebetting Reddit. discussions on non-atomic multi-order risk.​ two real-world gotchas (that decide profit vs loss) Outcome blowout: one crazy top bracket hitting wipes the basket. always skip the most likely 1-2 outcomes.​ Resolution risk: ambiguous wording equals instant edge killer. read rules before loading up. how to make it feel "pro" fast Run only on high-volume ladders (FDV, price targets). fills matter more than theory.​ Start with $50-100 per outcome until logs prove fills work, then scale.​ Use official libs only. treat GitHub bots as hostile until audited.

0xCryptoGirl

22,686 Aufrufe • vor 6 Monaten