正在加载视频...

视频加载失败

timelapse #85 (27.5 hrs): - currently cant rely on any other coding models except grok code fast 1 + grok 4 fast (for complex reasoning grok 4 fast is 20 cents for 1M tokens) - wrote qwen3-next trainer entirely from scratch to make it more managable - each piece...

283,820 次观看 • 11 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

In the latest Box AI Enterprise Evaluation, we tested xAI's Grok 4. We have seen how this new model moves beyond surface-level retrieval to tackle sophisticated business logic, perform precise calculations, make inferences based on qualitative patterns, and pinpoint critical contract clauses. 👉 Key Highlights: ↳ When analyzing company financial data, Grok 4 correctly performed multi-step tasks, followed sequential logic, and made accurate calculations to determine gross margins and performance comparisons. ↳ In reviewing information from text passages, Grok 4 showed advanced qualitative reasoning, comparing stylistic elements like tone, perspective, and vocabulary to correctly group passages and identify the number of authors. ↳ It excelled in extracting complex information from contracts, identifying detailed clauses like uncapped liability and profit sharing, as well as analyzing the implications of interdependent terms in agreements. 💡 Why It Matters: ↳ For Legal and Finance teams, Grok 4’s improved ability to handle calculations and interpret complex clauses makes it a powerful tool for in-depth contract review and financial analysis. ↳ For researchers, the model's advanced analytical capabilities can help deconstruct and synthesize information from dense technical papers. 👉 The Takeaway Overall, Grok 4 shows measurable advancement in sequential logic, numerical precision, and domain-specific language understanding. The model’s ability to blend quantitative and qualitative reasoning widens the range of workflows that can be automated inside Box.

Box

1,973,403 次观看 • 1 年前

Inspired by Grok as a developer and a heavy gamer for over 15 years, I spent some time last week building a few things. Thrilled to unveil my latest creation: an infinite runner game built almost entirely by Grok from xAI! This project showcases the incredible power of AI in game development. Grok handled everything—from designing the game mechanics to writing the code and even helping me debug issues along the way. I brought it to life using some amazing free assets from a treasure trove for indie developers. You can play the game now at Elon Musk, I’d be honored if you checked it out. AI is revolutionizing game development, and Grok is at the forefront with its outstanding capabilities. It’s more than a tool—it’s like a tireless co-developer. Grok grasps complex concepts, provides suggestions, and turns rough ideas into working code fast. For this infinite runner, it crafted smooth player controls, randomized obstacle generation, and an engaging scoring system, letting me focus on the overall vision. And cross_protocol, founded by Henry @CROSS is set to harness AI’s full potential in gaming, pushing the boundaries even further. This is just the beginning. With Grok’s help, I’m planning future projects: 1) Physics-based puzzle game where players tweak gravity and momentum to solve puzzles 2) 2D RPG with deep storytelling and branching dialogue 3) Fast-paced 3D shooter with immersive worlds Each genre requires unique skills, but Grok’s versatility makes it ideal for all of them. It adapts to any challenge—be it physics simulations, character AI, or level design—producing results that could rival a full dev team. AI like Grok is opening up creative doors I couldn’t tackle alone, and I can’t wait to see what’s next. Stay tuned for more.

J

99,803 次观看 • 1 年前

Introducing my newest app, TethrX, made with Grok 4.5 to access Grok Build on your phone. TethrX connects to Grok Build running on your own computer, so you can start a task from anywhere, follow Grok's reasoning as it works, approve every command before it runs, and review the code it writes. Your code never leaves your machine. The public TestFlight is now open, and a demonstration is below. You pair your phone by scanning a QR code, either on your local network or from anywhere through Tailscale. From there TethrX streams Grok's reasoning, tool calls, command output and file changes as they happen, and asks your approval before anything runs. Plan mode lets you read the plan before the work begins. When a task finishes you can review exactly what changed. TethrX lists the modified files in your project, shows a diff for each one, and lets you commit or discard the work without leaving your phone. The app supports slash commands, including /compact and any skills you have installed, along with voice dictation, queued messages and reusable prompts. Sessions can be searched and organised into folders, and you can pair several computers and switch between them. Siri can start a task or tell you what Grok is doing without opening the app, a home screen widget shows whether Grok is working, and a Live Activity tracks progress on your lock screen and Dynamic Island. Every session reports its context window, token usage and cost, the app can be locked behind Face ID, and your computer is kept awake for as long as a task is running. TethrX requires Grok Build installed and signed in on your computer, together with Node.js 20 or newer. A single command starts the bridge: npx tethrx-bridge TethrX is open source under the Apache License 2.0. Both the iOS client and the local bridge are available here: Grok 4.5 helped a lot, thanks to SpaceXAI for making Grok 4.5 exceptional.

Myrhe𝕩

13,099 次观看 • 1 个月前

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 次观看 • 20 天前

HERMES AGENT SUPPORTS 300+ MODELS. PICKING THE RIGHT ONE PER TASK IS THE DIFFERENCE BETWEEN $5/MONTH AND $50. STARTING OUT: Claude Sonnet 4.6. official recommendation from Nous Research. "the model this project was built and tested with." strong reasoning. reliable tool calling. mid-range pricing. PREMIUM TIER: Claude Opus 4.8. best coding benchmarks available. self-correcting reasoning. catches its own mistakes. 1M context. use for demanding tasks where quality matters. GPT-5.5. #1 Chatbot Arena. #1 GPQA Diamond reasoning (94.1%). #1 creative writing. 2M context. handles entire codebases in one pass. Grok 4.30. the only frontier model with live X firehose access. real-time social data, breaking news, market sentiment. connects via Grok OAuth. no separate API key. Grok-Composer-2.5-Fast (v0.17.0). Cursor's coding model. 200K context. available through your Grok subscription via OAuth. no extra cost if you already pay for Grok. MID-RANGE TIER: Claude Sonnet 4.6. best balance of quality and cost for daily use. strongest prose and tool calling in this tier. Gemini 2.5 Pro. Google Search grounding built in. cites sources. verifies claims. pulls current data. 2M context. best for research-heavy workflows. GPT-4.1. reliable tool calling. solid general reasoning. good middle ground when you need OpenAI compatibility. BUDGET TIER: Claude Haiku 4.5. fastest Anthropic model. cheapest paid Claude option. strong at classification, routing, simple queries. use for auxiliary tasks: compression, vision, web extraction, approval scoring. DeepSeek V4. best cost-to-quality ratio in the market. 90% cache discount on repeated context. use for sub-agents and bulk parallel work. DeepSeek V4 Flash. cheapest paid model worth using. 1M context. MIT license. self-hostable. use for cron jobs, monitoring, routine searches. MiniMax M3. Nous Research and MiniMax collaborating on optimization. 1M context via lightning attention. 59% SWE-Bench Pro. beats several premium models on coding. one of the most-used models inside Hermes. FREE / LOCAL: Qwen 3.5 27B via Ollama. 16GB VRAM. reliable tool calling. best free local model for Hermes as of mid-2026. Qwen 3 8B. 8GB VRAM. fits a $7 VPS. handles routine tasks at zero API cost. Llama 4 Maverick. best open-weight tool calling. 1M context. needs more VRAM but strongest local option. HOW TO ASSIGN MODELS: main model: Desktop app / Dashboard → Models → switch sub-agent model: set in Desktop app, Dashboard, or config.yaml: delegation: model: "deepseek/deepseek-v4" auxiliary models (compression, vision, web extract): Desktop app / Dashboard → Models → Auxiliary Haiku 4.5 or Gemini Flash work well here. saves significantly when your main model is premium. per-profile: each Hermes profile gets its own model. Scout on DeepSeek. Analyst on Sonnet. Briefer on budget model. Coder on Opus. per-cron-job: pin a specific model to any cron job. morning brief on Haiku. deep research on Sonnet. monitoring on DeepSeek Flash. each job uses only the model it needs. per-session: /model deepseek/deepseek-v4-flash hot-swap mid-conversation. no restart needed. FALLBACK CHAINS: if your primary model is unavailable, Hermes automatically switches to the next provider. rate limit or server error = next model in the chain. no failed runs. no manual intervention. set in Desktop app, Dashboard, or config.yaml: fallback_providers: - openrouter - nous - codex PROVIDER PATHS: OPENROUTER: 300+ models under one API key. pay per token. most flexible. NOUS PORTAL: 300+ models + Tool Gateway (web search, image gen, TTS, browser). one OAuth. one subscription. 10% off token-billed providers. CHATGPT SUB: GPT-5.5 + Grok via OAuth. included tokens with $20 subscription. OLLAMA: free. local. private. zero API cost. your hardware only. mix providers across profiles and tasks. Scout on OpenRouter. Analyst on Nous Portal. Coder on ChatGPT sub. Monitor on Ollama. THE RULE: premium for work that needs deep reasoning. mid-range for daily driver tasks. budget for volume and background work. free for monitoring and routine jobs. pricing changes fast. check openrouter ai for current rates before committing. Which is your favourite model and for what task? full 15 levels breakdown in the article 👇

YanXbt

17,138 次观看 • 2 个月前

veo 3.1 fast vs seedance 2.0 vs grok imagine vs happyhorse 1.1 four video models pulled from the openrouter video leaderboard by request count this week (skipping the duplicate google/bytedance variants to get four distinct labs): #1 veo 3.1 fast (Google DeepMind) – 45k requests #3 seedance 2.0 (bytedance) – 22k requests #5 grok imagine video (SpaceXAI) – 9k requests #8 happyhorse 1.1 (Alibaba Group) – 4k requests so we tested them. 3 prompts, text-to-video, 16:9 / 720p / 8s, real-player likeness fed in as reference images where the model allowed it. all run via AI/ML API in the run-up to the 2026 world cup final – argentina vs spain – we built three broadcast moments from that tie. each one has to be mechanically correct, not just pretty: • stadium flyover – 80k-seat bowl, argentina vs spain, one continuous descending aerial spiral, tifo + flares, golden-hour / floodlight split • penalty – lamine yamal (spain #19) vs emiliano martínez (argentina keeper): run-up, single strike, full-stretch dive, ball in the net. real faces via reference • free kick – messi 25m out, five-man wall, curl up and over the wall into the top corner. the wall has to face the ball with arms pinned down, like a real wall the takeaway up front: the gap that decides this isn't quality – it's moderation. three of the four refuse to render real footballers' faces (grok was the only one that took every reference), so most of the test had to be reshot "faceless" – camera behind the player. the price spread on top of that is ~4x overall results: cost #1 grok – $1.56 #2 veo 3.1 fast – $3.12 #3 happyhorse – $4.38 #4 seedance 2.0 – $6.00 generation time #1 grok – 4m 16s #2 veo 3.1 fast – 4m 26s #3 happyhorse – 8m 16s #4 seedance 2.0 – 10m 38s avg bitrate (picture density) #1 grok – 12.0 mbps #2 veo 3.1 fast – 11.1 mbps #3 happyhorse – 7.4 mbps #4 seedance 2.0 – 5.3 mbps real faces allowed ✅ grok – took every reference ❌ happyhorse – yamal ok, messi blocked ❌ veo – blocked ❌ seedance – blocked observations: 1. moderation is the whole story. three of the four blocked at least one real face – veo and seedance refused every reference outright, happyhorse took yamal but rejected messi. only grok rendered all of them. everything else had to be shot from behind so no face shows 2. grok is the outlier: cheapest, densest picture, fastest, and the only one that renders real faces. it won on every axis that mattered here 3. seedance is the anti-grok – 4x the cost, 2.5x the time, half the bitrate, and no real faces. worst value in the set 4. none of them understand football out of the box follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,453 次观看 • 1 个月前