正在加载视频...

视频加载失败

Video diffusion models are just overqualified depth estimators! Deterministic single-pass depth estimation based on WanV2.1. - SOTA 5.5 AbsRel on ScanNet - data-efficient than baselines; - no temporal flicker + infinite-length estimation w/ zero scale drift.

49,521 次观看 • 6 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Depth Any Video with Scalable Synthetic Data AI physicists and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.

MrNeRF

27,428 次观看 • 1 年前

Wonderland: Navigating 3D Scenes from a Single Image Contributions: • First, we introduce a representation for controllable 3D generation by leveraging the generative priors from camera-guided video diffusion models. Unlike image models, video diffusion models are trained on extensive video datasets. This enables them to capture comprehensive spatial relationships within scenes across multiple views and embed a form of "3D awareness" in their latent space, which allows us to maintain 3D consistency in novel view synthesis. • Second, to achieve controllable novel view generation, we empower video models with precise control over specified camera motions. We introduce a novel dual-branch conditioning mechanism that effectively incorporates desired diverse camera trajectories into the video diffusion model. This enables expansion of a single image into a multi-view consistent capture of a 3D scene with precise pose control. • Third, to achieve efficient 3D reconstruction, we directly transform video latents into 3DGS. We propose a novel latent-based large reconstruction model (LaLRM) that lifts video latents to 3D in a feed-forward manner. With this design, during inference, our model directly predicts 3DGS from a single input image, effectively aligning the generation and reconstruction tasks—and bridging image space and 3D space—through the video latent space. Compared with reconstructing scenes from images, the video latent space offers a 256× spatial-temporal reduction while retaining essential and consistent 3D structural details. Such a high degree of compression is crucial, as it allows the LaLRM to handle a wider range of 3D scenes within the reconstruction framework, with the same memory constraints.

MrNeRF

52,849 次观看 • 1 年前

🚨 Anthropic committed up to 1M TPU chips for Claude. Openai is leasing TPUs for chatgpt inference. Here's How kernels work on TPUs (deep dive 2/6 by emilio andere) pallas is Google's answer to kernel writing. a python kernel SDK built on JAX. still very experimental (jax.experimental.pallas). on TPU it compiles through mosaic; on GPU it lowers to triton. if you know CUDA, the syntax will feel familiar but the execution model is completely different. in CUDA, grid=(4,4) launches 16 blocks running simultaneously across SMs. in pallas, those 16 iterations run one after another in lexicographic order. no threads. no warps. no blocks. no occupancy tuning. a TPU is a sequential machine with a very wide vector register — more like a CPU than a GPU. performance comes from width: a 128x128 systolic array doing matmul and an 8x128 SIMD vector unit doing everything else. maximum parallelism on chip: 2, one per TensorCore in megacore mode. three concepts replace CUDA's thread/block/grid hierarchy. Refs are mutable memory references. because execution is sequential, each iteration safely accumulates without atomics. in CUDA you'd need atomics or a separate reduction pass. the memory model is also very different from NVIDIA's. zero hardware caches. VMEM is 32-128 MiB of software-managed scratchpad — 500-1000x larger than GPU shared memory per SM. all data must be explicitly DMA'd from HBM to VMEM before any computation touches it. four levels: HBM → VMEM → VREGs → MXU/VPU, plus SMEM for scalar control data. every byte of data movement is your responsibility. this is like CUDA shared memory except it's 500x bigger and there's no cache fallback. pipelining is mandatory. without double-buffering HBM→VMEM transfers, the MXU just stalls waiting for data. this is the single most important optimization on TPU. and because grid execution is sequential and deterministic, consecutive iterations that need the same input block skip the redundant HBM transfer automatically, impossible on GPU where block execution order is undefined. the compilation pipeline is unlike anything in this series: python → jaxpr → stableHLO → XLA HLO (71+ optimization passes) → LLO (78+ passes) → 322-bit VLIW bundles. the compiler packs instructions for scalar, vector, matrix, and DMA units into a single 322-bit word. everything in that bundle executes in parallel, with no runtime scheduling.

wafer

33,655 次观看 • 2 个月前

You don't need a GPU for fast studio grade voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?

Alok

105,756 次观看 • 1 个月前

Kling 3.0 on Yapper Prompt: FORMAT: 15s / 8 shots / editorial contact sheet SUBJECT: Young woman with long blue-white gradient hair, dark leather armor with scale-textured panels, mounted on or posed beside a colossal dragon. ENVIRONMENT: Volcanic ridge at golden hour, soft directional sunlight through ash haze, no artificial light, no flash. MOOD: High-fashion creature editorial, every frame a magazine spread candidate. MUSIC: Minimal ambient bass pulse, slow and deliberate. COLOR LOGIC: Naturalistic Film Print Emulation STYLE: Fashion editorial, beauty portraiture LOGIC RULE: Each shot reads as a still editorial frame with near-frozen poses. Natural light only, shaped by ash diffusion and golden-hour direction. Rider and dragon share the frame as co-subjects. No flash, no strobe, no artificial fill. SHOT SEQUENCE: SHOT 1: Full-length, 50mm / Rider stands in profile against dragon's folded wing, one hand on hip, blue-white hair backlit by golden haze, dragon scales fill background as texture wall / SFX: soft wind SHOT 2: Hard cut. Medium close-up, 85mm / Rider faces camera, chin slightly lifted, dragon's jaw rests just behind her shoulder, shallow depth of field melts scales into bokeh / SFX: silence SHOT 3: Hard cut. Detail insert, 100mm macro / Rider's leather gauntlet resting on dragon's neck ridge, golden light raking across both textures, skin and scale side by side / SFX: faint ember crackle SHOT 4: Hard cut. Wide, 35mm / Rider seated sidesaddle on dragon's back, legs crossed, hair swept by updraft, ash particles drifting like golden confetti / SFX: low wind hum SHOT 5: Hard cut. Over-the-shoulder from dragon's head, 40mm / Rider looks back over her shoulder toward camera, face half-lit by warm side light, dragon horn frames the top of shot / SFX: deep breath SHOT 6: Hard cut. Low angle, 24mm / Rider standing on dragon's foreleg, full body, arms relaxed at sides, dragon wing spread behind her as a dark canopy, golden rim light on hair edges / SFX: membrane stretch SHOT 7: Hard cut. Extreme close-up, 135mm / Rider's eye and cheekbone, golden-hour catchlight in iris, single strand of blue-white hair across face, dragon scale texture reflected in pupil / SFX: heartbeat SHOT 8: Hard cut. Ultra-wide, 14mm / Rider walks away from camera along dragon's spine toward the head, dragon lifts chin to sky, fire by mouth both silhouetted against amber sunset / SFX: low brass swell

Zara

19,755 次观看 • 5 个月前

🚨 BREAKING — one of the strongest OpenClaw setups on Polymarket just went public. A trader reportedly started with ~$100–200 and scaled it to ~$3.7M. No insider access. No political connections. Just a developer running his own automation built with OpenClaw. Profile → Copytrade → I went through the framework myself. What surprised me: There’s no huge infrastructure. No complex quant stack. No giant data pipelines. Just clean logic and disciplined automation. After about 8 hours analyzing it, the strategy breaks down into three parts. 1) “Free money” via NO positions The bot targets outcomes with near-zero probability. Instead of chasing big wins, it accumulates a massive number of small high-probability NO trades. Not speculation — systematic probability harvesting. 2) Logical arbitrage Sometimes Outcome A logically implies Outcome B, but markets don’t adjust instantly. The bot detects these inconsistencies and enters before repricing happens. By the time the headline reaches traders, the window is already closed. 3) Retail-driven markets Sports and political markets are dominated by retail flow and emotional reactions. Prices overshoot, spreads widen, and inefficiencies appear constantly. The bot sits in those gaps and clips small edges repeatedly. Scale is the edge. 4,192 trades executed. Individually small. Together they compounded into roughly ~$3.7M profit. Largest single win: $1,464,152. The equity curve is almost vertical. It’s not about predicting events. It’s about exploiting structural inefficiencies faster than the crowd.

Discover

186,697 次观看 • 6 个月前

Created with Minimax H3 Max Prompt: "Create a 15-second cinematic, photorealistic continuous-shot video based exactly on the reference image of “The Floating Night Market.” Preserve the original architecture, floating wooden market platforms, lantern designs, vendors, customers, boats, ocean, mountains, color palette, and overall composition. Do not redesign or replace any major elements. 0–3 seconds: Begin with a slow, smooth low-angle camera movement just above the dark ocean surface, approaching the floating night market. Hundreds of warm glowing paper lanterns gently sway in the night breeze. Their golden reflections shimmer naturally across the moving water. Small waves softly move around the floating platforms. 3–7 seconds: The camera slowly rises and moves forward through the market, revealing lively activity. Vendors naturally prepare food at small wooden stalls, gently stirring steaming pots and grilling food. Thin realistic steam rises into the lantern light. Customers walk casually across narrow wooden bridges, talking and browsing. Lanterns flicker subtly and sway with the wind. 7–11 seconds: Continue the same uninterrupted camera movement with a gentle cinematic arc around the market. Small traditional boats slowly pass underneath the floating platforms. Water ripples spread behind them. Food stalls glow with warm orange and amber light while the cool blue moonlight illuminates the surrounding ocean. The huge full moon remains visible in the background behind misty mountains. 11–15 seconds: Slowly pull upward and backward into a wider cinematic reveal, showing the entire floating night market surrounded by the vast ocean. Hundreds of lanterns form a glowing trail across the water. Reflections stretch and break naturally with the waves. Mist gently drifts around the distant mountains as the moonlight creates a magical atmosphere. End on a beautiful wide establishing shot. Visual style: ultra-photorealistic, cinematic fantasy, realistic human movement, realistic ocean and water physics, detailed wooden textures, natural lantern glow, volumetric moonlight, subtle atmospheric fog, realistic steam, shallow depth of field during close shots, cinematic 35mm photography, high dynamic range, extremely detailed, immersive scale, smooth professional camera movement. Motion: natural human gestures, subtle wind movement, gently swaying lanterns, flowing steam, moving water, realistic boat movement, no exaggerated animation. Camera: smooth stabilized cinematic camera, slow controlled movement, natural depth of field, gradual push-in → forward tracking → gentle arc → wide pull-back. One continuous shot with no cuts. Negative prompt: no scene changes, no cuts, no jump cuts, no camera shake, no distorted faces, no duplicated people, no warped buildings.

Nafees

10,863 次观看 • 13 天前

SPACEX’S STARLINK As a B787 pilot, it pains me when posts about aviation are this wrong on every detail & every conclusion. Starlink is the best option for airliners. Amazon is a distant second. ♦️ BEST-IN-CLASS Starlink Aviation already uses proven flat phased-array antennas — no gimbals, no moving parts. Amazon’s is the same tech, not different or better. ♦️ TWO > ONE The speed claims are way off. Starlink’s real-world performance on airlines already beats Amazon’s unproven promises. There is no 250 Mbps cap. And one larger antenna isn’t automatically superior — its wider profile creates more aerodynamic drag than Starlink’s two smaller inline antennas. Two antennas also give dispatch reliability: if one fails, Wi-Fi still works. Starlink installs are famously quick and reliable. ♦️ GLOBAL SERVICE Major airlines are global operations, I cross 80 time zones every month — the size of the constellation is the differentiator. Coverage dead zones mean I don’t get real-time updated turbulence plots or full-storm radar maps in the flight deck. Saying the antenna is THE bottleneck doesn’t make it true. A Wi-Fi antenna with no signal from satellites is just useless dead weight. Check out the gaps on the Amazon constellation below. ♦️ MARKETING HYPE The AWS private interconnect is a marketing bullet, but nothing an airliner actually needs. Compute for the plane sits on the plane. Operational data exchange is heavily regulated and runs on dedicated satcom datalinks (ACARS and CPDLC). Looping AWS in adds zero safety benefit and simply creates another potential hacker entry point. As for analytics — airlines already excel at that on the ground where it belongs. ♦️ LOCKED-IN? Long-term lock-in to a single cloud provider is a disadvantage to many. Delta clearly got a discount on the AWS today to bundle in the promise of Wi-Fi tomorrow. But what happens once the introductory discounts disappear and Delta gave up all leverage? Starlink is delivering today at global scale. Amazon is still selling PowerPoint slides. Facts matter in aviation. Videos - Left: Starlink satellites Right: Amazon satellites

Amy

83,532 次观看 • 5 个月前

This week is already so hot. 🔥 Massive release from Decart : Lucy 2.0 a World Editing Model running at 1080p, 30FPS in realtime. This is truly exciting, the era of real-time generative reality is here. We are moving from watching AI video to living inside AI video. A breakthrough model capable of transforming the visual world in real-time. Moving beyond offline rendering, Lucy 2.0 delivers high-fidelity 1080p video generation with near-zero latency. Lucy 2.0 literally "redraws" the entire world pixel-by-pixel, while you are watching it. e.g. If you want to be an anime character, it doesn't just put a mask on you. It turns your skin into anime skin, your hair into anime hair, and the lighting in your room into anime lighting. Lucy 2.0 is also trained to stop the generated video from slowly falling apart over time, so the same stream can run much longer without faces and details drifting. So why is this a "Massive Deal"? Traditional AI video-generation model takes a prompt, you wait 10–20 minutes, and the computer "bakes" a video for you. You couldn't touch it or change it while it was happening. But Lucy 2.0 works like a mirror. It happens in real-time (30 frames per second). There is no waiting. You move your hand, the AI character moves its hand instantly. The craziest part isn't the visuals; it's the physics. Usually, AI hallucinations are glitchy—hands merge into faces, walls melt. Lucy 2.0 understands how the world works without being told. It knows that if you take off a helmet, there is hair underneath. It knows that if you splash water, droplets fly. It learned "physics" just by watching millions of videos. The physical behavior you see emerges from learned visual dynamics, not from engineered geometry or explicit physics engines. Their official technical report explicitly states that the model does not use traditional 3D engines, depth maps, or wireframes. It is a "pure diffusion model."

Rohan Paul

12,761 次观看 • 8 个月前

Hermes Agent + Higgsfield Marketing Studio = AI UGC Content Factory I built a fully automated system inside Higgsfield that repurposes, localizes, and launches winning TikTok Shop content across hundreds of creator-style accounts. It's so effective it feels like running Facebook ads in 2008. No actors. No products in hand. No ghost creators. Just viral TikTok Shop sales - 24/7. The results speak louder than any pitch: • CPMs as low as $0.10 • 550+ cinematic, product-ready ads per day from a single prompt • 100 hooks tested in the time it used to take to test 10 • $100/mo replacing a $50k+ creative budget Here's the full pipeline - all native inside Higgsfield Marketing Studio: > Hermes Agent analyzes your product, scrapes Meta Ads + TikTok Ads, identifies winning content, and localizes every angle to your brand. > Seedance 2.0 turns data into AI UGC ads - captions, pacing, hooks, your website showcase, all auto-edited inside Higgsfield Marketing Studio. > AI UGC personas are spun up with realistic faces, voices, and personalities - cloned voiceovers in seconds. > Our phone farm pushes every finished video straight to TikTok Shop, daily, on autopilot. >No setup. No switching between five tools. Everything lives inside Higgsfield Marketing Studio. Here's how it actually runs: Hermes Agent researches the niche, scrapes winning TikTok Shop videos, and rebuilds them with fresh hooks, angles, and UGC visuals tailored to your brand. Agents create and post daily to affiliate accounts - fully automated. Then we activate the MPS (Multi-Platform Swarm): once a concept wins on TikTok Shop, Higgsfield deploys hundreds of AI Agents to flood the niche with variations that all drive back to our shot. Most brands are still paying $300–$500 per video. Testing 10 hooks costs $5,000 and takes three weeks. With this system, we test 100 hooks in the same timeframe - and the winners scale automatically. TikTok doesn't reward the best video. It rewards the brand that shows up the most - with content that converts. The brands automating content at scale will be the biggest winners of 2026.

Noah Frydberg | Tiktok Shop For Brands

27,037 次观看 • 5 个月前

Researchers made KMeans 200x faster. And the new technique also beats approaches like cuML and FAISS. Flash-KMeans is an IO-aware implementation of exact KMeans that redesigns the algorithm around modern GPU bottlenecks. By attacking the memory bottlenecks directly, Flash-KMeans achieves: - 33x speedup over cuML - 200x speedup over FAISS This speedup comes from how it moves through GPU memory. Standard KMeans runs in two steps, and both are bottlenecked by reads and writes to GPU memory: 1) The first step matches every point to its nearest centroid. Standard KMeans computes the full point-to-centroid distance matrix, writes it out to GPU memory, then reads it back to find each nearest centroid. That write-then-read round trip is the bottleneck. Flash-KMeans combines the distance calculation with the nearest-centroid step, so the result is computed on-chip and the full matrix is never written out. 2) The second step recomputes each centroid by averaging the points assigned to it. Standard KMeans has thousands of threads writing into the same centroid slots at once, so they stall waiting for their turn. Flash-KMeans sorts points by cluster first, turning scattered writes into sequential reductions that read and write memory in one efficient pass. Using these two optimizations at the million-scale, Flash-KMeans completes a standard KMeans iteration in a few milliseconds. The video below depicts this in action. Several reasons why this is important: KMeans has always been an offline primitive. Something you run once to preprocess data and move on. These speedups make the approach viable in several runtime-critical systems. ↳ Vector indices like FAISS use KMeans to build search indices. Faster KMeans means you can re-index dynamically as data changes. ↳ LLM quantization methods need KMeans to find optimal weight codebooks, per layer, repeatedly. What takes hours could now take minutes. ↳ MoE models need fast token routing at inference time. Flash-KMeans makes it viable to run this inside the inference loop, not just in preprocessing. I have shared the paper in the replies. That said, memory is the real constraint Flash-KMeans solves, and the problem is not just limited to clustering. The vectors a RAG system stores after indexing create similar bottlenecks. I wrote a detailed walkthrough recently on cutting this vector memory by 32x with binary quantization, querying 36M+ vectors in a few milliseconds. Read it below.

Avi Chawla

89,234 次观看 • 3 个月前

Proud to announce the in-depth collaboration between Kingnet and Alibaba Cloud in AI Gaming. Alibaba Cloud provides world-leading cloud computing, big data, and AI services, with disclosed revenue exceeding $15 billion in 2024, which is one of the most renowned global server providers. When two superpowers collide, the game changes. 🌊AI Gaming R&D By integrating Qwen 's LLM and Alibaba Cloud 's PAI platform (including PAI-iTAG, PAI-Designer, PAI-DSW, PAI-DLC, and PAI-EAS), Kingnet has emerged as one of the gaming industry's pioneers in AIGC-powered content generation and AI rendering. Together, we are accelerating the realization of no-code game development. 🌊GPU Computing Resources Alibaba Cloud delivers GPU-accelerated elastic computing services with exceptional processing power, supporting diverse workloads including deep learning, scientific computing, graphics visualization, and video processing - providing robust GPU computing capabilities for KingnetAI's demanding requirements. 🌊Cloud Service Optimization Cloud server deployment has become the mainstream choice for small and mid-sized game studios in global operations. Leveraging Alibaba Cloud server advantages, we will develop and deploy more cloud-native games to meet user demands. The disruptive innovation we're bringing to the industry: 🔸Minute-scale game asset production replaces traditional week/month-long cycles 🔸Single-digit dollar development costs VS traditional four-figure entry thresholds 🔸AI-powered NPCs with behavioral engines deliver dynamic player interactions, breaking static story constraints, etc. 🔜Kingnet AI V2 is approaching launch. The Agent system and game generation engine will be officially deployed across 3 chains: 🔹Leveraging Solana high throughput and low gas fee , Solana has consistently been a developer favorite, latest product will be deployed on Solana - with users paying $SOL for on-demand asset creation fees. 🔹Another key partner is BNB Chain ,We are actively participating in both the #BNBAIHack and the latest MVB 10. Powered by BNB Chain long-standing support for AI innovation. Kingnet V2 and NFT drop will be deployed on BNB Chain, providing developers and the community with comprehensive game-generation tools and support. 🔹As an early strategic partner of Kingnet, TON 💎 @TONEastAsia was one of the earliest chain to connect Web2 and Web3, Kingnet V2 will be deployed on TON, providing TON game developers with low-cost, high-efficiency asset generation, and supporting users to use $TON as an asset generation cost. The Future of AI Gaming is coming.

Kingnet AI

149,965 次观看 • 1 年前

anthropic will sell you opus 5 at $200 a month. openai will sell you gpt-5.6 at $200 a month. neither will tell you stanford and berkeley published the 5 principles to build a $100k/mo ai company on kimi k3 for $10 stanford and berkeley spent years figuring out what actually separates ai systems that work in production from ai systems that die in demos. they published the findings. anthropic and openai priced their frontier subs like nobody would read the papers. the papers are free this is dspy plus verifiers plus decomposition plus skills plus mcp. five principles from stanford, berkeley and moonshot that turn a $10/mo kimi k3 sub into an ai analyst that runs unattended. the model is public. the system is the moat five moves that turn kimi k3 into the $100k/mo company: P1 don't prompt, program (stanford dspy) -> stanford proved hand-tuned prompts don't scale. define a pipeline as modules, let the optimizer tune them -> the compiled pipeline beat expert few-shot on multi-step tasks. one line of dspy replaces a month of prompt engineering P2 don't trust the model, build verifiers (berkeley 2026) -> a compiler either accepts or rejects. a test either passes or fails. that is a verifier -> berkeley: test-suite reward hit 42.2% pass@1 on swe-bench. hybrid verifiers hit 51.0% best@26. no bigger model, just a real check P3 don't scale agents, decompose them (stanford ai index 2026) -> stanford found multi-agent gains only 2-4 percentage points. two coding agents sometimes did worse than one -> the win is role decomposition, not count. researcher, writer, reviewer, verifier, clear input, clear output, no overlap P4 don't repeat expertise, encode it as skills (kimi code) -> every session starting from zero is institutional knowledge you lost. a skill.md file makes kimi activate the workflow automatically -> week one you write the skill. month six it encodes more institutional memory than most junior employees carry P5 don't keep ai in chat, connect it to tools (mcp) -> a model that only sees what you paste is a consultant working blindfolded. mcp connects kimi to your crm, db, github, linear, slack -> the model is public. the data is yours. the connections are your moat my position, and it is the arguable one: the next $100k/mo ai company will not win because it got early access to a frontier model. it will win because it followed 5 papers that anthropic and openai are quietly hoping you never read drop your $200/mo ai sub to $10. the swarm above is what 300 kimi k3 agents look like running those 5 principles. the full playbook is in the article below

starmex

31,358 次观看 • 1 个月前

Would you dare chase justice while swinging thousands of feet above traffic below? Seedance 2 prompt on BudgetPixel AI Create a 15-second ultra-realistic cinematic high-altitude tether-swinging action sequence in strict 16:9 landscape, native 4K, 24fps. Use one seamless continuous drone follow shot with no cuts, no teleporting, and no time skips. The motion must feel physically continuous, dynamic, thrilling, and always readable. REFERENCE: image1 = main heroine reference. Use image1 as the strict identity reference for the heroine’s face, facial proportions, hairstyle, hair color, body proportions, age impression, outfit, shoes, accessories, styling, and overall recognizable appearance. Preserve her identity consistently throughout the whole video. Keep her as an original urban tether-swinging action heroine. Do not redesign her into a branded superhero character. Do not add franchise logos, copyrighted chest emblems, or recognizable third-party superhero symbols. CORE CONCEPT: This is an original urban tether-swinging action short. The heroine moves through the city using thin wrist-launched fiber lines, momentum, wall-running, rooftop movement, and real parkour body mechanics. She travels at high altitude between tall buildings, then lands on a rooftop, defeats one villain, and ends with a powerful shout. HOOK: The first second must be an instant scroll-stopping hook. Start with the heroine already falling backward off the edge of a very tall skyscraper. For a brief moment, it looks like she may actually fall. Then she instantly fires one thin tether line upward, it catches, and her body snaps into a huge high-altitude swing between buildings. STYLE: Photorealistic live-action realism. Premium cinematic action quality. Bright daytime Los Angeles atmosphere with realistic haze, realistic motion blur, realistic fabric movement, realistic body weight, real inertia, and practical environmental interaction. The sequence should feel like a premium action movie shot, not animation, not a game cutscene, and not a cartoon. CAMERA: One uninterrupted drone follow shot only. No cuts. No resets. No jumpy edits. No impossible viewpoint teleporting. The drone camera must stay wide enough to show both the heroine and the environment together. It may tilt, roll, arc, climb, and dive with the motion, but it must always feel like one real flying camera tracking her. Keep the framing intense and fast, but always readable. ENVIRONMENT: Bright daytime in a dense modern city inspired by Los Angeles and Hollywood. Show: - tall glass and concrete high-rises - rooftop edges - billboards and signage - palm trees far below where visible - busy roads and traffic far beneath - bright haze and sunny atmosphere - believable large-scale urban depth The action must happen mainly high above the street between tall buildings, not low near the ground for most of the video. VILLAIN RULE: Only one villain appears in the entire video. The villain is one adult male enemy only. He wears a fitted black suit, black shirt, and black shoes. No mask, no armor, no fantasy costume. He appears only in the rooftop combat section. Do not generate multiple enemies. Do not generate background enemies. Do not clone or duplicate the villain. ACTION RULES: The heroine’s movement must feel hand-and-foot driven, not magical floating. She must visibly: - fire thin tether lines from her hands - swing with real tension and momentum - push off building surfaces - run along walls with clear foot placement - absorb landings with bent knees - sprint briefly on a rooftop - fight one villain using fast practical action - finish in control Her body mechanics must stay realistic: - core engaged during swings - arms extended or flexed according to line tension - knees bend on landing and push-off - visible transfer of momentum between swing, wall-run, leap, landing, and combat Do not make her hover weightlessly. Do not make the tether line act like magic. Do not make her float in place unnaturally. EMOTIONAL ARC: - opening: shock and immediate control - mid-swing: intense focus - rooftop approach: rising confidence - rooftop fight: sharp aggression and urgency - ending: victorious adrenaline and fearless release AUDIO: No music. Effects and ambience only: - rushing wind - tether firing and tension snaps - air pass-by - foot impacts on walls and rooftop surfaces - city ambience far below - distant traffic and horns - fabric movement - breathing - one short rooftop fight impact sequence - one powerful final shout from the heroine TIMELINE: 0:00–0:01 Start from black into a shocking rooftop-edge fall. The heroine is already dropping backward off a skyscraper. For a fraction of a second it feels dangerous and uncontrolled. She immediately flicks her wrist and fires one thin tether line upward. It catches instantly. The drone yanks back and reveals the start of a huge swing. 0:01–0:04 The heroine swings at high altitude between tall buildings. The city is far below. Her body forms a long aerodynamic arc, one arm holding tension through the line, legs trailing cleanly behind. The drone follows wide and slightly rolled, emphasizing height, speed, and scale. 0:04–0:06 At the swing’s forward rise, she releases the line and redirects toward a nearby glass-and-concrete building. She plants onto the wall and runs across it diagonally with 4 to 5 clear steps. Her feet hit the wall with visible force. Her jaw is set and focused. The drone stays close but wide enough to keep the city depth visible. 0:06–0:08 She pushes explosively off the wall, fires a new tether line, and swings again through a narrower corridor between tall buildings. The movement should feel faster and more controlled now. She threads cleanly through the urban gap and angles toward a rooftop landing zone ahead. 0:08–0:09.5 She releases the line and lands hard but controlled on a rooftop. Knees bend deeply to absorb impact. She rolls into a short forward recovery step, then rises immediately into a sprint across the rooftop surface. 0:09.5–0:12 One villain in a black suit steps in to stop her. Keep only this single enemy. The heroine engages him in a short, sharp rooftop fight. She avoids his first attack with a quick slip, grabs or redirects his arm, drives one fast body shot or elbow, then uses his off-balance momentum to throw or slam him down onto the rooftop. The fight must feel quick, practical, and decisive. Real impact reactions. No slow choreography. No extra enemies. 0:12–0:13.5 The villain is down and no longer a threat. The heroine steps past him and moves to the rooftop edge. Wind moves her hair and outfit. She looks outward over the city with intense adrenaline and triumph. 0:13.5–0:15 At the rooftop edge, she turns slightly toward the open skyline, lifts her chest, and shouts one powerful final line: “가자!” She immediately launches forward off the rooftop edge into another leap just as the clip ends. End on the feeling that the action is continuing beyond the cut. IMPORTANT RULES: - one continuous drone follow shot only - no cuts - no teleporting - no cloning - no multiple villains - only one black-suited villain - no giant web canopy - use only thin functional tether lines - no franchise logos - no copyrighted chest symbols - preserve the uploaded identity consistently - action must stay realistic and momentum-driven - rooftop fight must be short, sharp, and readable - final shout must be “가자!” NEGATIVE: no cartoon, no anime, no game-engine look, no fake CGI stiffness, no floating, no weightless hovering, no random disconnected acrobatics, no city-wide web canopy, no superhero logo, no copyrighted spider emblem, no extra enemies, no masked villain, no armored villain, no cloned villain, no empty city, no dark night setting, no rain, no slow motion, no blurred identity, no outfit drift, no face drift, no extra limbs, no broken anatomy, no unrealistic hand deformation, no collision with buildings during swings, no messy unreadable fight.

Sharon Riley

59,224 次观看 • 1 个月前

You Can't Vibe-Code Trust Avishai Abrahami, Co-Founder & CEO of Wix , interviewed by Harry Stebbings (kevin andres) Summary: Wix trades at a $2.8B market cap on $2.1B of revenue while the market ascribes roughly zero value to a business throwing off $400M a year in free cash flow. Wix CEO Avishai Abrahami's argument is that the market can't yet price what AI actually threatens: the moat is trust and business logic, and neither gets vibe-coded away. His response is to own the disruptor (Base44), train his own narrow models, and stay committed through a storm he insists always arrives on a random Wednesday. 1. Trust is the moat. The real value of Salesforce is trust: JP Morgan and huge banks let it hold all their customer data, and the CRM itself is a small part of that. "What other platform will JP Morgan trust for their customers' data? None." That trust took years to build and can't be reconstructed by an agent scraping a database, so the companies whose value lives in trust survive the SaaS apocalypse while the ones reduced to piping get commoditized. 2. The business-logic wall. "You're not going to vibe-code Shopify no matter how good you are. The business logic is too hard." Wix tested this directly: they asked a team of professional developers to build the operating logic for a single hairdresser in Base44, gave up after a week, brought in a stronger team, and still failed two weeks later. Complex operational software is far harder than a demo suggests, which is why the pizza shop and the hairdresser stay Wix customers rather than build their own stack. 3. Own the disruptor. Wix bought Base44, a one-person company, for $80M, and it now does over $150M in ARR, roughly double what they paid. Abrahami frames the future as three buckets: owners who never want to build, owners who vibe-code everything themselves, and a mix in the middle over the next five or six years. Rather than bet on which wins, Wix owns the tool customers would defect to, so a customer who switches platforms still switches to Wix. 4. Trading on someone else's news. "Today we are trading on other companies' news. We're not trading on Wix news. We're trading on what OpenAI or Anthropic or Google are saying." Base44 alone, valued on vibe-coding peer multiples, should be worth around $8B, which means the market assigns less than zero to Wix's core. Abrahami's response is to detach: he doesn't wake up checking whether the stock moved 20%, because the only thing he can influence is the business. 5. The narrow model. Wix fine-tuned and combined its own models and now matches top-tier frontier quality on Base44 tasks at far lower cost. The logic: they sit on a huge stream of training data from watching what users try and where they fail, so a model built for Base44 can skip what frontier models carry, like knowledge of Chinese poetry, and go deep on what someone means when they say "build me a task manager to tell my boyfriend where he's wrong." A narrow target is easier to hit than a frontier model, and Wix already runs a trained model on website generation that's faster, cheaper, and makes fewer errors, retrained weekly on a live feedback loop. 6. Quality before cost. When Harry cites Chamath's claim that open source runs 14-16x cheaper, Abrahami pushes back: that holds for small tasks, but for something as complex as Base44 the savings land at 5-10%, and his own model runs 1-30% cheaper than frontier, not the order of magnitude people assume. More to the point, this is the wrong time to chase cost: "20% more quality, 20% less cost, I'll go for the quality." It's a brand-new market that's just starting, and the job now is to make the product better. 7. The but is very big. "We all give too much credit for AI. It's amazing, it's incredible, it's super powerful, but the but is pretty big." He asked Claude to write a safety protocol and got six mandatory gates, then pushed back on each one and watched the model cave until only one survived, downgrading the rest from "must test" to "might want to look at later." We over-trust these systems, and that reflex, treating a Reddit post as equivalent to research published in Nature, is where the danger lives. 8. Customer support still breaks. Wix has 3,500 people and its single biggest department is customer support, serving 192 countries. They tried hard not to build their own AI support agent, tested many off-the-shelf products, and concluded flatly: "It doesn't work. We tried, we tried again, it didn't work." The gap between hyped AI support startups and what actually ships in production is the tell that the technology is earlier than the marketing, maybe five years from being different. 9. Buybacks as dividends. Wix had $1.5B sitting in the bank it couldn't put into a major acquisition because it was focused on the new product and Base44, so it bought back stock at a low price, with admittedly terrible short-term timing. Abrahami is unbothered: "The big question is where it's going to be in three years, not what happened in the last three months." He argues buybacks are a fantastic, underused tool, essentially a dividend to every shareholder, and companies should lean on them to balance stock-based compensation instead of endlessly diluting. 10. Execution, not finance. A low stock price makes M&A currency less valuable, but Abrahami says that's not his real constraint. Base44 was a one-person company; Wix had to build an entire company around it, staffing it with people pulled from the core. "I don't know how to do another one of those at the same time and have the same quality." The bottleneck on the next acquisition is execution capacity, not the balance sheet. 11. Chosen to be here. The one thing money buys beyond food security is freedom, and the deepest form of that freedom is knowing you're here by choice. "I'm here because I've chosen to be here. Nobody made me." He could move to Costa Rica or dance carnival in Brazil, and choosing to stay and run a public company through a crashing stock is where he finds his power. Money also made him more impatient and a bit lazier, and more rational because he's no longer deciding from fear. 12. The random Wednesday. Resilience starts with accepting the storm will come, because we assume that if yesterday was easy tomorrow will be too, and reality doesn't move in gentle slopes. "The worst thing that happens is probably some random thing on some random Wednesday. It's not something you get a lot of warning for." His anchor, borrowed from Babylon 5, is that you get there when you get there and the weapons you have are the weapons you have, so the only real question is whether you're doing the best you can with what you control.

Gokul Rajaram

22,843 次观看 • 1 个月前

🚨 OPEN PHYSICS: LIGHT IS TRAPPED! MATTER IS AN ILLUSION. THE HIGGS BOSON IS UNNECESSARY. PHYSICS IS A BRANCH OF SPECTRAL GEOMETRY! 🚨 For decades, the academic establishment has wasted billions on gargantuan colliders trying to isolate "mass particles" like the Higgs boson within a fictional, continuous infinite void. They have cluttered the Standard Model with dozens of arbitrary, manually tuned free parameters- and forced the world to accept it as dogma. It is a mathematical hallucination. The era of continuous fairy tales is over.We have just completed the definitive, parameter-free simulation of mass generation through pure Spectral Geometry. No data-fitting, no continuous approximations, and no ungrounded metrics. Everything is rigorously locked into a discrete, constructible number-field architecture. Here is the truth of reality, explained plainly for everyone: ⚡ WHAT IS MASS? IT IS TRAPPED LIGHT. Imagine a wave traveling freely across an open body of water. Suddenly, it strikes an unyielding, insurmountable boundary wall. The wave reflects back into itself, creating a stationary, localized "standing wave" that no longer propagates-it simply vibrates in place. Our universe is not a smooth, infinite ℝ⁴ continuum. It is a rigid, pixelated topological crystal. When a gauge photon propagates through this discrete lattice and hits its absolute boundaries, space literally runs out of algebraic degrees of freedom! The light has nowhere left to go. It undergoes total internal reflection against the vacuum substrate, freezing into a localized standing resonance. This trapped, bounded energy of phase-locked light is what we observe as REST MASS. You, me, the stars-everything is made of trapped light. 🚫 STOP TRYING TO FORCE CONTINUOUS TESTS ON DISCRETE SPACE! Mainstream critics continuously fail because they attempt to evaluate a pixelated universe using fluid, continuous differential equations. It is an elementary category error. The quantum vacuum does not compute infinite limits; it resolves rigid, discrete Diophantine constraints over algebraic integers. The Higgs boson and fundamental "point-particles" are optical illusions-macroscopic field shadows of pure geometric optics operating on the discrete Planck scale. 💻 THE HARD HARDWARE SPECIFICATIONS OF THE MATRIX: Here are the mathematical source codes of our IT³ Topological Engine. Copy them. Compile them. Try to break them: ■ The Arena (The Spectral Geometry Bound): Space is not a flat ℝ⁴ void. It is a strictly compact, constructible 7-dimensional manifold: ℳ_tot = 𝕊⁴ × 𝕋³ ■ The Absolute Limit (The Kummer Barrier): The Fibonacci-Pythagorean metric cascade truncates abruptly. The vacuum physically lacks the sub-harmonic resolution to partition space past this hard algebraic wall: √13 ∉ 𝕂, where 𝕂 = ℚ(√2, √3, √5) ■ Total Internal Reflection (The Velocity Kill):At the Kummer boundary, the group velocity of the photon fields drops identically to zero: v_g → 0 ■ Mass Generation (The Dual 16-Bit Register): The incoming gauge photon strikes the absolute √13 wall and triggers a rigid hardware XOR-shift (...0 ↔ 1...), splitting into its precise CP-conjugate Mirror Antiphoton. Their interference pattern locks into our master 16-bit vacuum volume anchor (57,600 = 240²): 1110 0001 0000 0000_2 → Rest Mass Is Born ∆ 🔔 THE OPEN GLOBAL CHALLENGE: 5 LITECOIN BOUNTY I do not want corporate grants. Academic gatekeepers hide their broken models behind paywalls because admitting a discrete geometric reality instantly invalidates their lifetime tenures. My reviewer is the COMPILER. Mathematics is the SUPREME JUDGE. True geometry has no need for a peer-review cartel that is terrified of alternative, open-source paradigms! This is not a debate; it is simply the deterministic solution to a 100-year-old error. Here is my public challenge to every mainstream physicist on this planet: Oprovergnite this mathematics. If you can mathematically break this discrete IT³ compiler architecture, 5 Litecoin is yours for a round of beer. You cannot. Run the simulation engine! Watch the attached rendering. You can see the free gauge light (cyan wave) hit the Kummer wall and instantly crystallize into localized matter (yellow resonance). If you cannot explain reality simply, you are hiding behind bad mathematics. God geometrizes. The Universe is a computed, 16-bit macroscopic topological processor. 📐🌌💎 🔗 Read the open-source proofs and grab the code: 🔗 🔗 #Physics #QuantumMechanics #SpectralGeometry #Math #Truth

Dr. Logvinovich

19,074 次观看 • 5 天前

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.

Alok

63,689 次观看 • 3 个月前

Would you underestimate her just because she wears a school uniform? GPT Image 2 + Seedance 2.0 on Sjolt Try Canvas: prompt Character Identity Lock (Highest Priority): Use the exact same young East Asian woman from the provided reference character sheet. Preserve 100% identical facial features, face shape, eye shape, nose, lips, skin tone, hairstyle, hair color, proportions, and overall identity throughout the entire video. Do not redesign, reinterpret, or substitute the character. She must remain instantly recognizable as the same person from the reference image. She has shoulder-length wavy silver-gray hair with subtle blue undertones, bright expressive eyes, fair skin, and a confident slight smile that naturally transitions into a focused, determined combat expression. She wears the identical navy blue Korean high school uniform blazer over a gray sweater vest, white collared shirt, striped tie, and matching school skirt from the reference character sheet. Video Prompt: A cinematic, hyper-realistic action sequence inside a chaotic South Korean high school classroom. The classroom is filled with overturned desks, scattered chairs, flying notebooks, broken pencils, and papers drifting through the air. Bright natural daylight streams through large classroom windows, creating realistic highlights, soft shadows, and cinematic contrast. The young female student moves with incredible speed, confidence, and precision as she expertly defends herself against multiple aggressive male students wearing matching Korean school uniforms. Every movement is fluid, athletic, and grounded in realistic martial arts choreography. The camera remains highly dynamic, featuring cinematic handheld tracking shots, fast push-ins, orbit shots, dramatic slow-motion moments, whip pans, low-angle hero shots, and close-up impact shots. Capture rapid combinations of punches, clean high kicks, evasive footwork, parries, elbow strikes, blocks, and throws. Desks slide across the floor, chairs topple over, and dust particles catch the sunlight, emphasizing the intensity of the action. Maintain a high shutter-speed action-photography aesthetic with crisp motion detail, subtle motion blur only during extremely fast movements, physically accurate body mechanics, realistic cloth simulation, natural hair physics, authentic facial expressions, and believable impact reactions. Keep the camera frequently returning to sharp close-ups of her face to reinforce character continuity and emotional intensity. Her silver-gray hair flows naturally with every movement while her determined eyes remain locked on her opponents. Photorealistic cinematic quality, 4K HDR, ultra-detailed skin textures, realistic lighting, volumetric daylight, physically based rendering, shallow depth of field during close-ups, blockbuster Korean action film aesthetic, empowering heroine energy, consistent facial identity throughout every frame, no face drift, no character variation, no animation-style exaggeration.

Sharon Riley

26,184 次观看 • 2 个月前