
Artificial Analysis
@ArtificialAnlys • 111,244 subscribers
Independent analysis of AI
Shorts
Videos

Example generations from HappyHorse-1.0 compared to Dreamina Seedance 2.0, Kling 3.0 Pro, grok-video-imagine and PixVerse V6 (Text to Video with Audio): Prompt [1/4]: A hula hoop spinning on a kid's waist, gradually climbing to their chest, then dropping to knees, then clattering to the floor. They pick it up to try again.
Artificial Analysis36,741 görüntüleme • 3 ay önce

Following up on our Intelligence Index v4.1 release yesterday, in the video below, Daniel from our team shares a short overview of what's changed: 1. Three upgraded evaluations: Terminal-Bench 2.1, τ³-Bench Banking and GDPval-AA v2 2. Cost, time, and tokens per task: Understand the cost, time, and tokens of tasks across our Index and for individual evals, and how these trade off against Intelligence 3. Cached input token reporting: We now report the amount of cached tokens a particular model uses and how this influences cost
Artificial Analysis12,491 görüntüleme • 1 ay önce

Gemini 3.5 Flash is a step forward for Google on speed and agentic capabilities but comes at a trade-off of being higher cost than prior models We have measured up to ~280 output tokens/sec, placing it on the speed/intelligence Pareto frontier and well ahead of Gemini 3 Flash. It also shows a major uplift on agentic tasks, reaching ~1650 ELO on GDPVal-AA. The trade-off: cost is up ~5x versus Gemini 3 Flash, driven by higher token prices (3x higher than Gemini 3 Flash) and higher token usage. In this video, Declan Jackson, Member of Technical Staff at Artificial Analysis, breaks it down.
Artificial Analysis16,055 görüntüleme • 2 ay önce

We’re unveiling a new look for Artificial Analysis! We’ve come a long way since launching Artificial Analysis over 2 years ago. Today, we benchmark 400+ models, 50+ inference providers, and benchmark not only language models but also image, video, speech, music, hardware, and agents. Our mission to support the AI ecosystem with independent benchmarking remains the same, but our brand and website refresh is designed to better reflect how much we’ve grown and how much further we plan to go. A huge thank you to everyone who has been part of the Artificial Analysis community along the way: from developers choosing models and building agents, to labs, inference and hardware providers, and fellow independent researchers.
Artificial Analysis18,100 görüntüleme • 3 ay önce

Overview of our recent launch of Coding Agent benchmarks on Artificial Analysis and our first Youtube Video! We walk through the performance, cost, token usage and speed differences across different coding agents. This includes looking at Opus 4.7 in Claude Code's leading performance and Composer 2.5's strong positioning on the Coding Agent Index / Cost Pareto frontier. We have also launched our YouTube channel! Come say hi and subscribe:
Artificial Analysis10,623 görüntüleme • 2 ay önce

Announcing Artificial Analysis Video Arena - the first crowdsourced comparison for Text to Video models Text to Video models are accelerating rapidly and crossing quality thresholds every month. We created Video Arena to compare them using the only source of truth for visual media - human preference! Video Arena includes hundreds of videos from the leading video models, including: - Runway's Runway Gen 3 Alpha - Pika's Pika 1.5 - Luma's Dream Machine - MiniMax / Hailuo AI-MiniMax Hub - Kling AI's Kling 1.0 - Zhipu AI's CogVideoX-5B Voting is open now and we’ll be announcing the first leaderboard results within 24 hours. Any predictions? In the meantime, you can see your own ‘personal leaderboard’ of how you’ve ranked the video models after 30 votes. Link to the Artificial Analysis Video Arena in the below tweet! 👇
Artificial Analysis27,096 görüntüleme • 1 yıl önce
Daha fazla içerik yok.