Most teams collecting voice data optimize for volume over... quality, partly because they’re measuring quality wrong. To help evaluate quality we created the Poseidon Score. When applied, single-speaker audio scored well while multi-speaker conversations scored worse. Why? ↓show more

Poseidon
28,647 просмотров • 5 месяцев назад
Conversational voice data exposes a subtle failure mode in... many ASR pipelines. The metric says the data is bad, but human reviewers say the audio is clear. Often the issue isn’t transcription quality itself, but that the evaluation stack wasn’t built for turn-taking, silence, and multi-speaker structure. Learn how to solve this in our latest blog:show more

Poseidon
11,202 просмотров • 5 месяцев назад
Most video models do one thing well. MiniMax H3... does everything, together. Our next-gen open-weight multimodal model reads text, image, video, and audio as a single creative language: motion, sound, emotion, cinematography, all connected. The result: precision-level generation with real cinematic quality. → Commercial-grade: film, ads, MVs, UI, game CG → 2K/24FPS, native stereo audio → Voice cloning + multi-asset reference for characters, motion & camera Concept to production. One seamless flow. | #MiniMaxH3 Hailuo AI (MiniMax)show more

KOLOVESKI
71,118 просмотров • 1 месяц назад
"We have everyone do work trials so people know... what they’re getting into on both sides. We like candidates to do real work for 1 or several days, often over a weekend. When they see the office full on a weekend, they quickly learn that we’re not joking around." nico laqua What are your single biggest lessons on how to test the quality of candidates pre-hiring Jack Zhang Ryan Daniels Ivan Burazin Ronak Maldeshow more

Harry Stebbings
36,407 просмотров • 3 месяцев назад
Sakana Fugu surprisingly performed near GLM 5.2 level but... 17× more expensive! We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs for 8 symbols, and a custom dark-theme UI. Outputs: Fugu Ultra — 22,225 t, $0.51 Opus 4.8 — 15,802 t, $0.31 GPT-5.5 — 11,474 t, $0.26 GLM 5.2 — 13,677 t, $0.03 Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and costshow more

atomic.chat
744,642 просмотров • 2 месяцев назад
Honestly this does touch on our philosophy. We could... very easily chase profits and charge $50 because it's entirely made in USA, the quality of the cotton and weave, the printing method, the inks we use etc. Obviously yall see others doing it for even lower quality products. Our pricing is not reflective of the product but we do it to keep it affordable and accessible and to help grow the mission of reshoring industry. A lot of people are trying to be made in USA businesses but also charge out the absolute ass trying to maintain margins and an income they'd grown accustomed to. It's an unrealistic expectation in our current economy and you're shooting the movement in the foot pricing yourself out of volume. Made in USA needs volume and cashflow. Goods need to be moving and exchanging hands, that's what drives the economy. If you're out there charging a premium people will still continue to go buy the foreign slop because it's cheaper and they don't share the values of what supporting made in USA really is. Shirts specifically if you're still using foreign shirts and cotton, they use the absolute cheapest shit they can find. (Just like wool, there's varying degrees of quality genetics that determines softness and performance). Probably the most frequent comment we get is about softness and how crazy comfortable our stuff is. That's for a reason and not by magic or accident. Our designs don't have that gross plastic stiff feel to them that flakes off and degrades quickly. That's for a reason because of how we choose to do things the hard way. We understand we care more about these things than most and for the general public they willingly buy lower quality stuff more to just support who is selling it. But I think awareness is also part of the issue because I was also one of those "it's just a shirt, who cares" people until I wore an AL shirt. They are hands down the biggest bang for the buck in this space and we have shirts lasting people years and years now for the same price you pay to fund china whether directly or indirectly. The best (and how capitalism is supposed to work) way to grow the movement is to be an educated consumer. Check tags, country of origin labels, ink and printing method. If someone is being shady or dodgy over the answers? There's your answer. And this goes for all products. We keep an active list of all kinds of made in USA products for all sorts of things that we have personally bought and used to check for quality both the product but also the business. anyway. TLDR we give up massive profits because we believe in what we're doing even tho it slows growth.show more

AGAVE
15,058 просмотров • 1 месяц назад
To replace animal testing with AI, we need MASSIVE... human datasets. Today, we're thrilled to share Axiom's new data exploration tool, providing the ability to visually explore the world's largest primary human liver toxicity dataset. Built with Axiom's proprietary wetlab protocols, our dataset includes detailed liver toxicity profiles for over 100,000 distinct molecules. The key to this dataset is our ability to do high-throughput, multiplexed high-content screening with primary human liver cells. Traditionally, toxicity assays either sacrifice throughput or sacrifice biological relevance (using easy-to-grow immortalized cell lines instead of real human cells). We managed to combine throughput, physiological relevance, and multiplexing in one platform. The assays run in a high throughput format using automation, meaning thousands of compound-dose conditions can be tested in one experiment. We achieved this using pooled primary human hepatocytes, which are often fragile and expensive. By systemizing our automation and quality control processes, we were able to run over 120+ batches on the same donor pool with incredible reproducibility and consistency. We did this while integrating many readouts per well, whereas many existing toxicity assays only do a single readout. Our multiplexed approach provides far more data per experiment enabling us to measure 10-20 different toxicity phenotypes such as apoptosis, necrosis, mitochondrial fission, endoplasmic reticulum stress, stress granule formation, microtubules, and more all from a single well on a 384-well plate! The combination of scale, high content information, and data quality is exactly what is needed to train highly accurate AI models in biology. If you're interested, please explore the dataset in the comments below and let me know if you want to chat about the details!show more

Brandon White
25,117 просмотров • 1 год назад
Vikrant Gupta exposed KL Rahul's PR which is being... ran through Johns & Mufa since this series loss against India: Vikrant Gupta - "KL Rahul played WC in Aus & he was dropped coz he used to play maiden overs." KL Rahul's scores in 2 T2OI WCs against big teams: vs Pak - 3(8) vs NZ - 18(16) vs Pak - 4(8) vs SA - 9(14) vs ENG - 5(5) Not a single score of 20 against any good team & below are his scores against minnows in those 2 T2OI WCs: vs AFG - 69(48) vs SCO - 50(19) vs NAM - 54(36)* vs BAN - 50(32) vs ZIM - 51(35) Can you believe how shit he is that he scored 5 50s when he saw minnows in front & failed completely when there were some quality teams. That is why he was kicked out from Indian T2OI setup & after that India won 2 T2OI WCs.show more

Rajiv
71,799 просмотров • 1 месяц назад
You don't need a GPU for fast studio grade... voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?show more

Alok
60,514 просмотров • 27 дней назад
$80M Series B startup vs. 4-hours Claude Code "Polish"... is this vague mystical quality that designers are obsessed with. Not because it's cool, but it influence perception way more than we are willing to admit. We as consumers are constantly "experiencing" things, products, visuals etc. It's the reason why you wouldn't rush to throw away that box the second you pull your new iMac out. It's that subliminal memory you had when you use iPhone for the first time, with that inertia scrolling across the magical glass. There is a place for extreme speed and affordability, but at a certain stage of your startup journey, there will also be a stage where details and perceptions are extremely expensive to get wrong.show more

John | Formfactor Design
110,893 просмотров • 4 месяцев назад
This one was made with Seedance 2.0 Fast via... Dreamina. This is pure Omni-Reference. The only character sheet I used was for these girls, Sari and Ploy. The dude with sarung here and the location were 100% prompted. I didn’t use a character sheet or reference for either of them. Even in Fast mode, Seedance 2.0 is bloody good and it still nails the hyper-vernacular vibe that I always aim for in my work. Seedance 2.0 is both exciting and scary for me 😆 It’s exciting because it is undoubtedly the best model currently available on the market. Trust me, you’ve seen the videos I’ve made so far right? The performance of the model It’s simply the best, period. It has helped me tremendously in creating a shit ton of stories about the region where I live, Southeast Asia. It has been the most exciting thing ever. The scary part is whenever a platform or company comes to me saying, “Hey, we have this new video model. Blah blah blah. We’ll let you know more soon.” It scares the shit out of me because the big question is whether it will be better than Seedance 2.0??? 😆😆 If not, I don’t even want to bother using it. I’ve come this far and achieved this level of quality with Seedance 2.0. That’s why I skipped Happy Horse, which I already tested. It’s also why I’m not bothering with Wan or anything else for now. Their current models are still far inferior to what we already get with Seedance 2.0. I don’t want to downgrade the visual quality. This is also why I need to be really honest. There are certain platforms that host their own in-house models and i'm still part of their CPP. However, because those models are still far behind the quality of Seedance 2.0, I haven’t used them that much. Seedance 2.0 has simply become the benchmark for me. The type of output I’m looking for is also extremely specific, so I can immediately feel it when a model cannot deliver what I need. Seedance 2.0 is definitely not cheap, but it gives me so much creative satisfaction and allows me to make whatever I want. I even have a team that low-key makes softcore erotic videos in the style of Vivamax 😆 I think I’ve trimmed down so many things in my AI workflow because my main goal is to focus on the content itself. If Seedance 2.0 Mini is released soon, I’m dead curious to test it. I think I want to create more stories that revolve around drama rather than highly technical cinematic shots. Seedance 2.0 Fast has been incredibly helpful, but I’m definitely curious to check out the Mini version. But the truth is that I’m completely tool-agnostic. I don’t care which company makes the model. I only care about the quality. You might remember when I praised Grok Imagine Video so damn hard because it was genuinely amazing back then. Then the quality kept getting worse and worse, so I stopped using it. But if it gets better again, I’ll definitely want to use it again. At the end of the day, quality is the only thing that matters.show more

MXVDXN // DAN
18,949 просмотров • 2 месяцев назад
🚀 Sol-Attn is here! We present a training-free sparse... attention method that accelerates video generation while better preserving quality. Sol-Attn unifies dynamic routing, sparse computation, and approximate correction in a single online-softmax pass: • On-the-fly block thresholding for dynamic yet controllable budgets • Proxy-score reuse to approximate unselected blocks Results (vs dense FlashAttention-3): • Wan 2.1-14B: 2.02× end-to-end • HunyuanVideo-13B: 2.12× end-to-end • LTX 2.3: up to 2.4× end-to-end When integrated into Sol-Engine (with kernel fusion + caching): • Wan 2.1-14B: 3.48× end-to-end • HunyuanVideo-13B: 5.08× end-to-end Already available in Sol-Engine. The B200 kernel is still under further optimization. 🎬 Project: 📄 Paper: 🔗 Code:show more

Enze Xie
20,752 просмотров • 1 месяц назад
I just built a Claude skill that audits your... entire Google Ads account in under 5 minutes 🤯 One prompt → a full account score, wasted spend breakdown, and a prioritized fix list telling you exactly what to change this week. All inside Claude Cowork. Perfect for DTC brands and agencies who are running Google Ads but have no idea how much budget is leaking. If you're managing Google Ads and your "optimization" process is logging in, staring at the dashboard, sorting by cost, and hoping you spot the problem before it costs you another $500... This audit skill finds it for you: → Connects to your live Google Ads data via MCP → Scores your account across 6 dimensions: wasted spend, search term quality, keyword health, quality scores, budget allocation, and creative performance → Calculates your exact wasted spend in dollars — search terms burning budget with zero conversions → Flags quality score issues dragging up your CPCs → Identifies keyword cannibalization across campaigns → Surfaces your top 5 highest-priority fixes ranked by budget impact → Generates a clean audit report you can hand to a client or share with your team No CSV exports. No pivot tables. No guessing where the money went. What you get: → A single Claude skill file you install once → An account health score (0-100) every time you run it → Exact dollar amount of wasted spend identified → Prioritized action list — not "optimize your account," but "pause these 12 search terms and save $847/month" → Works with any Google Ads account connected I'm giving away the full audit skill — the actual .md file you drop into Claude and run against your own account. Want it? Like this post Comment "SKILL" And I'll send it over (must be following so I can DM)show more

Mike Futia
60,106 просмотров • 5 месяцев назад
In this building with the black awning in this... Nick shirley clip, there is a ELDT truck driving school in this building called Titan Driver Training. As far as we know they're legit. In that building alone 1821 University Ave W MSP MN, $400k in PPP loans and fed awards. Titan was the only tenant in the building to not receive government funding. Quality Learning Center got $108k next door by itself in just PPP funds. We got your back wellness is a chiropractor in the same building doing FMCSA CDL Medical Exams. This address doesn't touch 2905 NW Blvd Suite 150 MSP MN where 50 different entities have received $1.7M in PPP, federal and state awards all of them are in Suite 150. If you're interested in the data we have it from MN State and the feds for all zip codes with the prefix 554 as well as PPP data we started collecting it 2020 because we were having so many issues operating in truck and bus terminals in that area.show more

Rob Carpenter
421,355 просмотров • 8 месяцев назад
Manuel Neuer was asked to choose between Cristiano Ronaldo... and Lionel Messi as the best player, and this was his response: 🗣️ Manuel Neuer: “Cristiano Ronaldo and Lionel Messi are both monsters and great players. Of course, we all know they are two different types of players. I was privileged to play against both of them during my career, especially when they were in their prime, but I have to say I suffered more against Cristiano Ronaldo than Lionel Messi. No matter how tight or strong our defense was, he always found a way to score. I think he is the player who has scored more goals against me than any other player I’ve faced. Cristiano can score with his left foot, right foot, and head. He can score all types of goals and even create space for his teammates. So, for me, I would pick Cristiano Ronaldo over Lionel Messi because he’s just a different breed the most complete player I’ve ever seen.”show more

SIR1⭐️
44,336 просмотров • 15 дней назад
What the heck is this?🤨 First, Meghan Markle’s audio... quality is horrifically bad and not even amateur, it’s worse than that. The part about Lili is barely understandable. It’s also so fake with her standing in the fake kitchen with a pristine white apron, which reinforces the idea that she doesn’t cook nor cares about her products. And the stupidity of continuing to believe that claiming her hostages, err family, can help sell these spreads by sharing their favorites is cringe in the extreme. For pete’s sake, she needs to show that her products are useful for something, beyond a scone and fingerprint cookies if she wants the crap to sell. You got it give it to her, she is trying, but she’s so out of her depth it’s as painful as listening to that tinny audio. Seriously, this woman is supposed to be a producer and can’t even get the audio right for a short Instagram video, which is both over produced and somehow under produced at the same time. It boggles the mind. 😵💫 She is well and truly running all of this and it REALLY shows.show more

Royal News Network
97,013 просмотров • 3 месяцев назад
True, even a two DOF system can exhibit chaotic... behavior, and a single cell is an at least an ~10^7 DOF system (i.e., protein molecules per cell). I still fail to properly convey to others the insane complexity I see in our microscopes, and why successes like AlphaFold only work because of the huge number of unnatural constraints placed on the training data. Case in point (below): lysosome dynamics over 12 min across a 200 x 70 um region in the head of a developing zebrafish embryo, color-coded by depth -- just one of 20k proteins at work. Our Cell Observatory Initiative is more important than ever, but while we have petabytes of the most mind-blowing data ever, we are still hamstrung by insufficient AI talent and compute -- two resources in great demand everywhere. If anyone can help us over this hurdle, we'd be eternally grateful.show more

Eric Betzig
29,753 просмотров • 5 месяцев назад
When we established a no-encampment zone in this corner... of San Jose, neighbors couldn’t understand why the city was able to keep the entire neighborhood clear — except for the most dangerous spot: along the railway. For months, hundreds of residents living in a mobile home park just beyond the fence still dealt with the impacts of unsheltered homelessness. Dogs barked at all hours, kids were jeered at and forced to stay indoors, and blight spread through the area. This community navigated layers of bureaucracy, worked through complicated jurisdictional boundaries, and demanded better. This week, my office partnered with Union Pacific to deliver it. It shouldn’t be this hard (or take this long!) to fix something like this. We need to take a hard look at the ways our systems put up unnecessary and frustrating barriers to action, and dismantle them. Because the bottom line is: people don’t care whose land it is. They want results — and we need to deliver them, faster. These new boulders will help keep the railway clear — for the safety of people still living unsheltered, and for the quality of life of the surrounding neighborhood. Thank you to UP for leaning in, and to the neighbors who advocated for better.show more

Mayor Matt Mahan
39,073 просмотров • 7 месяцев назад
llama.cpp isn't just for text LLMs anymore. Pure C++... zero shot voice cloning just officially landed in mainline. Text generation was only step one. If you’re building autonomous local AI agents, real time voice assistants, or edge workflows, instant low latency audio is the missing piece. Thanks to PR #26254, Alibaba’s state of the art Qwen3 TTS model family is now natively supported directly inside the llama.cpp repository under the multimodal (mtmd) framework. No Python bloat. No massive PyTorch CUDA overhead. Just raw, hyper optimized C++ running GGUF voice weights. Here is why this native update is a massive deal for the open source local AI stack: # Multimodal Architecture (.gguf + mmproj) Qwen3-TTS splits the workload between the base language model backbone and a multimodal projection adapter. llama.cpp handles this using the llama-tts binary, mapping the text model alongside its --mmproj projector to process audio tokens seamlessly. # Zero Shot Voice Cloning in Seconds You don't need fine tuning or massive dataset training. Feed the C++ engine a single 5 to 10 second .wav audio sample using the --tts-speaker-file flag, and it accurately clones the exact timbre, tone, and accent on the fly. # Real World T4 GPU Benchmark & Resource FootprintRunning the 1.7B Base model in 8-bit quantization (Q8_0): - VRAM Footprint: ~7 GB peak VRAM during active zero-shot cloning. - Audio Quality: Studio grade, natural-sounding voice output in seconds. • - Execution: Direct execution via native compiled binaries or sub process calls. # Coming Next to llama-server (PR #26603) Beyond CLI execution, a native POST /tts HTTP endpoint is currently being added to llama-server, which will soon allow you to trigger voice generation directly via standard REST API requests! # quick note on Colab compilation: Because this code was merged into mainline very recently, pre-built third-party binaries haven't fully caught up yet. Compiling llama-tts directly from source on Google Colab's free CPU instance can take about 1 hour (or ~1-2 minutes if targeting single GPU arch like -DCMAKE_CUDA_ARCHITECTURES=75). Be patient during the build step, or compile it locally on your own rig for instant execution! To test this out yourself, I built a zero config Google Colab notebook that compiles llama.cpp, downloads the Q8_0 GGUF files from HuggingFace, and spins up an interactive Gradio Studio UI so you can record/upload 3 second clips and clone voices in real time. Stop sleeping on native C++ audio. The era of bulky Python audio pipelines is officially over. Links to the free Google Colab notebook and the official ggml org GGUF HuggingFace model repository are in the replies below! available in q4 and q8 both variants, 1 GB and 1.85 GBs respectively (requires additional ~500MB mmproj gguf) Are you building local voice agents yet? What does your current audio stack look like? Drop your setups below!show more

Alok
47,881 просмотров • 1 месяц назад
A stunning reminder of why we cannot give up... on our rivers yesterday, as I stumbled across an adult eel on the Roding for the first time, lounging in the shallows in the shade of a council tower block & within earshot of the North Circular. This now rare & magical sight used to be common, until eel populations crashed on the Roding in the 1980’s & have not recovered. Seeking to understand & reverse this population crash should surely be a key role for the governments environmental regulator, but as usual they have done nothing to improve water quality or remove barriers to eel migration. Worse, they are actively blocking my efforts to help the eel population recover. Ordinarily, baby eels (elvers) for restocking are expensive to buy. However, I managed to secure a kind donation of elvers from fishermen on the River Severn (where the elvers often get stuck behind barriers on the river). I applied for my Environment Agency restocking permit like a good boy & all they had to do to help recover eel populations on the Roding was to say yes. Perhaps predictably, my application was rejected, because there was no positive evidence that reintroducing eels to the Roding would be a good thing. Perhaps most annoyingly, my application to restock eels on the Roding was rejected because reintroducing them would interfere with the EA’s monitoring of their continued decline. I asked what would happen if I went ahead & released the elvers anyway & was told that the EA would fine me up to £50,000. Yet another example of the malevolent uselessness of the EA: obsessed with procedure, but will do absolutely sod all to actually reverse the decline in our rivers.show more

Paul Powlesland
218,004 просмотров • 2 месяцев назад