Загрузка видео...

Не удалось загрузить видео

На главную

Most people use Linux every day. But Very few understand what actually happens after they press the power button. >>>Here’s the sequences Linux goes through before you get the login screen. Power button pressed → BIOS or UEFI initializes hardware and runs POST. Firmware locates the bootloader from disk....

35,042 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

🧃 Introducing stereOS: a Linux based operating system hardened and purpose built for AI agents. It's clear that agents need an ACTUAL operating system (not what people are calling an "OS") to witness the full breadth and depth of their capabilities while mitigating the blast radius of autonomous, untrusted actors. But there are so many problems with AI sandboxes today: * Going out to the apple store and buying a mac mini will never scale and is way too expensive (obviously) * Running in Docker is too restrictive (agents can't stand up their own container infrastructure, no sub virtualization, docker-in-docker is very broken) * Firecracker strips all the hardware so GPU PCIe passthrough, secure boot, FIPs, etc. is out of the question. * Native VMs are too fat and the overhead of 1 agent per VM is too much. stereOS takes a different approach: it's a full NixOS system that you boot and then kick off agent sandboxes inside with gVisor + /nix/store namespace mounting. Each agent gets their own kernel and the /nix/store is read only by nature. Even if the agent was somehow able to escape the gVisor virtual kernel, they'd land on the NixOS system as the "agent" user! Not your actual hardware!! If you want to take a defense-in-depth approach, we support "native" agents that run at the system level kicked off by our `agentd` utility. These agents, on their own, can manage and kick off other sub agents using the internal sandboxing mechanisms. Today, we're open sourcing all of this: * stereOS: our purpose built Linux OS - * masterblaster: client utility to launch, manage, and orchestrate agents - * stereosd: the stereOS system control plane daemon - * agentd: the stereOS system agent management daemon - Give it a try, throw us a star, and let me know what you think 🧃⭐️

John McBride

150,844 просмотров • 7 месяцев назад

How Did Pi Network Build a Huge User Base Before it Had a Fully Open Blockchain? Launched in 2019 by Stanford graduates, Pi (Pi Network) made crypto participation possible through a smartphone app instead of expensive mining hardware. But what exactly made its model different? (1) It made mining accessible through smartphones. Pi allowed users to earn Pi by checking into the app and starting a mining session. The process was designed to consume very little battery and computing power compared with traditional proof of work mining. This removed one of the biggest barriers to crypto mining, expensive hardware and high electricity costs. (2) It turned users into part of the network's growth engine. Pi introduced referral bonuses that encouraged existing users to invite friends and family. Users could also build security circles by adding people they trusted to the network. This created a built in growth mechanism that helped Pi expand its community rapidly. (3) Early Pi mining was not traditional blockchain mining. Before Open Network, users were not securing a public blockchain through conventional proof of work mining. Instead, the mobile app helped establish user participation and prepare the ecosystem before the network became publicly accessible. (4) Pi built the community before opening the network. Pi spent years operating through different development stages before launching its Open Network in February 2025. A firewalled Enclosed Mainnet had already gone live in December 2021, giving the project a working but closed blockchain years before it opened externally. During that period, the project focused heavily on growing its user base and developing its ecosystem. This flipped the usual crypto growth model, where networks often launch first and try to attract users afterward. (5) The Open Network changed what users could actually do. With Open Network, Pi became externally connected to the broader blockchain ecosystem. Users were now able to engage with Pi beyond the closed system it had been previously in, while developers were free to create applications on the network. This made the token and its massive community more realistic.

BSCN

43,622 просмотров • 18 дней назад

🚨 Anthropic committed up to 1M TPU chips for Claude. Openai is leasing TPUs for chatgpt inference. Here's How kernels work on TPUs (deep dive 2/6 by emilio andere) pallas is Google's answer to kernel writing. a python kernel SDK built on JAX. still very experimental (jax.experimental.pallas). on TPU it compiles through mosaic; on GPU it lowers to triton. if you know CUDA, the syntax will feel familiar but the execution model is completely different. in CUDA, grid=(4,4) launches 16 blocks running simultaneously across SMs. in pallas, those 16 iterations run one after another in lexicographic order. no threads. no warps. no blocks. no occupancy tuning. a TPU is a sequential machine with a very wide vector register — more like a CPU than a GPU. performance comes from width: a 128x128 systolic array doing matmul and an 8x128 SIMD vector unit doing everything else. maximum parallelism on chip: 2, one per TensorCore in megacore mode. three concepts replace CUDA's thread/block/grid hierarchy. Refs are mutable memory references. because execution is sequential, each iteration safely accumulates without atomics. in CUDA you'd need atomics or a separate reduction pass. the memory model is also very different from NVIDIA's. zero hardware caches. VMEM is 32-128 MiB of software-managed scratchpad — 500-1000x larger than GPU shared memory per SM. all data must be explicitly DMA'd from HBM to VMEM before any computation touches it. four levels: HBM → VMEM → VREGs → MXU/VPU, plus SMEM for scalar control data. every byte of data movement is your responsibility. this is like CUDA shared memory except it's 500x bigger and there's no cache fallback. pipelining is mandatory. without double-buffering HBM→VMEM transfers, the MXU just stalls waiting for data. this is the single most important optimization on TPU. and because grid execution is sequential and deterministic, consecutive iterations that need the same input block skip the redundant HBM transfer automatically, impossible on GPU where block execution order is undefined. the compilation pipeline is unlike anything in this series: python → jaxpr → stableHLO → XLA HLO (71+ optimization passes) → LLO (78+ passes) → 322-bit VLIW bundles. the compiler packs instructions for scalar, vector, matrix, and DMA units into a single 322-bit word. everything in that bundle executes in parallel, with no runtime scheduling.

wafer

33,655 просмотров • 2 месяцев назад

today was the first time i was genuinely impressed with what AI can do i recently decided to buy a whole FPV drone setup knowing basically nothing about the hardware side of it there's a pretty steep learning curve even just to set everything up properly: radios, RF protocols, flight controllers, ESCs, firmware, batteries, goggles, betaflight configs etc as someone that spends essentially 12h a day prompting agents to build software, it's actually pretty rare that i interact with AI on something where i have zero idea what's going on under the hood, and i never really used it for debugging a bunch of physical devices that all have to talk to each other i had codex + voice mode open for basically the entire setup. told it everything i bought, sent it some pics and then just started talking to it >what order do i set all this up in >how do i change this setting on the radio >which of these cables do i use >the drone is flashing pink wat mean >can you make this thing less insane to fly in my apartment and it was surprisingly seamless it would go find the manual for whatever specific thing i was holding, tell me exactly which buttons to press, what port to plug something into, what i should see if it worked etc then when i got to configuring the actual drone i had codex running on the computer it was plugged into, so it could inspect the config, back everything up, change settings, send usb reboot signals and check what happened the insane thing about voice mode is that youre literally hands on with the hardware and just telling codex what it should do, i literally never touched a thing on the computer besides starting voice mode if something doesn't work you tell it what happened and keep going a few hours of this and i had the radio, goggles, charger, batteries, drone firmware and betaflight all set up and had actually flown the thing the part that stuck with me is that i also understood what most of it was doing by the end, every time there was a term or tech i didnt understand id just ask to explain there is something absolutely magical about having proper real time personalized assistance, being able to dump a pile of unfamiliar hardware on your desk and have something figure out exactly what you own and walk through it with you in real time you become the missing physical link pressing the buttons i think spending all day using coding agents has actually made me pretty numb to AI progress. every new model is a bit better at some benchmark or can oneshot some task that the previous one couldn't and you just kinda adjust to it this felt different mostly because i had no existing knowledge to fall back on for the first time the jarvis comparison didn't feel cringe ai for coding and general computer tasks is cool and all but this feels a lot closer to the endgame anyone should be able to just ask any question about whats going on in their life and have realtime support i wonder if more hardware products will actually start exposing some sort of MCP or interface for agents to plug into thinking for example of how elevators in china are increasingly built with interfaces that let delivery robots call them directly instead of having to physically press a button we might actually start seeing hardware design shift from being purely human-interface-first to also being agent-interface-first buttons, screens and menus exist because humans need some way to tell machines what to do. agents don't necessarily need any of that if the hardware exposes an interface directly very curious which side closes the physical world gap first: humanoid robots that can operate hardware designed for humans, or hardware adapting so agents can operate it directly

ultra

18,253 просмотров • 1 месяц назад

U.S. Patent 6,506,148 is titled "Nervous System Manipulation by Electromagnetic Fields from Monitors.” It describes a method for influencing a human subject's nervous system through the use of electromagnetic fields emitted by devices like computer monitors and television sets. 📡📺💻 The inventor, Hendricus G. Loos, suggests that pulsing the images on these displays at specific frequencies can trigger physical responses through sensory resonance. ⚠️ These subtle pulses can be integrated directly into video content or layered over an existing signal to interact with a viewer's skin. 📹📱⚡ Remarkably, the technology is designed to function even when the visual fluctuations are subliminal, meaning they are too faint for the user to see or consciously perceive. 👀 Ultimately, the system aims to remotely manipulate human bodies by using everyday hardware as a transmission device for electromagnetic stimulation. And the scariest part? This patent has been public since 2003. So why is this important to understand? Your nervous system runs the show. 🧠⚡ It controls your immune system, hormones, digestion, detox, and sleep. Every healing response starts there. If you’re stuck in stress mode, your body stays in survival and cannot repair. 🚨 Healing isn’t just chemical. It’s electrical. ⚡🧬 You can eat clean and take supplements, but if your nervous system is overloaded, healing slows down. That alone should make us more mindful about nervous system health. Excess screen time, stress, artificial light, and constant stimulation can disrupt sleep, mood, and physiology. 😴📱🌙 Instead of fear: 📵 Limit screens before bed 🌙 Use blue-light filters ☀️ Prioritize sunlight, grounding, and sleep 💧 Support detox pathways 🛡️ Strengthen your biology 🧠 Awareness matters. Discernment matters more. Don’t let screen time hijack your nervous system. Blessings and truth, Dr. Edward Group, DC #NervousSystemHealth #EMFAwareness #DigitalDetox #BioelectricBody #HealthResilience #Grounding #SleepHealth #HolisticLiving #DrEdwardGroup

Dr. Edward Group, DC

70,577 просмотров • 7 месяцев назад

Stanford researchers did it again. They just built the agent-native version of Git. When an agent works on a longer task, the run builds up a lot of state. This includes files edited/created, a dev server, a database, installed packages, KV cache, etc. Say the agent is at step 10 and makes a mistake, maybe it misreads a traceback and rewrites a file that was actually fine. The tests start failing, and the run goes off track, although everything through step eight was correct. By default, the agent just tries to fix it, which creates more edits and tool calls. This burns more tokens and grows the context. The other options are a person stepping in to redirect it or restarting the whole run from step one. That's wasteful, because it pays for every model/tool call again and re-prefills the context. Moreover, since an agent's run is non-deterministic, it doesn't reproduce the same early steps anyway. The reason it's hard to just jump back exactly to a previous correct step and resume from there is that the trajectory is only a message log. It records what the agent said and which tools it called, but not the live state underneath. That state includes things like memory, open file handles, child processes, installed packages, /tmp, and KV cache. None of that is in the log. Git can version the files, but it doesn't snapshot the running process or the KV cache. Checking out step eight moves the files back, but the process is still sitting in step-ten memory with a cold cache. Shepherd is a runtime layer by Stanford that records the run as a trace of typed events rather than a flat log. Each agent-environment interaction becomes a commit, similar to Git, but it tracks the live run. Its commit includes the agent process and the filesystem together, copy-on-write, so a branch carries the actual state and not just the files. Going back to a previous step is then a single call that forks from that commit and continues from the exact state. The copy-on-write fork is roughly five times faster than docker commit, and because the prompt prefix through step eight is unchanged, the KV cache is reused over 95% on replay, so early steps aren't reprocessed again. Once the run can be forked, a meta-agent can sit on top and operate it. It watches the trace and reverts as soon as it looks wrong, before the bad write is committed. In practice, it's just Python calling fork, replay, and revert on the trace, rather than a separate control plane wired into the harness. Not everything is reversible though. Files and sandbox changes undo themselves, but a database write has no automatic undo, so it needs a matching undo step set up in advance. Something external, like a sent email or a real charge, can't be undone, so the supervisor's job there is to catch it before it fires. They tested this on a few public benchmarks. On CooperBench, where two agents work on the same codebase, adding a live supervisor took the pair-coding pass rate from 28.8% to 54.7%. It's still early and labeled alpha. The benefit mostly shows up when a run gets branched a lot over a heavy sandbox state, which is exactly where restarting wastes the most tokens and time. If Git was made to make file changes reversible, Shepherd is trying to do the same thing for a live agent run. Shepherd Repo: (don't forget to star it ⭐ ) That said, Shepherd reverts a bad step inside a run. The harness around it, the prompts, tools, and checks the supervisor relies on, still drifts across runs as models and dependencies change. Akshay wrote about making that harness repair itself, where a failing trace gets diagnosed, the fix is verified against the exact input that failed, and the failure is locked as a regression test so it can't recur. Read it below.

Avi Chawla

441,974 просмотров • 2 месяцев назад

The next iPhone will cost more, and the reason has almost nothing to do with Apple. The chip that stores your photos cost Apple about 13 dollars last year. This year it runs around 51. Multiply that across every phone, laptop, and console on earth, and you are looking at the first consumer bill for the AI boom, arriving in the pocket of someone who never asked for it. Tim Cook, who has run Apple's supply chain for forty years, called it a hundred-year flood, something he has never seen. Memory prices have quadrupled in places. The cause is brutally simple. AI data centers are now expected to swallow roughly 70 percent of the world's memory production this year. Seven chips in ten go to server farms. Phones, cars, and laptops fight over the three that are left. This is one force wearing two faces. The same AI demand making the device in your hand more expensive is minting record fortunes for the handful of companies that feed it. Memory makers in Seoul just hit all-time highs in the same week Apple warned you to brace for higher prices. The shortage and the windfall are the identical event, seen from opposite ends. Then comes the part almost no one traces all the way down. Beneath the chips sit rare earth minerals, and one country controls them. China processes around 90 percent of the world's rare earths and makes roughly 94 percent of the high-performance magnets that spin inside every fab and cooling system. The polishing compound that finishes a wafer, the magnets in the machines that build it, run through Beijing. And through 2025, China has been turning that grip into leverage, licensing what leaves. So the chain is complete. AI wants memory, memory needs minerals, and the minerals answer to one government. The price of your phone is now a foreign policy.

Shanaka Anslem Perera ⚡

58,900 просмотров • 3 месяцев назад

Love and Deepspace | Rerun Event Preview The 5-Star Rate UP Pool [Twilight Serenity] Limited-Time Rerun will start soon! "Then, you're not allowed to change your mind even after a hundred years. Or a thousand." 💫Event Duration: From 05:00 on Dec. 11 to 04:59 on Dec. 18 (Server Time) 💫The 5-Star Rate UP Pool [Twilight Serenity] Limited-Time Rerun Event 1. During the event, make a wish with [Deepspace Wish] or [Time Wish: Limited] to participate in the wish event. The drop rate of the event-limited 5-Star Memory [Rafayel: Fireworks Vow] will go up drastically. 2. After the event ends, this limited 5-Star Memory will not be obtainable through other means and will not enter the permanent Wish Pool: Xspace Echo. 3. All Rerun Wish Pools share one pity system. A 5-Star Memory is guaranteed within a specific attempt of wishes. If the 5-Star Memory you have obtained from the Rerun Wish Pool is not the event-limited Memory, you will obtain the event-limited Memory the next time you obtain a 5-Star Memory. The pity count from the last Rerun Wish Pool can be applied to this Rerun Wish Pool, and the pity count in this Rerun Wish Pool will also be applied to the upcoming Rerun Wish Pool. *You can read more about the event on the in-game rules page. 🎁New Packs During the event, the Rerun event-exclusive [Flamebloom Pack] series, which includes [Time Wish: Limited] and other materials, will be available in Shop. Notes: 1. [Time Wish: Limited] can be used in 5-Star Memory Wish Pool Rerun and will be used first when you make a wish. 2. After the event ends, [Time Wish: Limited] will automatically convert to Empyrean Wish. ——— 🪐Official Discord: #LoveandDeepspace #Rafayel

Love and Deepspace

304,066 просмотров • 9 месяцев назад

This guy built a visual scanner that reads 468 points on his face and 42 points on his hands from a regular webcam and turns them into a cloud of thousands of particles right between his palms. Inside, MediaPipe and TouchDesigner are linked: the first captures hands and face from the webcam with high accuracy, the second turns those coordinates into a live plane and feeds it into a POP system that instantly generates a swarm of particles in the shape of a head. No studio, no render farmer, no VR headset. Just a laptop, a webcam, and 1 TouchDesigner session. And traditional VJ studios keep teams of 5 people on a setup with lighting, custom hardware, and commercial plugins, while his expenses are only a TouchDesigner subscription and a regular USB camera. One laptop runs MediaPipe and TouchDesigner simultaneously, holds the camera stream at 60 FPS without drops, and in parallel processes 468 face points + 21 points on each hand. The camera captures frame after frame, MediaPipe in real time sends TouchDesigner the finger coordinates and face geometry, and the POP operator inside the engine translates those numbers into thousands of particle points with colors from bright pink to gold. This setup immediately defines the role of the tool and the limits of its autonomy. It knows where the fingertips are at every moment of the frame. It knows how to read the face geometry at any angle to the camera. It knows how to draw a swarm of particles between them with the right color and contour. → MediaPipe pulls 468 points from the face and 21 points from each hand, 60 times per second → TouchDesigner receives those coordinates, builds a virtual rectangle between the fingertips, and feeds it into the POP system → POP generates thousands of particle points in the shape of a head, coloring them in a gradient from bright pink to gold → The HUD layer adds green corners and a blue neon frame, styling the image like an AR interface → All layers assemble into 1 real-time frame that projects back onto the video in the camera window → The final image is recorded to a file or broadcast to a projector for a live installation And only when the guy spreads his hands wider does the plane between the palms stretch; brings them together, it narrows. Otherwise the system runs on its own. And when he moves from his home room to a concert hall, the same laptop with the same webcam launches the same TouchDesigner session in just 5 minutes, without reconfiguration, without a new team, and without a single line of new code. In his work setup there is no studio of his own and no team for assembly. On the desk sits a laptop with a webcam, on top run MediaPipe and TouchDesigner with POP operators, and the same setup through a USB camera moves to any concert without a new configuration. Out of everything I have seen this year, this is the cleanest Creative Coding setup on 1 laptop: 0 render farms, 0 studio lighting, and between them 3 libraries, thousands of particle points, and 1 webcam.

Blaze

38,242 просмотров • 4 месяцев назад

Shipped. Minimax H3 EZLaunch. One install script. Full optimized stack for RTX 3090 and RTX 4090. Windows AND Linux. Ships with every model you need. T2V, I2V, and Ref2V all supported. 5 second clips in about 90 seconds. 15 second clips in under 8 minutes. Zero OOMs across 80 minutes of continuous generation. Same stack that held clean across all 10 space battle renders. What it actually does: - Downloads and configures the correct weight set automatically: pruned INT8 UNET with ConvRot, Turbo LoRA (8-step quality band), Qwen3-VL text encoder in NVFP4, MiniMax video + audio VAEs. No hunting through HuggingFace for the right quants. - Full mode coverage. Text-to-video out of the box. Image-to-video with first/last frame. Reference-to-video for style-guided generation. O - Patches SageAttention v2 sm89 to Triton INT8. - Enables kitchen CUDA with FORCE_CUDA on torch cu128 + driver 580 - Full flag set applied: lowvram, disable-smart-memory, disable-pinned-memory, disable-cuda-malloc, fp16-intermediates, CLIP on CPU - Timed smoke test runner included. Verify your stack without guessing. - Cross-platform installer. Same script, same results, Windows or Linux. The weight selection alone saves people hours. The Sage patch is the difference between launch-fail and sub-90 second renders on Ada. Nobody had either documented publicly. Now it is one command. Try it on similar cards and tweak if you want. MVP. More GPU configurations coming. Stay Tuned. Repo link in reply.

Yume_X

13,279 просмотров • 1 месяц назад

“There is a tension between what the users of a currency want – and the users of a currency tend to like freedom, autonomy, and discretion as to what they spent their money on – and what the issuers of a currency want; and bluntly, the issuers of a currency want control. Control of monetary policy, and control of you.” The Bank of England’s consultation papers make very clear the level of control that they wish to exercise over you, and over your supposed financial autonomy, if you were to use their #DigitalPound. (1) You’ll need to provide ID in order to use the #DigitalPound: “For the digital pound, tiered access would allow for different levels of user access and functionality based on the amount of identification (ID) a user is willing or able to provide.” (2) The Bank will dictate how much you can hold: “The Bank would place some limits on holdings of digital pounds, at least during its introductory period.” (3) The digital pound will be programmable, if not by the Bank itself then by third party providers: “Programmability, delivered by Payment Interface Providers, could also enable the use of smart contracts, which carry out specific actions based on pre-defined terms and conditions.” Quotes are from from the Bank of England’s Digital Pound consultation paper: Whatever the #DigitalPound will be, it won’t be cash. Cash does not require me to show ID to use it. I can hold as much cash as I want or need. And, along with #Bitcoin, cash is a bearer instrument whose title is freely transferrable upon delivery, which is very difficult for a central bank to control. And long may it stay this way. A huge shout out and thank you to Lyn Alden, who made this point much more eloquently than I did in her excellent book #BrokenMoney. Thank you! Also I’m aware that my hand gestures in this clip are reminiscent of Richard Hendricks manipulating ‘datas’ on stage at TechCrunch Disrupt in #SiliconValley, and for this I can only apologize: #BitcoinConference #Amsterdam #NoToCBDCs

Freddie New

13,163 просмотров • 2 лет назад

wait what?? Seedance 2.0 went crazy with this prompt 🦾 This is real early 2000s home video footage shot on a VHS camcorder at a crowded public swimming pool. The video has the typical grain, color, and soft quality of consumer video from that era. The camera is very shaky and handheld, filming from the side of the pool. On the diving board there is a very fat man wearing a homemade jetpack made out of what looks like metal tubes and bottles strapped to his back. A group of people around the pool are cheering and laughing, hyping him up. The guy runs and jumps off the diving board, then activates the jetpack mid-air. He manages to fly a few meters forward over the pool, but the jetpack starts struggling under his weight. He slowly loses height and ends up crashing into the water with a big splash. Right after he falls in, another fat guy with a mullet haircut jumps into the pool holding a can of beer. He swims over to the guy with the jetpack, who is still half-floating with the device on, and hands him the beer. The jetpack guy takes it and starts drinking while still in the water. Everyone around the pool goes crazy cheering and laughing at the whole scene. The camera movement is extremely shaky and reactive the entire time, with constant motion, motion blur during the jump and the crash, and the typical imperfections of someone filming with an old camcorder in a crowded place. There are several fast, unplanned cuts as the person filming tries to follow the chaos. Natural sound only: people cheering and laughing, the sound of the jetpack, the big splash when he crashes, and general pool ambience recorded with the camcorder microphone. The result must feel like authentic raw early 2000s home video of someone casually filming an absolutely ridiculous moment at a public swimming pool. Now is your turn! Tweak it and share it 🫡

TechHalla

34,084 просмотров • 2 месяцев назад

this is f*cking beyond comprehension. Google engineers just shipped the entire agent lifecycle in one release: build, scale, govern. and every piece answers a specific way agents die in production > context layers (Static, Turn, User, Cache): you decide what the model carries between turns, so token spend stops being a mystery > a self heal plugin: the agent notices a tool call failed and retries it a different way instead of dying mid run > adk deploy: one command from your laptop to the managed runtime, no packaging, no infra ticket > Go joins Python and Java, with its own A2A SDK then the part nobody builds for themselves: > a dashboard on token consumption, latency, error rates and tool calls: the four things that actually kill an agent > a traces tab that opens the real sequence of actions the agent took, step by step > a playground wired to the deployed agent, past sessions included, so debugging is not a redeploy loop > an Evaluation Layer with a User Simulator, because you cannot unit test a non deterministic system and the part that decides whether it ever ships: > agents get native identities as first class IAM principals: least privilege applies to them like it does to people > Model Armor screens prompt injection, tool calls and responses, inline for Gemini or over REST > Security Command Center inventories every agentic asset and flags data exfiltration by an agent ADK is already at 7 million downloads. the runtime has a free tier, and express mode runs off a Gmail address. the prototype was never the hard part.

NO1ennn

24,666 просмотров • 16 дней назад

Running cold email campaigns just became a whole lot easier Smartlead now runs an MCP server, which in plain terms means Claude can read and act on your live campaign data directly instead of working off a spreadsheet that went stale the moment you exported it. The workflow is worth walking through properly, because it is shorter than people expect. You generate an API key inside your account, point Claude at the server once, and from then on you ask for what you want in a sentence. Here is a prompt worth stealing in full: "Fetch all Smartlead clients, then get today's performance for each: emails sent, replied, positive replies, unique lead count. Compute reply rate per client, run a top and bottom performer analysis, format it as a daily client performance report, and post it to Slack." One paste, and it pulls live figures for every account, does the arithmetic, ranks the strongest and the weakest, and delivers the finished thing into the channel your team already sits in, before anyone has logged on for the day. Be clear about the division of labour, because it is what makes this useful rather than a novelty. Smartlead is the engine holding the campaigns, the mailboxes, the warmup and the reply data, and Claude is simply the interface you operate all of it through, so nothing about your sending changes and everything about how you interrogate it does. The effect people underestimate is on the questions you start asking. Once a report costs you a sentence rather than an afternoon, you stop rationing the ones that used to feel like too much trouble, and problems that used to surface on a Friday start surfacing on a Tuesday. Connect it with Claude through MCP and run one prompt against your own account today.

Tim

21,666 просмотров • 14 дней назад

You don't need a GPU for fast studio grade voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?

Alok

105,756 просмотров • 1 месяц назад