Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Just like the classic Notepad Ctrl+Click RCE behavior, terminals on Windows, Linux, and macOS also support clickable file/URI handlers. printf "\x1b]8;;file:///C:/windows/system32/calc.exe\x07Click here\x1b]8;;\x1b\\n"

19,802 görüntüleme • 4 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

You don't need a GPU for fast studio grade voice cloning anymore. Qwen3 TTS (1.7B Q4_K_M) + mainline llama.cpp is officially the fastest way to generate zero shot voice clones using 100% pure CPU execution. Following up on my last post where we ran the Q8 model on a GPU, we just took local C++ voice synthesis a massive step further. The open source community quantized Alibaba's SOTA Qwen3 TTS model down to Q4_K_M GGUF, completely freeing local audio pipelines from dedicated graphics hardware. Here is the real world benchmark and hardware breakdown of running SOTA voice cloning on CPU: # Architecture & Model Setup Using Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf paired with the 8 bit multimodal projector (mmproj-Q8_0.gguf), llama.cpp executes the entire pipeline in pure C++. No PyTorch, no CUDA dependencies, and no VRAM bottlenecks. # Real-World Memory Footprint - Baseline RAM: 1.6 GB system idle. - Peak Generation RAM: 8 GB RAM during active voice synthesis. - Requirement: Any basic machine with at least 8 GB of system RAM can run this easily. # Real World CPU Benchmarks - Google Colab Free Tier (Throttled 2 Core CPU): Synthesizes a 5 sec studio quality audio clip (~8 words) in 45 seconds. - Modern Consumer CPU (Intel i5/i7 13th/14th Gen or AMD Ryzen 7000/9000): generation should drop to 5 to 20 seconds (nearly 1:1 real-time generation speed!). # Zero Shot Voice Cloning Quality Pass any 5 to 20 second .wav audio sample to the C++ engine using the --tts-speaker-file flag. It yields clean, natural sounding cloned speech with virtually zero quality loss compared to unquantized FP16 weights. To make testing seamless, I built an updated zero config Google Colab notebook. It pulls the official pre built llama.cpp CPU binaries (zero compilation time!) launches a live Gradio web app right in your browser. Record a 5 second clip from your mic (or drop a .mp3, .wav file), type text, and generate cloned audio on CPU. Native C++ audio models are making edge based, offline AI voice agents a reality. Links to the free Q4 CPU Colab notebook and the Q4_K_M GGUF HuggingFace repository are in the replies below! Which models have you been running on your CPUs? What CPU hardware are you using for local inference?

Alok

105,756 görüntüleme • 1 ay önce

Run Gemma 4 26b MTP on 8 GB VRAM GPUs at 25+ tokens/second. Flags included! local llm space is moving at terminal velocity. only 3 days ago google released gemma 4 26b a4b qat quants. more efficient than before, ran on 8gb vram at 20 tok/sec. and now just a few hours ago, mainline llama.cpp merged a massive update and we just shattered our own record. decode throughput went 25-40% up on the same 8 GB VRAM setup! Before MTP: 20 tps -> After MTP: 28 tps! llama.cpp just officially merged PR #23398 ("add Gemma4 MTP"), bringing native Multi-Token Prediction (MTP) support to Gemma 4 models. By running speculative drafting on the same 8GB VRAM RTX 4060 setup, my decode throughput on a 64k context instantly leaped to a blistering 25–27 tokens/sec thats 25-30% increase with the same hardware. Here is the architectural catch you need to know: Unlike the Qwen 3.5 and 3.6 series, which bake the MTP heads directly into the base GGUF, the Gemma 4 MTP head is not built in. You must download a separate, specialized MTP drafter GGUF (the assistant model) to act as the speculator. (I've dropped the download link in the replies). copy and try the exact flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.7 --spec-draft-model gemma-4-26b-A4B-it-assistant-Q4_0.gguf -c 64000 -v n-max 4 and p-min 0.7 is also worth checking out. benchmark on your setup and workflow. if you have a single 8 gb vram nvidia rtx 4060, 3060, 3070, 2080, 2070, grab the MTP drafter GGUF link in the comments and try it yourself. Check it out even if you have asmaller or a larger gpu, such as a single rtx 3090, 4090, 3060, 2060. MTP works for all gemma 4 sizes such as gemma 4 12b, gemma 4 31b etc. but remember to grab the correct mtp draft assistant models respectively. what are you benchmarking today

Alok

200,913 görüntüleme • 3 ay önce

1️⃣ I did everything I could for Ritual I made all kinds of contributions. For 8 months, I contributed to Ritual, but I am extremely disappointed with their system. I made 181 contributions on X, bought Canva Premium to design infographics, and sent almost 17,000 messages. Despite all this, I have been permanently ignored for 8 months. Two of my friends, Ashish.Base.eth (❖,❖) (230 contributions) and KUNDAN (150 contributions), are facing the same treatment. Because of this unfair and favoritism based system, I decided to post on Twitter to expose it. When we raised our voice after 6 months of hard work and 150+ contributions, the so called Indian community handlers and famous team members labeled it as FUD and continued ignoring us. 2️⃣ When we asked the Indian community handlers why we were not given roles after 8 months, they said: “We can’t do anything, the team decides.” But when we contacted the team or mods, they said: “Go back to the Indian community handlers, they will help you.” If these handlers and big mods really can’t do anything, then why are they even here? They are here to build connections This system has been making fun of me and other hard working contributors. 3️⃣ Now look at the people who received roles based on “quality contributions.” Everything is available with proof. One person (Yaneul) got a role with only 15 very low quality contributions (mostly food and event screenshots). He even got Ritualist with just 25 contributions. This clearly shows either connections were used or it could be a mod’s second account. Another person (placboeffect) got a role with just 9 contributions. Many people received roles at 35, 40, 50, 55, 60 contributions, while I had 181 contributions. This is what they call “quality contributions.” 4️⃣ Now decide yourself: Who contributed more? Who worked harder? People contributing for 8 months to 1 year are ignored, while others get roles at 9 or 25 contributions. Yes, we may have wasted our time, but we did work hard. 5️⃣ I didn’t want to come to X. but after contributing like a mad person and getting ignored, I don’t want to waste my time anymore. 6️⃣ I don’t know if the Josh (❖,❖) admin is aware of how mods and team are damaging this project Claire (❖,❖) she just help her country mates doesn't care about other people. 7️⃣ What’s the worst that can happen? They will ban me, and my 8 months of contributions will be wasted. There was never respect for real contributors anyway. If I had good connections with mods, I wouldn’t have struggled this much. It’s not like Ritual is my life, or Web3 has only Ritual, or they gave me any role to take back. Now I will openly speak on X without fear. 8️⃣ Some fearful people will read this post. They know everything that’s happening, but they won’t comment because they are scared. I understand their fear. --- 9️⃣ Best of luck to those who keep working blindly despite seeing all this. And special luck to those who have roles because of good connections with mods. --- 🔟 Last The so called Indian community handlers are only mods by name. One of them is a paid event manager, doing the job only for money. Another handler sends one message after 10 days. These are the people leading us in Ritual, who don’t even know what they’re doing. Jez ritual/acc (❖,❖) dunken(ritual/acc) (❖,❖) Stefan | Mad Scientist (❖,❖) Hinata Sir W A R D E N you are mod in Donut and Dhillon saab π² you are mod at data haven, Choudhary (Ø,G) ꧁IP꧂ you are mod at Capx and sir Botsan (capx arc) you are Admin of Capx You should learn from ritual mods how to ignore real contributors and how to promote favourtism

Legend

12,873 görüntüleme • 7 ay önce

Another WTF moment. A developer just open-sourced a coding agent harness that boots 245x faster than Claude Code. It's called jcode. You launch it and the first frame renders in 14 milliseconds. Claude Code takes 3,436. One active session uses 27.8 MB of RAM. Claude Code uses 386.6. Run ten sessions in parallel and jcode holds at 117 MB while OpenCode swells to 3.2 GB. Each agent has a semantic memory graph instead of a scratchpad. Every turn gets embedded as a vector. The graph is queried on every turn for related memories, and a sideagent verifies the hits before injecting them into context. Consolidation runs in the background to check for stale or conflicting facts. No manual /remember calls. No token burn on lookup tools. The provider list is 30+ deep. Claude, ChatGPT, Gemini, GitHub Copilot, Azure, OpenRouter, DeepSeek, Groq, Mistral, Perplexity, Fireworks, Ollama, LM Studio, and any OpenAI-compatible endpoint you point it at. Ran out of tokens on your first ChatGPT Pro sub? /account swaps to the second. Then there's Swarm. Spawn two agents in the same repo and the server manages them. When agent A edits a file agent B has been reading, agent B gets pinged and can check the diff. Agents can DM each other, broadcast to the room, or spawn their own worker teams for parallel tasks. Groups, channels, and completion statuses are handled automatically. The UI has live side panels that render mermaid diagrams inline. To make it fast, the author wrote a Rust mermaid renderer 1800x faster than the JavaScript one, then wrote a custom terminal called Handterm because no existing terminal could do smooth partial-line scrolling. Self-dev mode is where it gets wild. Tell your agent to enter self-dev and it starts editing jcode's own source code, rebuilds the binary, reloads it live, and keeps working across your existing sessions. You can also resume broken sessions from Claude Code, Codex, OpenCode, or pi directly inside jcode. Anthropic's cache goes cold at the 5-minute mark and you're staring down a big cache miss on your next turn? The UI warns you before you spend the tokens. Written in Rust. MIT licensed. Runs on macOS, Windows, Linux, and Termux. Sitting at 11.2k stars with a native iOS app coming.

Brady Long

207,858 görüntüleme • 1 ay önce

BREAKING: SpaceXAI has released a major new update for Grok Build (v1.0.14) Grok v1.0.14 is a reliability and workflow update for the Grok CLI. It makes OIDC token refresh proactive, lets PostToolUse hooks send feedback back to the model after tools run, adds per-turn token and cost tracking via grok usage, and cuts Windows downloads by about 70%, with a large set of fixes that tighten sandboxing, subagent handling, hooks, and startup. Features: • OIDC token refresh is now proactive by default for better reliability. • PostToolUse hooks can now provide feedback and context to the model after tool execution. • SDK-registered PostToolUse hooks now provide model-facing feedback. • grok usage now shows persisted per-turn token and cost data. • Retry status in composer and title now shows a short reason for the retry. • Models can now declare a different identifier for each reasoning-effort level instead of always sending the same id. • Prompt suggestions now respect remote configuration and default to the current session model. • Windows CLI downloads are now ~70% smaller using the same compressed sidecars as macOS and Linux. Bug Fixes: • grok inspect now correctly shows Claude bypass locks as advisory rather than enforced. • Subagent sessions no longer leak threads or file descriptors when the parent is busy. • Cold startup no longer performs duplicate remote settings fetches. • Compaction failures due to context size now degrade input instead of retrying identically. • --sandbox strict now restricts writes to ~/.grok/sessions only. • Subagent spawning now waits longer on a busy coordinator and shows clearer retry guidance instead of "unreachable". • Failed task and todo tool calls now appear in the transcript instead of disappearing without a trace. • Composer status row no longer collapses or flashes when using double-Enter to send now. • Session close is no longer delayed by a single slow hook; each SessionEnd hook now has its own timeout. • Hook removal in the extensions modal no longer offers actions that the handler will refuse. • Interjections during a turn are now delivered atomically or not at all. • Subagent tasks no longer get incorrectly cancelled when the parent session is waiting for completion. • Workflow detail view now closes the overlay on X or outside click instead of returning to the run list. • Resuming subagents now succeeds for larger transcripts that still fit the model context with headroom. Performance: • Startup now fetches remote settings only once per boot instead of potentially twice. • First message on large repositories no longer waits on repository status scan. • Large session memory no longer blocks the agent during turn completion or subagent spawning. • Signed-in CLI starts faster by serving remote settings from a local cache on warm boots. Download Grok Build: Update to the latest Alpha release: grok update --alpha Update to the latest Stable release: grok update

DogeDesigner

52,351 görüntüleme • 24 gün önce

I told you to claim your free 16GB NVIDIA GPU for learning Local LLMs. Now I’m going to show you how to double its inference speed without touching the hardware. Google Colab gives you an enterprise grade NVIDIA Tesla T4 GPU for free, roughly 4 hours every single day. It is the absolute perfect sandbox for learning AI engineering, testing inference flags, and pushing massive context windows. The local AI timeline is moving way too fast. If you aren't using Multi Token Prediction (MTP) yet, you are leaving massive performance on the table. I just pushed DeepMind’s Gemma 4 26B to 64.9 t/s on this exact free tier. Let's look at the raw benchmark data running on an Ubuntu Linux environment with the latest compiled llama.cpp binaries and quantized GGUFs from Unsloth via HuggingFace: # Qwen 3.5 9B (Dense): Base: [ Prompt: 626.7 t/s | Generation: 21.0 t/s ] With MTP: [ Prompt: 539.1 t/s | Generation: 24.8 t/s ] # Gemma 4 26B QAT (MoE): Base: [ Prompt: 634.2 t/s | Generation: 48.3 t/s ] With MTP: [ Prompt: 572.1 t/s | Generation: 64.9 t/s ] If you are paying attention, this single Colab notebook reveals 3 massive observations about the current state of local LLMs: # 1. The MTP Speedup (Software Overclocking) Standard autoregressive decoding guesses one token at a time. MTP acts like a highly optimized, built in speculative decoder. It predicts multiple future tokens at once and the main model verifies them in parallel. The result? Zero accuracy loss and a massive throughput increase. Gemma jumped from 48 to 65 t/s just by flipping a flag. # 2. The MoE Paradox (Bigger is Faster) How does a 26B parameter model absolutely destroy a 9B model in raw speed on the exact same hardware? Architecture. Qwen 3.5 9B is a dense model. it activates all 9 billion parameters for every single token. Gemma 4 26B is a Mixture of Experts (MoE) model. It routes data efficiently, activating only 4B parameters per token. You get the reasoning capabilities of a 26B model with the compute cost of a 4B model. 3. Thinking Efficiency When I ran the exact same complex prompt on both models, the larger MoE spent significantly fewer "thinking" tokens to arrive at the correct answer. A smarter model doesn't just give better answers; it gets to the point faster, saving you compute cycles and preserving your context window. # Want to run this yourself? Here are the exact llama.cpp CLI commands. For Qwen (MTP is baked into the main model): ./llama-cli -m Qwen3.5-9B-UD-Q4_K_XL.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 For Gemma (Using a separate lightweight draft model): ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --model-draft mtp-gemma-4-26B-A4B-it.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Stop waiting for a $3,000 rig. Boot up Colab, pull these models, and start building your stack. I’ve put together a completely free, cell by cell Google Colab notebook that automates this entire workflow so you can test it yourself in 5 minutes and learn. Link to the notebook is in the comments below. Experiemt with different MTP parameters, context windows and post your results in the comments.

Alok

170,442 görüntüleme • 2 ay önce

AgentLinter is here! Is your agent sharp & secure? I built AgentLinter, a linter for and agent config files. Here's why. Whether you're vibe-coding or agent-coding, your AI's output quality comes down to one thing: how well you wrote your But managing these files properly? Way harder than it looks. 🎯 The Silent Failure Problem Vague instructions like "write good code" let the agent interpret however it wants. Output gets inconsistent, but nothing throws an error. The failure is silent. Anthropic's own docs say write "Use 2-space indentation" not "Format code properly." But as the file grows, spotting these with your eyes alone is nearly impossible. 🔐 The Security Problem People hard-code API keys and tokens directly into or and commit them, way more often than you'd think. AgentLinter stats show 1 in 5 workspaces has exposed credentials. .gitignore doesn't catch secrets buried inside markdown files. 💥 The Consistency Problem Multiple config files = contradictions. says "be a friendly assistant," says "concise, direct tone." The agent gets confused. references files that don't exist. Past 5 files, these conflicts triple. So I thought: is code. Code has ESLint. Why doesn't this have a linter? 🔍 What AgentLinter Does It diagnoses your agent config across 8 categories: 1) Structure: file organization 2) Clarity: instruction specificity 3) Completeness: missing definitions 4) Security: exposed secrets 5) Consistency: cross-file contradictions 6) Memory: session handoff 7) Runtime Config: gateway/auth settings 8) Skill Safety: dangerous shell commands & injection patterns Each scored 0–100 with concrete fix suggestions. Write "be helpful" and it tells you to specify response length, tone, and format. Find an API key? Instant CRITICAL alert to rotate. 🔒 Privacy-First & 100% Local Everything runs on your machine. Files never leave. Only the results are shared, and you can turn that off in settings. This matters — these files can contain system prompts, security rules, and personal context. Fully open source, MIT license, 100% free. 🛠️ Multi-Tool Support Works with Claude Code, Cursor, Windsurf, and Clawdbot. Detects for project mode, or clawdbot.json for agent mode and adjusts diagnostics automatically. 🚀 Get Started with one line npx agentlinter Node.js 18+, no config needed. Run it, check your score, fix what needs fixing. Happy vibe-coding & happy agent life! 🤙 Website: Github:

Simon Kim

48,629 görüntüleme • 7 ay önce

Release: LichtFeld Studio v0.5.3 is out! With 316 commits merged into master, this release is a huge step forward for LichtFeld Studio. What's new in v0.5.3 • Vulkan viewer/rendering migration: New Vulkan viewport pipeline, pass graph, VkSplat renderer, Vulkan point-cloud renderer, 3DGUT/VkSplat support, improved alpha/depth composition, tighter CUDA/Vulkan interoperability, and device matching on multi-GPU systems. • RAD + LOD workflow: Added RAD file export/import, RAD LOD viewer, Spark-style GPU LOD selection, GPU-driven page prefetching, a bounded VRAM pool, out-of-core PLY-to-RAD LOD conversion, and RAD import/export speedups of approximately 3–5×. • HiGS / macro-tile inference: Added a macro-tile inference path for the Vulkan viewer, including macro sorting, batched rasterization, composition, and capacity management. • Asset Manager: Added and significantly enhanced the Asset Manager with thumbnails, SH information, faster synchronization, import-from-URL support, docked mode, data-loading popup integration, and general UI cleanup. • Viewport export: Integrated viewport export directly into the application as a toolbar/overlay tool, added fast render_view_u8-style readback paths, fixed high-resolution clipping issues, improved orthographic export parity, resolved 32K image/video export problems, and added post-export GPU resource cleanup. • Selection and tooling: Added and reworked selection toolbar controls, the Select menu, ring selection, color eyedropper, distance-from-center selection, faster point-cloud and zoomed-out selection paths, Vulkan measurement tool fixes, and drag-and-drop scene graph improvements. • UI/RmlUi platform work: Major RmlUi redesign efforts, hot reloading for RML/RCSS/Python UI files, reactive UI/store integration, viewport toolbar flyouts, improved histogram interactions, input settings enhancements, custom TRS gizmos, and numerous panel, tooltip, and localization fixes. • Windowing and UX: Added borderless window support, title bar drag/maximize/restore behavior, work-area-aware maximize functionality, resize responsiveness and performance improvements, and DPI/UI scaling fixes. • Training and data features: Added adaptive depth loss and depth gradients for the EWA rasterizer, mask loading/application fixes, a new combined Ignore+Segment mask mode, --add-splat, --freeze, improved checkpoint and training state handling, and training speed and VRAM optimizations. • COLMAP/equirectangular support: Added SPHERICAL/equirectangular camera model support and canonical EQUIRECTANGULAR handling, along with fixes for undistortion and camera export. This release will be available to all supporters as a Windows binary via approximately in about an hour. At the same time, LichtFeld Studio remains committed to being free and open source under GPLv3 and can also be built directly from source. Please consider supporting the ongoing development of LichtFeld Studio through a donation via the portal or the supporters page. Thank you to everyone who supports this project financially, contributes code, reports bugs, provides datasets, helps with the website, and contributes in countless other ways. A special thank you to our foundational sponsor Core11 and our Gold Sponsor Volinga, whose support has helped make the current state of the software possible. Thank you as well to every donor and to all of our new Bronze Sponsors. Looking ahead to v0.6 For the next major release, work will focus primarily on stability and user experience. This includes improved cleanup workflows and the ability to modify training parameters while training is in progress. I would also like to introduce a native .licht project format that allows users to save and restore their complete editor state. You can find links to our main sponsors below. Please also visit our website to discover all our Bronze Sponsors. Hint: We do not yet have a Silver Sponsor or Platinum 😉

MrNeRF

26,219 görüntüleme • 3 ay önce

llama.cpp isn't just for text LLMs anymore. Pure C++ zero shot voice cloning just officially landed in mainline. Text generation was only step one. If you’re building autonomous local AI agents, real time voice assistants, or edge workflows, instant low latency audio is the missing piece. Thanks to PR #26254, Alibaba’s state of the art Qwen3 TTS model family is now natively supported directly inside the llama.cpp repository under the multimodal (mtmd) framework. No Python bloat. No massive PyTorch CUDA overhead. Just raw, hyper optimized C++ running GGUF voice weights. Here is why this native update is a massive deal for the open source local AI stack: # Multimodal Architecture (.gguf + mmproj) Qwen3-TTS splits the workload between the base language model backbone and a multimodal projection adapter. llama.cpp handles this using the llama-tts binary, mapping the text model alongside its --mmproj projector to process audio tokens seamlessly. # Zero Shot Voice Cloning in Seconds You don't need fine tuning or massive dataset training. Feed the C++ engine a single 5 to 10 second .wav audio sample using the --tts-speaker-file flag, and it accurately clones the exact timbre, tone, and accent on the fly. # Real World T4 GPU Benchmark & Resource FootprintRunning the 1.7B Base model in 8-bit quantization (Q8_0): - VRAM Footprint: ~7 GB peak VRAM during active zero-shot cloning. - Audio Quality: Studio grade, natural-sounding voice output in seconds. • - Execution: Direct execution via native compiled binaries or sub process calls. # Coming Next to llama-server (PR #26603) Beyond CLI execution, a native POST /tts HTTP endpoint is currently being added to llama-server, which will soon allow you to trigger voice generation directly via standard REST API requests! # quick note on Colab compilation: Because this code was merged into mainline very recently, pre-built third-party binaries haven't fully caught up yet. Compiling llama-tts directly from source on Google Colab's free CPU instance can take about 1 hour (or ~1-2 minutes if targeting single GPU arch like -DCMAKE_CUDA_ARCHITECTURES=75). Be patient during the build step, or compile it locally on your own rig for instant execution! To test this out yourself, I built a zero config Google Colab notebook that compiles llama.cpp, downloads the Q8_0 GGUF files from HuggingFace, and spins up an interactive Gradio Studio UI so you can record/upload 3 second clips and clone voices in real time. Stop sleeping on native C++ audio. The era of bulky Python audio pipelines is officially over. Links to the free Google Colab notebook and the official ggml org GGUF HuggingFace model repository are in the replies below! available in q4 and q8 both variants, 1 GB and 1.85 GBs respectively (requires additional ~500MB mmproj gguf) Are you building local voice agents yet? What does your current audio stack look like? Drop your setups below!

Alok

47,881 görüntüleme • 1 ay önce

And here it is, 🚨🚨Documents obtained by investigative journalist Yehuda Miller through the Freedom of Information Act, CONFIRMS, BEYOND THE SHADOW OF A DOUBT, The Department of Justice and the FBI uncovered a massive 2020 ballot fraud operation based in Michigan and operated in multiple swing states including Washington DC, Chicago, IL, Georgia, Iowa, South Carolina, Pennsylvania, and Florida, funded by Joe Biden’s 2020 presidential campaign. GBI strategies BUSTED for submitting fraudulent voter registrations during 2020 election cycle. Following a raid, Michigan law enforcement discovered caches of pre-paid gift cards, semi-automatic rifles with silencers, four modified pistols with ammunition inside. and disposable burner phones. Throughout the 2020 election period, these Democratic cartel election committees provided more than $4,000,000 to this criminal organization: Biden for President: $450,000 Democratic Senatorial Campaign: $2,117,605 DNC Services Corp: $1,031,856 Democratic Party of Iowa: $493,100 The investigation was initiated following the observation of a Muskegon, Michigan, clerk who noticed an individual depositing 8,000 to 10,000 completed voter registration applications at the city office on October 8, 2020. This same individual returned multiple times, registering an additional 2,500 voters. Thousands of these registrations displayed identical handwriting with fraudulent addresses and phone numbers. Additionally signatures did not match those on file with the criminal Secretary of State Jocelyn Benson. 🚨🚨And the icing on the cake. + 100,000 ballots cast in the Nov 3, 2020 election were also falsified at Jocelyn Benson’s own election headquarters. FOIA: “I am following up with some information on one of the most egregious situations that I encountered. There are many examples of significant election integrity issues, but this example involves actions of GOVERNMENT OFFICIALS from both the local and state levels that are shocking.” Dawn N. Ison Department of Justice Assistant U.S. Attorney | public corruption unit, “We received this complaint. Guess what the criminals do, when the FBI and the DOJ got hold of this investigation? Passed it right back to this criminal who were running it, 📝who the fuck was she planning to shoot?

🇺🇸RealRobert🇺🇸

500,240 görüntüleme • 2 yıl önce

Top 25 AWS services explained EC2 – Your server. But in the cloud. You pay even when it’s doing nothing. (Sound familiar?) Lambda – EC2’s lazy cousin. Only wakes up when there’s work. No work, no bill. ECS – “I want containers but Kubernetes gives me anxiety.” EKS – Kubernetes. For people who enjoy suffering professionally. Auto Scaling – Your app gets famous overnight. This makes sure it doesn’t die from the attention. S3 – A bucket that never fills up. Jeff Bezos’s gift to humanity. EBS – A hard drive for your EC2. Loyal. But only to one instance at a time. EFS – EBS but for people who like sharing. Multiple instances, one file system. FSx – EFS but for enterprises who need Windows compatibility and a bigger invoice. Snowball – When your internet is too slow to upload data, AWS ships you a literal box. VPC – Your private neighborhood inside AWS. Strangers not allowed. Route 53 – The GPS of your app. Tells traffic where to go. ELB – The bouncer at the club. Splits traffic so no one server gets overwhelmed. CloudFront – Your content, cached globally. Because nobody likes a slow website. Direct Connect – A private highway between your office and AWS. No public internet drama. RDS – A managed database. AWS handles backups so you can sleep at night. DynamoDB – NoSQL at insane speed. Schema? We don’t do that here. Aurora – RDS on steroids. Faster, smarter, slightly more expensive. Redshift – A warehouse for your data. Not clothes. Petabytes of analytics data. ElastiCache – RAM for your app. Because hitting the database every time is embarrassing. IAM – The bouncer for your entire AWS account. Get this wrong and you’re headlines. KMS – Locks your secrets in a vault. AWS holds the key. You trust them. Mostly. Cognito – “Login with Google” but you built it on AWS. GuardDuty – The security camera that never blinks. Watches for sketchy behavior 24/7. WAF – Stops hackers at the door before they touch your app. Bookmark it.

Akhilesh Mishra

35,971 görüntüleme • 7 ay önce

my 8 GB VRAM gaming laptop is absolutely going to hate me for this. but I still did it. ran a 31b dense model (Gemma 4 31b Q4) with only 8 GB VRAM last week I ran Gemma 4 26B A4B a mixture of experts model on my RTX 4060 and hit 25–28 tokens/sec using llama.cpp's new MTP support. smooth. snappy. but MoE has a secret: it only activates 4B parameters per token despite having 26B total. that's why it flies. so the real question started haunting me. what if I throw a full, no tricks, every parameter fires on every token, 31B DENSE model at the same machine? # Hardware: GPU: NVIDIA RTX 4060, 8 GB VRAM RAM: 16 GB CPU: Intel Core i7 H Laptop. Gaming. Modest. The model: gemma-4-31B-it-qat-UD-Q4_K_XL.gguf (model's unsloth huggingface link in the comments) This is Google DeepMind's flagship dense model in the Gemma 4 family that can run on single consumer GPU. It packs a hybrid attention architecture, supports up to 256K context natively, and is QAT (Quantization Aware Training) optimized, meaning it retains far more quality than standard post training quants at the same bit depth. This is NOT the MoE. This is 31 BILLION dense parameters, every single one of them loaded. # the flags I used: -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -cnv --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 8 --spec-draft-p-min 0.6 -c 6000 -v Multi Token Prediction (MTP) is still active here. Separate draft GGUF required, same as the 26B setup. # Results: → Decode: ~3 tokens/sec → Prefill: ~2 tokens/sec → Context: 6000 tokens → Hardware crying quietly in the corner: yes so is 3 tps actually usable? For real time back and forth chat? Not ideal. You're not having a fluid conversation at 3 tps. but slow ≠ useless. And this is where it gets genuinely interesting. think about how senior devs actually work in a real team. But when something is architectural, deeply complex, or needs serious reasoning? they walk down the hall and escalate to the senior. That's exactly the local AI agent architecture this unlocks: → Fast orchestrator model (Gemma 4 26B MoE at 25+ tps) handles routing, simple queries, tool calls, memory. The junior dev. → Gemma 4 31B dense is the senior, called only when the fast model genuinely hits a wall. Hard multi step reasoning. Complex code generation. Deep architectural decisions. The agentic loop stays fast. Only the hard hops touch the 31B. That's a legitimate production grade local AI architecture on a budget hardware. (requires 2 8gb gpus) other workflows where 3 tps is completely fine: - overnight batch jobs. summarize documents, extract structured data, review code. Fire it off. Sleep. wake up to results. - One shot deep reasoning - Silent code audit loops, you write and test, the 31B reviews diffs and flags issues in the background between your sprints - Any workflow where output quality > output speed A few weeks ago, nobody was running a 30B+ dense model on a single consumer GPU with 8 GB VRAM. At all. Now we're doing it on an Intel i7-H gaming laptop with a NVIDIA RTX 4060, thanks to llama.cpp + QAT quants + MTP speculative drafting. Google DeepMind said the Gemma 4 31B targets "consumer GPUs and workstations." They were not exaggerating. The hardware bar to run serious frontier class models locally keeps dropping. the tools are here. the models are here. you just have to be willing to abuse your laptop a little. what workflows would you actually run on a local 3 tps 31B dense model? genuinely curious. drop it below.

Alok

63,689 görüntüleme • 3 ay önce

FACT CHECK: Here at the first trial, the Commonwealth’s own expert witness, Ian Whiffin, confirms the necessity & importance of hash values for the sake of “hash verification”, a necessary step in authenticating the data & being able to verify that it hasn’t been altered or manipulated. In fact, Whiffin actually gives this testimony in response to a question about when the data have been altered or tampered with, if there’s a way for the forensic examiner (him) to detect it, and/or verify its authenticity and integrity. Remarkably, despite the DFIR industry standard methodology of hash verifying a digital forensic extraction, like that of Jen McCabe’s iPhone, prior to conducting any analysis on it with any forensic tools, Ian Whiffin testified that notably, for his work on this case, not only did he abandon this standard methodology, but he also admitted that the forensic extraction of Jen McCabe’s iPhone, which he received from the Commonwealth, was stripped of its hash value. Perhaps more remarkably, this stunning fact apparently didn’t raise any red flags for Ian Whiffin when conducting his analysis in this case, where he’s providing testimony in a murder trial. One must ask themselves why that is? However, defense expert Richard Green, in his affidavit, states that: “Typically, forensic examiners are provided with the raw image file and the associated: hash value documentation together. After validating the hash value, I would then accept that the data has not been manipulated. Here, however, the hash documentation was not provided with the raw image of the cell phone. Instead, it was withheld from the defense. As a forensic examiner having received hundreds of imaged phones over the course of my decades-long career, this was unprecedented.” Contrary to Mr. Whiffin’s approach, upon initially receiving a purported extraction of Jen McCabe’s iPhone without a hash value to authenticate and verify the integrity of the data, Mr. Green promptly requested the hash value and corresponding GrayKey supplemental files from the Commonwealth in order to conduct his analysis. After making this demand, and when the Commonwealth had to produce the hash verification data for Jen McCabe’s iPhone, remarkably, the Commonwealth also produced—for the first time, and over a year later on February 8, 2023—the Full File System Extraction of Jen McCabe’s iPhone (see “Notice of Discovery VIII,” attached). Unlike the initial purported “extraction” produced by Trooper Nicholas Guarino, this one contained Jen McCabe’s incriminating 2:27am Google search and all of the manual deletions of her communications, among other incriminating evidence, surrounding the murder of Officer John O’Keefe (see defense’s Rule 17 motion from April 12, 2023, attached). So, this begs the question: If Ian Whiffin knows the importance of hash verification in validating the authenticity of the data he’s working with in the first place, then why didn’t he take the same actions as defense expert Richard Green did to responsibly and reliably provide analysis in this case? If Whiffin ought to be deemed an expert, qualified to provide analysis and testimony at trial, then why did he abandon his industry’s standard methodology of hash verification in this case? Even Cellebrite knows this is a no-no! What say you? #KarenReadTrial #Cellebrite #DFIR

Olivia

20,211 görüntüleme • 1 yıl önce

In July 2014, Lars Mittank, a 28-year-old from Germany went on a vacation with his five of friends to Varna, Bulgaria. The week went by really fast," said Paul Rohmann, one of Mittank’s friends. However, the day before the group was scheduled to fly back home, Lars got in a fight with other German nationals at a bar. The fight stemmed from a disagreement over football. Lars was absent for most of the evening but reappeared at his hotel the next morning, recounting to his friends that he had been assaulted by a group of four men, resulting in injuries to his jaw and a ruptured eardrum. Afterwards, Lars visited a doctor who recommended against flying due to his injury. Additionally, the doctor prescribed an antibiotic for him. Despite his friends' desire to stay with him, he insisted he could manage alone. He encouraged them to stick to the initial travel schedule and fly home on July 7, which they did. Lars checked out of the hotel with his friends and then checked into a new hotel. But this is where things turn really strange. The day following his friends' departure, Lars displayed odd and paranoid behavior. From his hotel, he contacted his mother, Sandra Mittank, speaking in hushed tones, expressing fears that individuals were attempting to harm or rob him. He also urged her to cancel his credit cards. The CCTV in the hotel captured him pacing up and down the halls, looking out windows, and hiding in an elevator. Around 1 am, he departed from the hotel and came back approximately an hour later. His whereabouts during this time remain unknown. In the morning, he once again called his mother, telling her that the people pursuing him were getting closer. Lars arrived at Varna Airport on July 8, 2014, the day he intended to board his flight back to Germany. He sought advice from the airport doctor, Dr. Kosta Kostov, who later characterized his demeanor as "nervous and erratic." According to Kostov, he told Mittank that he was fine and could return home. Lars also expressed concerns about the medication he was taking. Shortly after, a construction worker entered the office, causing Lars to exclaim, "I don't want to die here. I have to get out of here." He abruptly fled the office, abandoning his luggage containing his wallet, cell phone, and passport. These final moments of Lars Mittank can be observed in the CCTV footage provided below. Upon exiting the airport, he can be observed in the footage jogging away from the airport, scaling a fence, and bolting into a meadow. Following this, he completely disappeared.

Morbid Knowledge

34,821,015 görüntüleme • 2 yıl önce

It’s just a video player on localhost:8080. That’s what I keep telling myself. Minimax H3 on Fish Creative Prompt Fixed-camera desktop screen recording of a real Windows 11 Chrome window filling the entire frame. Do not restyle the browser as a mockup, poster, or floating card. Keep the exact chrome from the reference still for every frame. Browser frame (locked, zero motion): - Tab title: Local Video Player - Address bar URL: localhost:8080 - Bookmarks bar left to right: Apps, Work, Personal, Google, GitHub, Docs, YouTube, then Other bookmarks on the right - Chrome controls: back, forward, refresh, extensions, profile avatar “M”, puzzle piece, shield, three-dot menu - Windows 11 taskbar along the bottom: search “Type here to search”, File Explorer, Google Chrome, Visual Studio Code, system tray, time 10:30 AM, date 5/20/2024 Page layout (labels, colors, and positions never change): - Left: large HTML5 video canvas - Under the canvas: play/pause triangle, timecode starting 00:00 / 00:10, purple seek bar, speaker, purple volume slider, fullscreen icon - Right sidebar on near-black: - “Video Player” with purple clapperboard icon - Controls: Space Play / Pause; ← → Seek ±5 seconds; ↑ ↓ Volume ±10%; F Toggle Fullscreen; M Mute / Unmute - Playlist: purple play arrow, “Sample Video”, 00:10 - Playback Rate: 0.5x, 1x filled purple and selected, 1.5x, 2x Character identity (must stay the same person in every pose): Photoreal young woman, long wavy blonde hair, gray school-style blazer over a white collared shirt, gray pleated mini skirt, black knee-high socks. Same face, same proportions, same outfit. Location stays the green grassy hill under a pale overcast sky. No costume change, no second character, no extra props. Motion rules: - Camera never moves. Browser chrome, sidebar, bookmarks, taskbar, and all UI text stay perfectly still. - Only the video canvas animates. Treat the canvas as a short fashion / pose test clip of the same girl on the hill. - Playback runs continuously at 1x. Do not pause the player. Do not flash a big pause icon. Do not cut away from the hill. - Timecode and the purple progress thumb advance smoothly from 00:00 to 00:10. Pose sequence inside the canvas (hold each pose ~1.5–2 seconds, then transition with natural body motion, hair, and cloth, not a hard cut): 1. Start pose (match the reference still): leaning forward, both hands on her knees, looking into camera with a pout. 2. Rise and stand tall: she straightens up, hands leave her knees, stands facing camera with arms relaxed at her sides, chin slightly lifted, wind moving her hair. 3. Hands-on-hips: she plants both hands on her hips, weight on one leg, slight hip cock, blazer and skirt settling, confident look into camera. 4. Turn / three-quarter: she twists her torso to her right, looks back over her shoulder toward camera, hair swinging, one hand lightly holding the hem of the skirt so it does not lift unnaturally. 5. Crouch / ready: she drops into a low athletic crouch, fingertips brushing the grass, still looking up at camera, same pout softening into a small smirk. 6. End pose: she stands again, steps one foot forward, both hands clasp in front of her thighs, then eases back toward the original lean-with-hands-on-knees so the last frame rhymes with the first still. Transitions must look like the same live-action take: weight shifts, knee bend, hair lag, grass and sky stay consistent. No teleport, no outfit swap, no extra readable text burned into the video. Audio: quiet outdoor hill ambience only. No music, no voiceover, no UI click spam. Style: photoreal live-action screen capture of a local demo page running in Chrome. Not animation, not a trailer, not a poster. Duration: about 10 seconds. Aspect: 16:9 desktop.

Sharon Riley

30,886 görüntüleme • 20 gün önce