last week we open sourced: anydoc: parse 14 document... formats, ~5ms per page pdf-inspector: parses + classifies pdfs, no waiting on OCR 14k stars each, built in rust. now powering Firecrawl /parse, which adds OCR for the hard ones what should we open source next?show more

Nicolas Camara
100,899 просмотров • 1 месяц назад
introducing anydoc now your agents get 100x faster local... parsing for pdf, docx, pptx & 10 more formats - sub-5ms md conversion - 500 docx files in 1.7s - top quality across all 13 formats - rust-based - open source already powering Firecrawl /parseshow more

Nicolas Camara
2,537,857 просмотров • 1 месяц назад
China open-sourced a peanut-sized OCR that parses entire 100-page... PDFs in one shot.. It's called Unlimited-OCR. Only 3B params. Runs locally. Every other OCR tool chops your doc into pages and loses the thread. this one reads the whole thing in a single pass. → One-shot "long-horizon" parsing (32K context window) → Multilingual, out of the box → 93% on the standard parsing benchmark (+6 over baseline) → <0.11 error rate past 40 pages → Runs 100% locally on your own hardware → Works with Transformers, vLLM, SGLang, Docker, Ollama, llama.cpp Traditional cloud OCR (Textract, Google Vision, Azure Doc Intelligence) costs $1.50–$15 per 1,000 pages. This runs on your machine. For free. Forever. Baidu built it explicitly to push DeepSeek-OCR one step further. Already at 1.9M downloads on Hugging Face and most people have no idea it exists yet. 100% open source.show more

Superman
1,105,000 просмотров • 1 месяц назад
Baidu just open-sourced an OCR model that reads entire... 40-page documents in one shot. It's called Unlimited-OCR. 3 billion parameters but only 500 million active during inference. Runs 100% locally on your machine. Why this matters: traditional OCR tools chop documents page by page. Tables that span two pages break. Reading order gets lost. Cross-page context disappears. Unlimited-OCR processes the whole document at once. 32K context window. Text, formulas, tables, reading order all preserved across pages. Output comes out as clean structured Markdown. → 93% accuracy on the standard benchmark. +6 points over the baseline. → Error rate stays below 0.11 even past 40 pages. → Multilingual out of the box. → 2.12 million downloads on Hugging Face last month. 14,600 GitHub stars. For context: Amazon Textract, Google Cloud Vision, and Azure Document Intelligence all charge per page. This runs locally for free.show more

Vaibhav Sisinty
415,690 просмотров • 1 месяц назад
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

AlphaSignal
92,137 просмотров • 5 месяцев назад
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

Jafar Najafov
13,630 просмотров • 5 месяцев назад
🚨 Alibaba just open sourced a GUI agent that... lives inside your webpage and controls it with natural language. It's called Page Agent and it's not a browser extension. It's pure JavaScript no Python, no Puppeteer, no headless browser, no screenshots. Just one script tag and your web app understands natural language. Here's what it actually does: → Embed it with a single tag or npm install → Control any web interface with plain English commands → Text-based DOM manipulation no OCR, no vision models needed → Bring your own LLM (GPT, Claude, Qwen, anything) → Ships a built-in UI with human-in-the-loop support → Turn 20-click ERP/CRM workflows into one sentence → Optional Chrome extension for multi-tab agent tasks → Works on any web app SaaS, admin panels, internal tools Companies are charging $30/month for AI copilots built on this exact idea. This is 3 lines of code. Your users. Your interface. The AI copilot layer for every web app just got open sourced. 1.6K stars. 100% Open Source. (Link in the comments)show more

Ihtesham Ali
135,634 просмотров • 6 месяцев назад
NVIDIA just made AI detect objects 10x faster by... deleting one step. It's called LocateAnything, and it removes the biggest bottleneck no one else was fixing in vision-language models. Normally a model builds each bounding box one coordinate token at a time. 100 objects means thousands of tokens before an answer. NVIDIA scrapped that: their Parallel Box Decoding predicts the whole box in a single forward pass, as one atomic unit. → 12.7 boxes/sec on one H100 → 10x faster than Qwen3-VL → +3.8% F1 on LVIS, accuracy up, not down → 3B params, runs on one consumer GPU Treating the box as one unit keeps its coordinates tied together, which is why accuracy climbed instead of falling. One model handles detection, GUI grounding, OCR, and document understanding, ready for computer-use agents, robotics, and document pipelines. 100% open source, weights, code, demo, and paper all live.show more

Alvaro Cintas
202,075 просмотров • 2 месяцев назад
We built an interactive 3D module to teach kids... about what temperature does to water. - You adjust the temperature in real time - You see the impact on the state of water and what's changing at a molecular level - You can unlock two secret states in there if you pass a quick quizz This is the first of a giant series. We're turning the open source Marble App curriculum into a full bank of these interactive lessons. One for each topic kids learn in primary school. If you are a parent or primary school teacher, let us know in comments which module you want to see next. Link to play with the water lab below 👇show more

Lionel Mora
118,441 просмотров • 2 месяцев назад
"Last week my husband and I went to the... shelter to adopt a baby. We had a plan. One puppy. That was it. But life had something else waiting for us. At the shelter, we met two tiny Presa Canario mix puppies who had just lost their mama in a tragic accident. They were so small… but somehow, they already understood what it meant to lose everything. We picked one up. And that’s when everything changed. She didn’t relax in our arms. She reached. Right for her brother. Her tiny paws wrapped around him like she was holding onto the only thing she had left in the world. And he… he didn’t hesitate. He buried his face into her neck, holding on just as tightly. No barking. No crying. Just quiet fear… and a bond that was impossible to ignore. Presa mixes are known for loyalty—but this? This was something deeper. This was survival. This was love. We looked at each other… and we knew. There was no way we were separating them. So we didn’t. We adopted both. This photo was taken on the ride home. Two little souls, finally safe, curled up together in our arms… still holding on to each other like they promised they always would. And in that moment, watching them breathe in sync… We realized something powerful— Sometimes the best decisions aren’t the ones you plan. They’re the ones your heart makes for you. Credit & Love ~ Sarah Williams, Coloradoshow more

ماعز-
99,315 просмотров • 1 месяц назад
People often ask what a normal week at a... dog rescue looks like, and the truth is, it’s rarely glamorous. One day you’re climbing headfirst into a drainage tunnel at 9 AM, and the next you’re out feeding street packs in a rainstorm. Sometimes it means showing up to work feeling completely unwell but refusing to quit, or relying on 10-minute car naps between rescues just to keep your eyes open. The exhaustion is real, and some weeks completely drain you but then you look at that last photo. Seeing Tina’s Hospital finally coming together makes every single rainstorm and brutal day completely worth it. This is exactly what we are working so hard for, and we wouldn't trade it for the world. ❤️show more

Happy Doggo
24,241 просмотров • 2 месяцев назад
claude fable 5 is live. spawn 5.0 was built... with it: 1,687 prompts, 102 sessions, my job shifted from architecture to judging taste. what we built, each of which would've been at least a month with a whole team on opus: — a from-scratch physics engine (mantle) that rivals rapier, testable in a day as a side thread — clustered froxel lighting: 8 lights to 1,000+ fully dynamic ("it's stupid fast" — creator of threejs) — realtime diffuse GI on webgpu, on your phone (landing in the next update) — million-particle gpu vfx with a shader architecture beyond what unreal and unity ship — mmo-scale netcode plus, for good measure: an (alpha) native ios app. all of it in about a week. all of it running on mobile. and the list doesn't include half of what we shipped — full changelog in the thread below. and the part the benchmarks won't show you: this model has a wonderful personality. it's genuinely funny. i laughed so hard i cried multiple times, mid-physics-rewrite. the genius and the character aren't separate features.show more

jacob
55,515 просмотров • 3 месяцев назад
🚨 One photo of your face. That's all someone... needs to become you on a live video call. In real time. Right now. The tool is free and open source. It's called Deep-Live-Cam. One image. One click. You become anyone on a live webcam feed. No training. No datasets. No waiting. Instant. Your face. Your expressions. Your mouth movements. All stolen from a single photo. Here's what this thing does: → Upload one photo of any face → Turn on your webcam → You are now that person. Live. In real time. → It matches your pose, your expressions, even your lighting → Mouth masking so the swapped face moves its lips when you talk → Multi-face mapping. Swap different faces on different people in the same call. → Virtual camera output. Plug it into Zoom, Google Meet, Teams. Nobody knows. → Works on NVIDIA, AMD, Intel, and Apple Silicon Here's the part that should terrify you: Your boss could be on a Zoom call with someone wearing your face right now. A scammer could call your parents looking exactly like you. A stranger could take your LinkedIn photo and become you in a video meeting. IShowSpeed's reaction when he saw it: "What the F**! This shit is crazy!" SomeOrdinaryGamers: "That's fucking freaky dude... that's so wild." This was the #1 trending repo on GitHub the day it launched. 1,600 stars in 24 hours. 80K+ stars today. No one is ready for what this means. And it's already out there. 100% Open Source.show more

Nav Toor
306,490 просмотров • 5 месяцев назад
we spent $2,352 promoting a health app on tiktok.... here's what happened. most brands are paying $8 CPM for ads nobody watches. this app converted $0.09 per install instead. we scrapped the whole thing and ran it through Content Rewards with an Adaptive AI slideshow agent instead. 25,638 downloads. $6,635 MRR. 30 days without paid ads. what changed: before: -> paying creators flat fees with no guaranteed results -> creators posting inconsistent formats with no proven structure -> briefs going back and forth for days before a single post went live -> 15-20 slideshows a week if they were lucky after: -> brands drop the agent link in their Content Rewards campaign brief -> creators open it, type a topic, slideshow is ready in minutes -> 200+ creators posting the same proven format simultaneously -> 25,638 downloads in 30 days -> $6,635 MRR. $0.09 per install. 2.82x ROI on spend same $2,352 and a 16x-38x better revenue outcome than Meta. the agent writes the story. Void AI generates the slides. exports a PNG pack and MP4 ready to upload. comment "SLIDES" + RT and i'll send you the agent so you can run this for your next Content Rewards campaign (must be following so i can DM)show more

adam
61,840 просмотров • 4 месяцев назад
For the past three to four years now, I... usually go on holiday twice a year — one international solo trip and one local family vacation. I do go for my solo vacation in spring and our family holiday is during the summer when the kids are on school break. Schools have finally shutdown for summer break last week and the kids are already all pumped up for our family vacation which will be coming up in few weeks from now. We haven't yet gone for our family summer vacation for this year but babe has already started making plans for our 2027 summer holiday. Last weekend, she brought up the idea of just the two of us going on a boat cruise next summer. She wants it to be just me and her without the kids. She suggested that we bring in a family member on invitation from Naija, such that there will be a family at home with the kids while we both go on a lovey-dovey boat cruise to western Europe. She advised that we should start saving up for the adventure by September. This woman and long term planning are five and six, and that virtue has been very helpful in helping us avoid financial pressure. We do everything with ease, thanks to her wisdom. For instance, two weeks before the end of the session that ended last week, we have already bought the kids school uniforms for 2026/27 academic session that will be starting in September; thanks to my wife's wisdom. If she hasn't suggested it, it wouldn't have even crossed my mind at all. I am awake now, staring at this beautiful woman sleeping peacefully next to me and I feel nothing but joy in my heart and total love for her. I am really blessed to have someone like her as my better half. I am the often chaotic and the more expressive person while she is the calm, quiet and better organised one. We perfectly compliment each other. It'll be our 12th anniversary in exactly a month's time from now, and I really look forward to our anniversary getaway weekend. Cheers to forever, baby.show more

Olamide Obe
13,140 просмотров • 1 месяц назад
🚀 Update Next Scene V2 only 10 days after... last version, now live on Hugging Face 👉 🎬 A LoRA made for Qwen Image Edit 2509 that lets you create seamless cinematic “next shots” — keeping the same characters, lighting, and mood. I trained this new version on thousands of paired cinematic shots to make scene transitions smoother, more emotional, and real. 🧠 What’s new: • Much stronger consistency across shots • Better lighting and character preservation • Smoother transitions and framing logic • No more black bar artifacts Built for storytellers using ComfyUI or any diffusers pipeline. Just use “Next Scene:” and describe what happens next , the model keeps everything coherent. 🧩 Try it directly in ComfyUI, or check the thread to launch it on fal . Open-source, no restrictions, made for filmmakers, animators, and dreamers. ComfyUI #AIcinema #LoRA #Flux #Qwen #ComfyUI #AIart #GenerativeVideo you can test on comfyui or to try on you can go here : and use my lora link : start your prompt with "Next Scene:" and lets go !!show more

Lovis Odin
43,366 просмотров • 10 месяцев назад
Following the amazing reaction to the Marble Curriculum yesterday,... we've decided to make it open source 🛰️👇 Everything a child learns in primary school. 1,590 concepts. 3,221 connections across 8 subjects, from Math and Science to Computing and Life Skills. Anchored in the US and UK curriculums, standard by standard (NGSS, Common Core, DfE). What you will find in the repo: every concept as structured JSON with its age band and the evidence a child must show to master it. Every prerequisite link marked hard or soft, with a written rationale. It's a true DAG you can compute learning paths on. Open license, you can build whatever you want with it. Now is a unique time in history to be building in education. Getting AI and kids education right is likely one of the hardest and most important problems to crack over the next decade and we need as many smart and creative minds behind it. We think a common solid basis, accessible to all and that can be built upon, is critical to move fast. That's why we're making this curriculum open source. It's not perfect but we know it's a robust basis, and we believe that sharing it openly is the fastest way to progress in this field. If you're building in education, share this around you and tell us in comments if you find this useful and if you want to contribute. We'll keep working and investing on it Marble App. Credit goes to Guillaume Boniface-Chang for building this. I just made it look pretty. Links below 👇show more

Lionel Mora
1,326,396 просмотров • 2 месяцев назад
Announcing Centaur 2.0! Centaur is frontier, agentic infrastructure that... you own. Centaur is like Claude Tag, but open source and on steroids. Centaur 2.0 can connect to everything you have access to, and can reason over it next to where you work, either in Slack, Discord, Teams or in your local Codex or Claude Code via an MCP. But context isn't useful if it can be accessed by anyone, so we rebuilt Centaur from the ground up for security. Centaur obviously doesn't have access to secrets because we're leveraging egress proxies. Centaur 2.0 takes that a step further enabling administrators to configure who can access what from where. This means that Centaur can have access to private information that you normally wouldn't feel comfortable giving it access to (e.g. DMs or sensitive channels and docs) but only expose it to authorized principals. It also means you can add Centaur to external channels and leverage it as a virtual colleague that doesn't live only in your Slack, for example, but also your Slack Connect channels! This is extremely powerful as we start moving to a world where agents cross organizational boundaries. Also in case anyone's wondering, yes we rewrote it in Rust! Centaur is now way way more stable at durably executing threads, and its workflow engine is now based on Absurd. We have been operating Centaur since January, officially launched and Open Sourced it in May, and now we're full speed towards making it the #1 open source agentic infrastructure. To succeed at that, I'm thrilled to welcome Matthew Slipper to the Paradigm team who will be leading all our Applied AI work, while continuing to maintain and extend Iron Proxy as the leading secure secret access for agents. Welcome Matt, it's an honor after all these years of knowing you! Read the full blogpost below, and apply to join our team!show more

Georgios Konstantopoulos
144,133 просмотров • 1 месяц назад
🚨BREAKING: Google just merged Gemini and NotebookLM into one... unified workspace and it changes everything about how you use AI for deep work. It's called Notebooks in Gemini and it's the personal knowledge base that power users have been begging for. You create a notebook for a project, drop in your files, PDFs, and documents, give Gemini custom instructions, and every chat you have stays organized in one place. No more hunting through old conversations. No more re-uploading the same files every session. The wildest part is the sync. Anything you add in Gemini automatically appears in NotebookLM. Anything you add in NotebookLM automatically appears in Gemini. One source of truth. Two powerful apps. Zero friction switching between them. So you can start a research notebook in Gemini, ask it questions all week, then flip to NotebookLM to generate a Cinematic Video Overview from the same material. Next morning, open Gemini and ask it to write a full report on exactly what you just watched. That workflow used to take three apps and a lot of copy-pasting. Now it's one notebook. Rolling out this week to Google AI Ultra, Pro, and Plus subscribers on web. Mobile and free users coming soon. What do you think?show more

Mayank Vora
136,834 просмотров • 5 месяцев назад
The final Epoch is now complete. DNA becomes fully... deflationary. 3 years ago, we bet on a revenue-share model before it was the meta. While the market was chasing high inflation tokens, we were building a sustainable system for the first truly on-chain DAO on Multiversᕽ. It was an experiment, and it is now battle-tested. The proof is on-chain. Over the last 4 epochs, the DAO generated enough revenue to buy back over 6m $DNA from the total supply. Even with EGLD dropping ~90% since our token fair-launch (with no raise, pre-mine, allocations, or hidden unlocks) $DNA sits at roughly the same USD price today as it did on Day 1. The DNA/EGLD pair has outperformed the market and remained healthy, proving the model works even in the toughest conditions. Now, the dynamic flips. Emissions have officially ended. We are entering a full deflationary cycle. No more DNA will be printed, but the DAO will keep buying. We expect the value of DNA to reflect this scarcity, while the power of your Subject X NFTs remains the key to accessing that value. What’s next? We switch our focus fully to strengthening the MultiversX ecosystem. We are here to build tools, products, and open-source tech that grow the chain. When EGLD grows, our products grow. When our products grow, the DAO earns more. When the DAO earns more, $DNA gets stronger. Thank you to everyone who made this journey possible. History is written. Now we continue building.show more

Project X 🧬
11,333 просмотров • 8 месяцев назад