GutenOCR: A Grounded Vision-Language Front-End for Documents Fine-tunes Qwen2.5-VL-3B... and 7B using a grounded OCR evaluation protocol, substantially improving region- and line-level OCR as well as text-detection recallshow more

Niels Rogge
45,811 Aufrufe • vor 7 Monaten
Announcing llama-ocr – a free + open source OCR... tool! It takes documents (images for now) & outputs markdown, and does really well for complex receipts, PDFs with tables/charts, ect... Powered by Llama 3.2 vision on Together AI & available on npm today!show more

Hassan
443,929 Aufrufe • vor 1 Jahr
Fine-tune DeepSeek-OCR on your own language! (100% local) DeepSeek-OCR... is a 3B-parameter vision model that achieves 97% precision while using 10× fewer vision tokens than text-based LLMs. It handles tables, papers, and handwriting without killing your GPU or budget. Why it matters: Most vision models treat documents as massive sequences of tokens, making long-context processing expensive and slow. DeepSeek-OCR uses context optical compression to convert 2D layouts into vision tokens, enabling efficient processing of complex documents. The best part? You can easily fine-tune it for your specific use case on a single GPU. I used Unsloth to run this experiment on Persian text and saw an 88.26% improvement in character error rate. ↳ Base model: 149% character error rate (CER) ↳ Fine-tuned model: 60% CER (57% more accurate) ↳ Training time: 60 steps on a single GPU Persian was just the test case. You can swap in your own dataset for any language, document type, or specific domain you're working with. I've shared the complete guide in the next tweet - all the code, notebooks, and environment setup ready to run with a single click. Everything is 100% open-source!show more

Akshay 🚀
126,213 Aufrufe • vor 10 Monaten
Mistral OCR 4 turned a handwritten calculus exam into... clean LaTeX! We gave it a photo of a hand-written exam page. The model read the handwriting and rebuilt every formula into structured digital text Output: Time: 5.1s · Cost: $0.09 Formulas came through exactly right - the hard part was nailed. The graph, unfortunately, it didn’t redraw. But that’s the telling part: most OCR tools just dump the text and quietly drop the figure. OCR 4 caught the plot, boxed it, and tagged it as a chart. It doesn’t get redrawn, but it gets read and accounted forshow more

atomic.chat
420,530 Aufrufe • vor 2 Monaten
Baidu just open-sourced an OCR model that reads entire... 40-page documents in one shot. It's called Unlimited-OCR. 3 billion parameters but only 500 million active during inference. Runs 100% locally on your machine. Why this matters: traditional OCR tools chop documents page by page. Tables that span two pages break. Reading order gets lost. Cross-page context disappears. Unlimited-OCR processes the whole document at once. 32K context window. Text, formulas, tables, reading order all preserved across pages. Output comes out as clean structured Markdown. → 93% accuracy on the standard benchmark. +6 points over the baseline. → Error rate stays below 0.11 even past 40 pages. → Multilingual out of the box. → 2.12 million downloads on Hugging Face last month. 14,600 GitHub stars. For context: Amazon Textract, Google Cloud Vision, and Azure Document Intelligence all charge per page. This runs locally for free.show more

Vaibhav Sisinty
415,690 Aufrufe • vor 1 Monat
NVIDIA just made AI detect objects 10x faster by... deleting one step. It's called LocateAnything, and it removes the biggest bottleneck no one else was fixing in vision-language models. Normally a model builds each bounding box one coordinate token at a time. 100 objects means thousands of tokens before an answer. NVIDIA scrapped that: their Parallel Box Decoding predicts the whole box in a single forward pass, as one atomic unit. → 12.7 boxes/sec on one H100 → 10x faster than Qwen3-VL → +3.8% F1 on LVIS, accuracy up, not down → 3B params, runs on one consumer GPU Treating the box as one unit keeps its coordinates tied together, which is why accuracy climbed instead of falling. One model handles detection, GUI grounding, OCR, and document understanding, ready for computer-use agents, robotics, and document pipelines. 100% open source, weights, code, demo, and paper all live.show more

Alvaro Cintas
202,075 Aufrufe • vor 2 Monaten
Ready to blast off to a whole new level?... 🚀 🖱️The #ROGSpeedNova wireless performance streamlines each step within the data transmission protocol, as well as refining data packet size and frequency for optimized process efficiency. 👉🏻Learn more:show more

ROG Global
46,716 Aufrufe • vor 2 Jahren
this is the BEST vision language model I have... ever tried! Aria is a new model by Rhymes.AI: a 25.3B multimodal model that can take image/video inputs 🤩 They release the model with Apache-2.0 license and fine-tuning scripts as well 👏 I tested it extensively, keep reading to learn more 🧶show more

merve
176,125 Aufrufe • vor 1 Jahr
Turkish Foreign Minister Hakan Fidan on Iran: The war... in our region must end as soon as possible. For months, we have made great efforts to establish a negotiating table. Today as well, we say that diplomacy is the solution to problems, and we continue to work in that direction.show more

Clash Report
28,130 Aufrufe • vor 6 Monaten
From our humble beginnings in 2011 to today's reimagined... experience - watch as Goodnotes evolves through the years. What started as a simple digital paper app has transformed into a powerful platform for capturing ideas in any form. The new Goodnotes is here - with Whiteboards, Text Documents, and AI superpowers all in one unified place. No longer just Goodnotes 6, but simply Goodnotes - where your ideas come to life. Our journey continues with a fresh look and a bold new vision. Download now and experience the next chapter of note-taking.show more

Goodnotes
33,219 Aufrufe • vor 11 Monaten
Have a little writing app I'm cooking that's designed... just for my workflow. As you zoom out, every card flips to a one line summary so you can quickly scan and eval the argument one level up. Using Ollama local models to make these so they're quick and free. So simple, so useful.show more

Maggie Appleton
38,311 Aufrufe • vor 2 Monaten
How can we address the scarcity of data required... for specialized AI? Learn about Simula, a framework that reframes synthetic data generation as dataset-level mechanism design. By using reasoning to architect datasets from first principles, Simula enables fine-grained control over coverage, complexity, and quality. More →show more

Google Research
142,710 Aufrufe • vor 5 Monaten
A new chapter starts today. As MegaETH mainnet approaches,... Avon’s identity needs to reflect the system we’ve been building: a calm, transparent, and predictable place to lend or borrow. Our first brand was put together quickly during the early build phase. It let us move fast, but it never fully captured the level of clarity and focus we were aiming for. The new brand is built around those principles. Simple. Transparent. Grounded. A visual identity that matches Avon’s role as the credit and liquidity layer of MegaETH. It sets the tone for how lending should feel in a high-throughput, real-time environment: clear, controlled, and reliable. This is the foundation for what Avon is becoming as we head into mainnet.show more

Avon
56,614 Aufrufe • vor 9 Monaten
In the final chapters of the Book of Revelation,... the end of the age is described as a massive confrontation between the forces of good and evil. The text says the kings and armies of the world gather for a final battle in a place called Armageddon. Yes, a place. According to the vision recorded by John of Patmos, strange signs appear in the sky, angels descend, and powerful figures emerge as the crazy battle between "godly" and worldly powers. The scene is filled with dramatic weird imagery including thunder, fire, and cosmic chaos as gods and earth seem to collide. The story goes on with Christ returning with heavenly armies to smash the forces against him, including the figures known as the Antichrist and the false prophet. After the battle the enemies of God are defeated and a period of judgment begins. The vision then moves toward a renewal of creation, where the old world passes away and a new reality emerges described as a new heaven and a new earth. The final chapters kinda show a symbolic picture of restoring justice after a time of horsesh*t and conflict. Are we on the cusp?show more

Jason Wilde
14,651 Aufrufe • vor 6 Monaten
We’re excited to introduce Text-to-LoRA: a Hypernetwork that generates... task-specific LLM adapters (LoRAs) based on a text description of the task. Catch our presentation at #ICML2025! Paper: Code: Biological systems are capable of rapid adaptation, given limited sensory cues. For example, our human visual system can quickly adapt and tune its light sensitivity to our surroundings. While modern LLMs exhibit a wide variety of capabilities and knowledge, they remain rigid when adding task-specific capabilities. Traditionally, customizing these models requires gathering large datasets and performing often expensive, time-consuming fine-tuning for specific applications. To bypass these limitations, Text-to-LoRA (T2L) meta-learns a “hypernetwork” that takes in a text description of a desired task, as a prompt, and generates a task-specific LoRA that performs well on the task. In our experiments, we show that T2L can encode hundreds of existing LoRA adapters. While the compression is lossy, T2L maintains the performance of task-specifically tuned LoRA adapters. We also show that T2L can even generalize to unseen tasks given a natural language description of the tasks. Importantly, Text-to-LoRA is parameter-efficient. It generates LoRAs in a single, inexpensive step, based solely on a simple text description of the task. This approach is a step towards dramatically lowering the technical and computational barriers, allowing non-technical users to specialize foundation models using plain language, rather than needing deep technical expertise or large compute resources.show more

Sakana AI
403,354 Aufrufe • vor 1 Jahr
The Hidden Language of Diffusion Models paper page: tackle... the challenge of understanding concept representations in text-to-image models by decomposing an input text prompt into a small set of interpretable elements. This is achieved by learning a pseudo-token that is a sparse weighted combination of tokens from the model's vocabulary, with the objective of reconstructing the images generated for the given concept. Applied over the state-of-the-art Stable Diffusion model, this decomposition reveals non-trivial and surprising structures in the representations of concepts. For example, we find that some concepts such as "a president" or "a composer" are dominated by specific instances (e.g., "Obama", "Biden") and their interpolations. Other concepts, such as "happiness" combine associated terms that can be concrete ("family", "laughter") or abstract ("friendship", "emotion"). In addition to peering into the inner workings of Stable Diffusion, our method also enables applications such as single-image decomposition to tokens, bias detection and mitigation, and semantic image manipulationshow more

AK
41,830 Aufrufe • vor 3 Jahren
THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T... RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.show more

RetroChainer
11,100 Aufrufe • vor 1 Monat
NEW: Photon Matrix built an Iron Dome for mosquitoes.... The Blue Laser Mosquito Air Defense is built for patios, campsites, and outdoor dinners, using sensors to find mosquitoes mid-air and zap them with a visible blue laser. -Scans a 20-foot, 90-degree field in front of the unit -Targets flying insects as small as 2 mm -Uses LiDAR, radar, and AI vision to confirm targets -Internal aiming system points the laser at the mosquito -Outdoor kit adds a rotating base for 360-degree coverage -Runs about 4–5 hours from a power bank -Blue flash shows when it fires Pricing starts at $1,088.show more

Ritwik Pavan
3,921,248 Aufrufe • vor 18 Tagen
We’re thrilled to announce our partnership with @Reployai ,... a cutting-edge AI ecosystem that bridges developers and projects to the blockchain using natural language processing (NLP). This collaboration aligns with Aurk AI’s mission to revolutionize accessibility and functionality in AI and blockchain. Through our partnership, Aurk AI and Rai will deliver next-level solutions for AI and blockchain enthusiasts, ensuring scalability, usability, and innovation at every step. Stay tuned as we unlock new possibilities with this powerful collaboration!show more

AURK AI
44,263 Aufrufe • vor 1 Jahr
Show-o One Single Transformer to Unify Multimodal Understanding and... Generation discuss: We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model.show more

AK
124,085 Aufrufe • vor 2 Jahren