Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing FlyOCR 🪰 - I trained a fly brain to read a PDF It uses the full MaleCNS v1.0 fruit fly connectcome. The architecture is inspired by doomfly by The fly splits a pdf image into individual glyphs, maps pixels into receptor activations, runs simplified current-based dynamics across the...

36,648 görüntüleme • 2 gün önce •via X (Twitter)

11 Yorum

Jerry Liu profil fotoğrafı
Jerry Liu2 gün önce

if u don't want to use a fly brain to read a PDF, you could always check out an actual OCR solution like llamaparse

Utkarsh Agrawal profil fotoğrafı
Utkarsh Agrawal2 gün önce

@nftechie_ I trained it to be a MMA fighter

Aethmere-OS profil fotoğrafı
Aethmere-OS2 gün önce

@nftechie_ A useful control: shuffle the connections while preserving neuron degrees, then train the same readout. That could help show how much the fly wiring contributes to OCR.

pwguler profil fotoğrafı
pwguler2 gün önce

@nftechie_ a fruit fly reading a pdf is the only benchmark that cannot be gamed: the fly has no incentive to please the evaluator

Imaan Sultan 🇵🇰🇸🇦 profil fotoğrafı
Imaan Sultan 🇵🇰🇸🇦2 gün önce

@nftechie_ that’s so wild holy

Mesut De profil fotoğrafı
Mesut De2 gün önce

@nftechie_ congrat

Zaid Rais | Design Engineer profil fotoğrafı
Zaid Rais | Design Engineer2 gün önce

@nftechie_ The fly connectome mapping is the fun part but the real test is whether receptor activations generalize past clean glyphs to a scanned page with noise and skew.

First Sauce Labs profil fotoğrafı
First Sauce Labs2 gün önce

@nftechie_ This feels like a good example of when the "AI" part just means calling a well-known component a few new things.

Oussama profil fotoğrafı
Oussama2 gün önce

@nftechie_ 87% on real financial glyphs before even touching MNIST territory suggests the connectome is doing more useful work than the benchmark framing implies.

Kevihaiceth 💹🧲 profil fotoğrafı
Kevihaiceth 💹🧲2 gün önce

@nftechie_ Training a fly brain to read PDFs is wild

klownShowz profil fotoğrafı
klownShowz2 gün önce

@nftechie_ Lord of the PDF flies.

Benzer Videolar

Web scraping will never be the same. (100% open-source visual search at scale) PixelRAG is a retrieval system that skips HTML parsing completely. Instead of scraping a page into text and embedding chunks, it screenshots the page and retrieves the image. A vision-language model reads the answer straight off the pixels. Why that matters: parsing is where web RAG quietly loses information. - A single HTML-to-text parser can drop 40%+ of a page. - Tables, charts, and layout get flattened or thrown out. - Swapping parsers alone can move accuracy ~10 points on the same docs. PixelRAG indexes the page a person actually sees. The team built a visual index of all of Wikipedia, 30M+ screenshots, and it still beats the strongest text RAG baseline by 18.1% on text-only QA. The repo also ships a Claude Code plugin that gives Claude eyes. It lets Claude screenshot any URL and read the rendered page instead of scraping the DOM. So you can hand it a live page, an arXiv paper, or your local site and ask what it actually looks like. One setup script. No MCP server, no backend. How the pipeline works: - Renders each document (web, PDF, image) to image tiles. - Embeds them with Qwen3-VL-Embedding, LoRA fine-tuned on screenshots. - Builds a FAISS index and serves a search API. A stronger reader model lifts accuracy with no re-indexing, since the index is just pixels. Everything is open-source under Apache-2.0. GitHub repo: Talking about RAG, I recently wrote an article on a new approach that makes retrieval much more efficient by cutting corpus size by 40x, reducing tokens per query by 3x, and improving vector search relevance by 2.3x. The article is quoted below.

Akshay 🚀

947,346 görüntüleme • 2 ay önce

$BRAIN is live. the open-source fly-brain bridge has a ticker now, and the ticker is becoming the input for a physical robot. CA: 9LuwgFQemAoV9rgVBBtwbRSxBmamcRRbysEK8yL2pump what’s already done: the neural bridge is working. camera input becomes activity across eight virtual neural populations, and that activity becomes left and right motor commands IMU data feeds movement back into the model. smoothing, speed limits and a 500 ms watchdog are already built in the synthetic demo runs locally today. the ESP32 scaffold is ready for hardware integration the entire project is open source and MIT licensed github: what’s next, in order: the on-chain listener goes live. every $BRAIN transaction becomes a stimulus sent directly into the neural model then the physical robot connects. the wallet address determines which neurons activate, the amount determines the strength of the impulse, and the brain converts that reaction into movement then the livestream. every transaction, neural impulse and physical response visible in real time then user commands. spend $BRAIN to trigger a stronger reflex, request movement, temporarily control the robot, name a neuron or place your name on the stream the loop, plain: a transaction enters the neurons react the robot moves a Proof of Reflex is created with the transaction hash, neural activity map and video of the physical response the current brain is a small fly-inspired simulation. the bridge and local demo work today. the blockchain connection, physical robot and livestream are being built now the internet becomes its sensory organ. the blockchain becomes its nervous system. $BRAIN makes the body move.

Frank

203,581 görüntüleme • 1 gün önce

There's been an unfortunate incident in LA with a Uhaul plowing into a crowd of anti-Khameini protestors This man should never have been able to get near the crowd with a uhaul but some info about the situation seem to be - signage on the truck is anti both the shah and the current ayatollah. is this his actual position or is this camouflage to have gotten into the protest to perpetrate an attack? "no Shah, No regime, No Mullah" Mullah is a religious leader so possibly referring to the current leader and not a king like Pahlavi Timeline appears to be - anti-Khameini protestors try to rip signs off his vehicle and are bashing on the windows and eventually his passenger side window is broken - guy in uhaul then stutterstops forward into the crowd, eventually accelerating further, then stuttering again, then full stopping down the road - there is a man surfing on top of the uhaul in the 3rd video, below I have posted another video showing the man on top of the uhaul trying to take the posters off the side, so he is likely part of the anti-Khameini protestors - uhaul driver is taken into custody by police is this a case of police not having the street sufficiently blocked off and so a guy was able to get a uhaul in here? He should not have been able to drive a uhaul this close to a massive protest crowd There are a lot of people saying this is a terrorist attack, it is possible it could be one but I don't think there's enough information to accurately assert that at this time The chronology of events also shows it is possible that the driver was in fear of his life since protestors were banging on the uhaul, windows, and removing signs+ eventually breaking his window Whatever turns out to be the actual case, it is an unfortunate event and as of right now a seeming silver lining is that no deaths have been reported

Kirsche 🥥 🧁

41,247 görüntüleme • 8 ay önce