Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing FlyOCR 🪰 - I trained a fly brain to read a PDF It uses the full MaleCNS v1.0 fruit fly connectcome. The architecture is inspired by doomfly by The fly splits a pdf image into individual glyphs, maps pixels into receptor activations, runs simplified current-based dynamics across the...

36,648 Aufrufe • vor 2 Tagen •via X (Twitter)

11 Kommentare

Profilbild von Jerry Liu
Jerry Liuvor 2 Tagen

if u don't want to use a fly brain to read a PDF, you could always check out an actual OCR solution like llamaparse

Profilbild von Utkarsh Agrawal
Utkarsh Agrawalvor 2 Tagen

@nftechie_ I trained it to be a MMA fighter

Profilbild von Aethmere-OS
Aethmere-OSvor 2 Tagen

@nftechie_ A useful control: shuffle the connections while preserving neuron degrees, then train the same readout. That could help show how much the fly wiring contributes to OCR.

Profilbild von pwguler
pwgulervor 2 Tagen

@nftechie_ a fruit fly reading a pdf is the only benchmark that cannot be gamed: the fly has no incentive to please the evaluator

Profilbild von Imaan Sultan 🇵🇰🇸🇦
Imaan Sultan 🇵🇰🇸🇦vor 2 Tagen

@nftechie_ that’s so wild holy

Profilbild von Mesut De
Mesut Devor 2 Tagen

@nftechie_ congrat

Profilbild von Zaid Rais | Design Engineer
Zaid Rais | Design Engineervor 2 Tagen

@nftechie_ The fly connectome mapping is the fun part but the real test is whether receptor activations generalize past clean glyphs to a scanned page with noise and skew.

Profilbild von First Sauce Labs
First Sauce Labsvor 2 Tagen

@nftechie_ This feels like a good example of when the "AI" part just means calling a well-known component a few new things.

Profilbild von Oussama
Oussamavor 2 Tagen

@nftechie_ 87% on real financial glyphs before even touching MNIST territory suggests the connectome is doing more useful work than the benchmark framing implies.

Profilbild von Kevihaiceth 💹🧲
Kevihaiceth 💹🧲vor 2 Tagen

@nftechie_ Training a fly brain to read PDFs is wild

Profilbild von klownShowz
klownShowzvor 2 Tagen

@nftechie_ Lord of the PDF flies.

Ähnliche Videos

Web scraping will never be the same. (100% open-source visual search at scale) PixelRAG is a retrieval system that skips HTML parsing completely. Instead of scraping a page into text and embedding chunks, it screenshots the page and retrieves the image. A vision-language model reads the answer straight off the pixels. Why that matters: parsing is where web RAG quietly loses information. - A single HTML-to-text parser can drop 40%+ of a page. - Tables, charts, and layout get flattened or thrown out. - Swapping parsers alone can move accuracy ~10 points on the same docs. PixelRAG indexes the page a person actually sees. The team built a visual index of all of Wikipedia, 30M+ screenshots, and it still beats the strongest text RAG baseline by 18.1% on text-only QA. The repo also ships a Claude Code plugin that gives Claude eyes. It lets Claude screenshot any URL and read the rendered page instead of scraping the DOM. So you can hand it a live page, an arXiv paper, or your local site and ask what it actually looks like. One setup script. No MCP server, no backend. How the pipeline works: - Renders each document (web, PDF, image) to image tiles. - Embeds them with Qwen3-VL-Embedding, LoRA fine-tuned on screenshots. - Builds a FAISS index and serves a search API. A stronger reader model lifts accuracy with no re-indexing, since the index is just pixels. Everything is open-source under Apache-2.0. GitHub repo: Talking about RAG, I recently wrote an article on a new approach that makes retrieval much more efficient by cutting corpus size by 40x, reducing tokens per query by 3x, and improving vector search relevance by 2.3x. The article is quoted below.

Akshay 🚀

947,346 Aufrufe • vor 2 Monaten

$BRAIN is live. the open-source fly-brain bridge has a ticker now, and the ticker is becoming the input for a physical robot. CA: 9LuwgFQemAoV9rgVBBtwbRSxBmamcRRbysEK8yL2pump what’s already done: the neural bridge is working. camera input becomes activity across eight virtual neural populations, and that activity becomes left and right motor commands IMU data feeds movement back into the model. smoothing, speed limits and a 500 ms watchdog are already built in the synthetic demo runs locally today. the ESP32 scaffold is ready for hardware integration the entire project is open source and MIT licensed github: what’s next, in order: the on-chain listener goes live. every $BRAIN transaction becomes a stimulus sent directly into the neural model then the physical robot connects. the wallet address determines which neurons activate, the amount determines the strength of the impulse, and the brain converts that reaction into movement then the livestream. every transaction, neural impulse and physical response visible in real time then user commands. spend $BRAIN to trigger a stronger reflex, request movement, temporarily control the robot, name a neuron or place your name on the stream the loop, plain: a transaction enters the neurons react the robot moves a Proof of Reflex is created with the transaction hash, neural activity map and video of the physical response the current brain is a small fly-inspired simulation. the bridge and local demo work today. the blockchain connection, physical robot and livestream are being built now the internet becomes its sensory organ. the blockchain becomes its nervous system. $BRAIN makes the body move.

Frank

203,581 Aufrufe • vor 1 Tag

There's been an unfortunate incident in LA with a Uhaul plowing into a crowd of anti-Khameini protestors This man should never have been able to get near the crowd with a uhaul but some info about the situation seem to be - signage on the truck is anti both the shah and the current ayatollah. is this his actual position or is this camouflage to have gotten into the protest to perpetrate an attack? "no Shah, No regime, No Mullah" Mullah is a religious leader so possibly referring to the current leader and not a king like Pahlavi Timeline appears to be - anti-Khameini protestors try to rip signs off his vehicle and are bashing on the windows and eventually his passenger side window is broken - guy in uhaul then stutterstops forward into the crowd, eventually accelerating further, then stuttering again, then full stopping down the road - there is a man surfing on top of the uhaul in the 3rd video, below I have posted another video showing the man on top of the uhaul trying to take the posters off the side, so he is likely part of the anti-Khameini protestors - uhaul driver is taken into custody by police is this a case of police not having the street sufficiently blocked off and so a guy was able to get a uhaul in here? He should not have been able to drive a uhaul this close to a massive protest crowd There are a lot of people saying this is a terrorist attack, it is possible it could be one but I don't think there's enough information to accurately assert that at this time The chronology of events also shows it is possible that the driver was in fear of his life since protestors were banging on the uhaul, windows, and removing signs+ eventually breaking his window Whatever turns out to be the actual case, it is an unfortunate event and as of right now a seeming silver lining is that no deaths have been reported

Kirsche 🥥 🧁

41,247 Aufrufe • vor 8 Monaten