- Physics - Data Science - Postgraduate Adjunct Lecturer... at Pan-Atlantic University, on data science and NLP - 9 Published Papers on African NLP - 3 Papers under review - Research paper reviewer at top conferences - Founded Tonative, a community that curates African Language Datasets for AI models.show more

Dearly Beloved
26,752 просмотров • 3 месяцев назад
In partnership wtih 3 pan-African data science & ML... communities —SisonkeBiotik, Ro’ya, & DS-I Africa— we hosted an Africa-wide Data Science for Health Ideathon, focused on using Google’s open health AI models to tackle real-world health challenges. More →show more

Google Research
10,031 просмотров • 9 месяцев назад
OpenAI's Deep Research is getting a run for its... money. Deep Lake was just released, and it's a different take on an AI system that can do deep research on your own data. You can use Deep Lake to build AI search with reasoning on your private and public data. (Look at the attached videos to get an idea of how it works.) If you want to research proprietary and sensitive data, Deep Research won't help you because it's limited to public data. Deep Lake, however, will allow you to use your private data. On top of that, Deep Lake supports multi-modal retrieval from the ground up. It uses vision language models for data ingestion and retrieval so that you can connect any data (PDFs, images, videos, structured data, etc.) You can even use mixed-data queries! Deep Lake can search your data from S3, Dropbox, and GCP. It learns from your queries over time, making the results as relevant to your work as possible!show more

Santiago
171,340 просмотров • 1 год назад
I used Jev to classify 1,018 AI research papers.... The result: $0.08 total cost and 256ms median end-to-end latency per paper. The pipeline was: 1. Summarize each paper with DeepSeek V4 Flash 2. Send the title + summary + 24 possible topics to Jev 3. Use Jev to classify each paper 4. Visualize everything on The summaries cost $3.99 on Together AI. The classifications cost $0.08 on TypeSafe AI. So for just over $4 of inference, I ended up with a pretty useful way to explore the top AI research papers from the past year. I think this is where things are heading: different models for different parts of the workflow, instead of using one model for everything. I’m running evals on the Jev classifications before replacing the current ones, but the site is already live:show more

Hassan
337,656 просмотров • 13 дней назад
Making way for the future of health care! With... Texas Longhorns having a home at Moody Center, we’re saying goodbye to the Drum. The demolition sets up the expansion of Dell Medical School and the establishment of the University of Texas at Austin Medical Center with MD Anderson News. But we aren’t simply building a traditional academic medical center. We have an opportunity that is unique in Texas and only possible at a few places in the world to build an academic medical center that is linked to a top research university and that is driven by innovations in technology, digital health, data science, artificial intelligence, robotics, material science and moreshow more

Jay Hartzell
196,972 просмотров • 2 лет назад
a $400k research analyst at a hedge fund gets... through maybe 15 papers a month a Grok Bot read 3,900 of them last night and pulled out 6 with a tradable rule inside analyst's actual job was never insight, it was filtering - deciding which of ~200 new quant papers a week deserves a backtest that filter used to cost a salary. now it's an api key and one instruction: keep only papers with an executable rule guy running it is 26, no finance degree. built it because arxiv publishes an rss feed nobody in retail bothers to read overnight it pulls every new q-fin paper, strips anything without a testable signal, pushes survivors straight into a backtest 11 weeks. 4,100 papers scraped. 9 came out of sample alive 9 out of 4,100 isn't the bot failing, that ratio is what published quant research looks like once you demand it hold on data it has never seen what kills the other 4,091 is deflated sharpe. run 4,000 variants and something always looks brilliant purely by luck desks have known this since 2014. Bailey's paper is free, sitting on the same arxiv the bot scrapes every night one of the 9 is live on 3 exchanges right now. 0.9 sharpe, nothing cinematic, just real papers were public the whole time and so was the math that filters them. gap was never access, it's that nobody had 4,000 hours to read Grok Bot has 4,000 hours. it spent them last night while you sleptshow more

Livsun
13,674 просмотров • 1 месяц назад
Introducing ml-intern, the agent that just automated the post-training... team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.show more

Aksel
1,268,589 просмотров • 5 месяцев назад
BlackRock Sees Blockchains Becoming AI Payment Rails BlackRock published... a new whitepaper on September 22 examining AI and digital assets. The paper, titled “The Machine-Native Economy,” explores how AI agents could use blockchain infrastructure. BlackRock argues that autonomous systems will need financial rails for economic activity. Stablecoins could support frequent machine payments for data, software, and computing resources. The research also examines tokenized assets and digital markets for computing capacity.show more

BSCN
39,644 просмотров • 6 дней назад
Building a personal knowledge base for my agents is... increasingly where I spend my time these days. Like Andrej Karpathy, I also use Obsidian for my MD vaults. What's different in my approach is that I curate research papers on a daily basis and have actually tuned a Skill for months to find high-signal, relevant papers. I was reviewing and curating papers manually for some time, but now it's all automated as it has gotten so good at capturing what I consider the best of the best. There are so many papers these days, so this is a big deal. You all get to benefit from that with the papers I feature in my timeline and on DAIR.AI. The papers are indexed using tobi lutke qmd cli tool (all of it in markdown files along with useful metadata). So good for semantic search and surfacing insights, unlike anything out there. I am a visual person, so I then started to experiment with how to leverage this personal knowledge base of research papers inside my new interactive artifact generator (mcp tools inside my agent orchestrator system). The result is what you see in the clip. 100s of papers with all sorts of insights visualized. I keep track of research papers daily, so believe me when I tell you that this system is absolutely insane at surfacing insights. This is the result of months of tinkering on how to index research and leverage agent automations for wikification and robust documentation. But this is just the beginning. The visual artifact (which is interactive too) can be changed dynamically as I please. I can prompt my agent to throw any data at it. I can add different views to the data. Different interactions. I feel like this is the most personalized research system I have ever built and used, and it's not even close. The knowledge that the agents are able to surface from this basic setup is already extremely useful as I experiment with new agentic engineering concepts. I feel like this knowledge layer and the higher-level ones I am working on will allow me to maximize other automation tools like autoresearch. The research is only as good as the research questions. And the research questions are only as good as the insights the agents have access to. Where I am spending time now is on how to make this more actionable. I am obsessed about the search problem here. The automations, autoresearch, ralph research loop (I built one months ago) are easier to build but are only as good as what you feed them. Work in progress. More updates soon. Back to building.show more

elvis
467,434 просмотров • 6 месяцев назад
🚀 Introducing EgoExo Forge - built on top of... Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tunedshow more

Pablo Vela
36,542 просмотров • 1 год назад
Robots don’t just need better brains. They need WAY... more real-world data. 🤖 And collecting high-quality dexterous robot data at scale is one of the hardest problems in physical AI. A fascinating approach is emerging: Wearable human demonstrations + structurally matched dexterous robots. Instead of humans directly teleoperating a robot, Chinese embodied AI startup X Square Robot's TwinDEX system captures the motion, contact, and visual information needed to train the robot — while keeping the data closer to the hardware that will actually execute the task. Early results show promising performance on tool use, fine manipulation, and complex contact-rich tasks. If this approach scales, it could change how we build real-world robotics datasets. The next frontier of physical AI might not be bigger models. It might be better data.show more

The Daily Ai
33,341 просмотров • 27 дней назад
Trained on zero real-world data. Learned to walk, pick... up boxes, and follow multi-step instructions... in the REAL world. ( 📌 Paper below) Researchers from Amazon FAR, Berkeley, Stanford, and CMU scanned real rooms with an iPhone, rebuilt them as 3D Gaussian Splatting scenes, then generated 48,000 synthetic trajectories of a Unitree G1 walking, grasping, and placing objects inside those virtual replicas. They rendered the robot's first-person camera view from each run and paired it with the matching language instruction and motion data. That's the dataset every humanoid team needs and nobody has: synced egocentric video + language + kinematics, at scale. Instead of collecting it in the real world, they manufactured it. They trained a vision-language-kinematics policy on that synthetic data alone, then deployed it on the physical G1 across five task types: navigation to a named object, lifting boxes of three different sizes with no per-size tuning, chained multi-step tasks, robustness to mid-task layout changes and flickering lights, and multi-minute long-horizon runs. No real-world fine-tuning at any point. Real-world interaction data has been the hard limit on humanoid learning... slow, expensive, and small. If scanning a room once and synthesizing thousands of labeled interactions holds up as a general recipe, that limit moves. Data stops being the bottleneck robotics teams have to solve for. 📌 Paper: Project: ——- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
12,950 просмотров • 2 месяцев назад
🔥 Nebius AI R&D is hiring AI Research Interns... for short, high-impact RL projects. Exclusive to X right now — no LinkedIn mass postings yet. In 2019, I was a fresh dental grad with 3 months of runway left, begging for an AI shot. I know the grind. We’re looking for sharp early-career folks (students, grads, career-switchers) to join us and work on: > Agent trajectories analysis at scale > Long-horizon tasks for coding agents > Pushing open RL environments > Any other data / RL env / eval project that will benefit open-source community What you get: 💰 Fully paid internship (3-6 month) 📦 100% open-source shipping 📄 Co-author research papers ⚡️ Access to Nebius compute infra 🌍 Remote-friendly (EU/US) or Amsterdam/London/other office. If you’ve done any cool AI/ML/RL stuff, dm me with your most impressive project + 1-sentence summary + cv Sharing appreciated!🤝show more

Ibragim
33,571 просмотров • 5 месяцев назад
Google is doing some great AI work in India... that is actually helping farmers on the ground and honestly, almost nobody is talking about it. It’s surprising how much we focus on every new model drop, while some of the most meaningful AI work is happening quietly in the real world. Google DeepMind’s AnthroKrishi team built two AI models using satellite imagery: - Agricultural Landscape Understanding (ALU) maps farm boundaries, trees and water bodies, with historical data going back 15 years. - Agricultural Monitoring & Event Detection (AMED) uses multispectral data to monitor crops, sowing and harvesting, identifying 11 major crops. And this is already being used at scale: - 5M+ farmers in Telangana through its agricultural DPI - 140M+ hectares covered by Terrastack - 2.6M hectares of irrigated land in Karnataka - CarbonFarm is using it across 12 countries - FAO is integrating the models into its global agricultural data platform with $2.5M from Google These models are now providing agricultural insights across 6 African countries, and the Agricultural Landscape Understanding layer is one of the most popular layers on Google Earth globally.show more

AshutoshShrivastava
58,766 просмотров • 1 месяц назад
This #CVPR2026 paper from our research team is trending... #1 on Hugging Face 🤗 Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to act. Trained on 138M high-quality samples, LocateAnything decodes bounding boxes in parallel instead of one coordinate at a time, improving localization accuracy while dramatically increasing throughput for visual grounding and detection. Project page:show more

NVIDIA AI
341,460 просмотров • 4 месяцев назад
🌍 The brand-new CV VC African Blockchain Report is... now live! The report, co-published by Absa Corporate and Investment Banking, depicts a clear message: Africa’s blockchain future is already here. Download the full report now: Our annual African Blockchain Report offers a data-rich view into the continent’s rise as a global blockchain frontier. Globally, blockchain made up just 3.2% of VC funding in 2024. In Africa? 7.4%. More than double. That’s not a coincidence. Some key takeaways from the report include: 🔹 Blockchain accounted for 12.7% of all African VC deals and 7.4% of funding, a growing share despite tighter capital markets 🔹 Africa’s share of global blockchain deals rose to 2.3%, even as global venture funding became more selective 🔹 Seed rounds dominated, attracting 34% of blockchain-focused funding, a clear signal of investor faith in early-stage innovation 🔹 Centralized Blockchain Financial Services led by funding share (41%), followed by DeFi (30%), and Data Verification (20%) 🔹 Nigeria led by deal count, while Seychelles ventures secured the highest funding shares at 31.7% 🔹 Median deal size for blockchain ($2.8M) was nearly double the all-sector African median Despite tighter capital conditions driven by both global and local challenges, funding still moved. African blockchain startups focused on practical applications, sharpening their impact across finance, infrastructure, and regulatory-compliant data solutions. Discover more about the African funding landscape and Web3 ecosystem in our report:show more

CV Labs
38,239 просмотров • 1 год назад
🚨 BREAKING: HHS Sec. Robert F. Kennedy Jr. effective... immediately is SHIFTING AWAY FURTHER from cruel, inhumane and unnecessary animal testing — announcing a direct final rule from the FDA and actions from NIH GOOD! NO MORE Dr. Fauci-style beagle abuse 😡 KENNEDY: "We are moving HHS toward a new era of biomedical research that puts human biology at the center of science." "We are modernizing outdated regulations, investing in human-based technologies, and breaking down barriers that have kept researchers dependent on animal models when better tools are available." "Under President Trump, we will use the best science to protect patients, accelerate medical breakthroughs, and end unnecessary animal testing.” The FDA rule removes references to animal testing and the NIH is focusing on investments in proper research 👏🏻 Overall, 20 new initiatives have been launched to surge human-oriented to reduce reliance on animal models Thank you, Secretary Kennedy! Another common sense W.show more

Eric Daugherty
37,606 просмотров • 8 дней назад
We are entering an extremely exciting era for open-weight... models. Kimi K2.6 now feels like a top agentic model. I took it for a spin via Fireworks AI fast inference APIs. Kimi K2.6 has impressive agentic capabilities, design skills, and the ability to synthesize large amounts of information. I built a little Skill that produces survey papers on any AI research topic you want. (see example in the clip) You can use the skill to tell your agent to generate a survey on whatever topic and watch it go to work. The artifact was fully generated by Kimi.ai's Kimi K2.6. It's cheap and fast. Next step for me is to explore ways to continue integrating the capabilities of these models on use cases like automating my LLM knowledge bases and augmenting my agent memory capabilities. Stay tuned for more.show more

elvis
47,678 просмотров • 5 месяцев назад
Since Professor V Kamakoti has been nominated in the... committee to overhaul our education system, he is being mocked by opposition leaders for his beliefs. Today, I sat to check his academic credentials . Alongside serving as the Director of IIT Madras, his academic and professional accomplishments include: • 250+ Research Papers and Patents: He has published over 250 research papers in leading international journals and conferences and holds numerous national and international patents. • Development of the SHAKTI Microprocessor: He led the design and development of SHAKTI, India’s first indigenous, open-source microprocessor. Microprocessor Development Programme (MDP): He heads this flagship initiative of the Ministry of Electronics and Information Technology (MeitY), aimed at strengthening India’s semiconductor and chip-design ecosystem. • Record Patent Filings – “One Patent a Day”: As Director of IIT Madras, he introduced the vision of “One Patent a Day.” Under his leadership, the institute filed a record 417 patents during 2024–25. • Role in the National Security Advisory Board (NSAB): As an NSAB member, he provides strategic advice to the Government of India on cybersecurity, telecommunications security, and digital infrastructure. • Leadership of the AI Task Force: He chaired the Artificial Intelligence Task Force constituted by the Ministry of Commerce and Industry, helping shape India’s AI policy framework. • Information Security Education and Awareness (ISEA): He oversees initiatives under the ISEA programme to build India’s cybersecurity workforce and promote public awareness of information security. • Research in VLSI and Hardware Security: His primary academic contributions lie in VLSI design, algorithms, and strengthening cryptographic security at the silicon hardware level. • Institutional Reforms at IIT Madras: Since assuming office as Director in 2022, he has helped IIT Madras retain its top position in the NIRF rankings while launching new online degree programmes and expanding the institute’s startup ecosystem. In recognition of his outstanding contributions to science and technology, he has been awarded the Padma Shri (2026), the DRDO Academic Excellence Award, and the IBM Faculty Award. Here is in his native village, moving on roads in a Prabhat Feri invoking the Grama Devta.show more

Rahul Kaushik
186,175 просмотров • 2 месяцев назад
Every enterprise deploying AI is making a bet on... where their data infrastructure lives. Regulators, customers, and compliance teams now want clear answers about where data is stored and processed before they sign. Announced today at BoxWorks London: Box Zones expands to 10 regions worldwide: • 3 new Zones: Switzerland, Singapore, Israel • In-region compute for France & Canada • Roadmap: in-region AI processing later in 2026 According to Gartner, 35% of countries will be locked into region-specific AI platforms by 2027. Organizations that build their content infrastructure with residency at the center now, will be better positioned to win regulated customers and deploy AI responsibly. Read the full announcement:show more

Box
2,949,694 просмотров • 3 месяцев назад