Загрузка видео...

Не удалось загрузить видео

На главную

Kled Version 3 is coming. Over $20M+ in rewards will be paid directly to users from leading AI labs across robotics, legal services, image and video generation, world modeling, and more. In the last seven days, we’ve received inbound data requests from several decacorn AI labs and enterprises for...

124,728 просмотров • 8 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

We’re excited to finally introduce Kled Special Tasks, the final major feature included in the V2 app update. Users will now have access to a fully interactive terminal where they can view and complete domain specific upload tasks directly from enterprise buyers. These tasks can be region locked and person specific. For example, PhD students at Stanford might be prompted to upload their coursework or research materials and get paid for it. Our first domain specific task will focus on homework collection from high school and college students across Europe and the United States. Students will verify their emails and academic credentials directly within the app. We’ve built labeling workflows to ensure all uploaded content meets our criteria, and participants will receive weighted payouts based on the value of their submissions. We’ve already built a network of over 3,800 students from Stanford, MIT, UIUC, Rutgers, and Duke who will be actively onboarded to contribute content. Kled will work hand in hand with our research division, HADES, to justify the large scale purchase of this homework content. Several enterprise buyers have already expressed interest, each confirming that academic data from students represents a growing multi year industry requiring a continuous flow of fresh material. Kled Special Tasks also gives us the ability to internally identify valuable content types, issue calls for specific datasets, and collect 1,000-2,000 unique samples per task. We can then package these datasets into specialized data packs that our sales team will use to pitch directly to AI labs and enterprise clients with matching data needs. This will be one of our most powerful tools for expanding Kled’s buyer network. All of this will be fully available in the V2 update. We’re excited to show just how advanced our segmentation and data validation software has become as we bring this release to market.

Kled AI

74,618 просмотров • 10 месяцев назад

Dupe has just made a $550,000 investment into $KLED, and are partnering up to enrich hundreds of millions of commerce data points. Dupe is one of the fastest-growing shopping networks on the internet. Approaching $100 million in GMV this year and on track to 5× that within the next 12 months. Over the last four months, Kled has expanded beyond data collection to full scale data enrichment/labeling, building enterprise infrastructure that turns raw data into structured, insight-rich training sets for next-generation AI. Unlike traditional shopping platforms, Dupe is platform-agnostic, it sees shopping behavior across the entire internet. This gives rise to a massive opportunity: to understand not just what users buy, but why they buy. Through Kled’s enrichment layer, we’re mapping shopping journeys in granular detail: – Identifying product discovery patterns – Understanding brand affinities – Measuring historical price sensitivity and intent – Predicting cross-category purchase paths Dupe will use this enriched dataset to power shopping LLMs capable of anticipating needs, personalizing recommendations, and reducing friction from discovery to checkout. In the coming months, Kled and Dupe will continue deepening this collaboration as we use this data to enhance their user experience. We’re excited to push past our limits and create the perfect labeling infrastructure for this data. Kled will continue to compete with not only the data collection giants but also the data enrichment unicorns that are worth billions of dollars.

Kled AI

80,171 просмотров • 11 месяцев назад

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 просмотров • 1 год назад

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 просмотров • 1 год назад

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,015 просмотров • 10 месяцев назад

After 8 months of building in stealth and testing our infrastructure on 10000+ hours of real-world data and hundreds of unique environments, we're bringing FPV Labs into the open today. FPV Labs started with the following bet - if human data proves to be the underlying factor that determines scaling laws in general-purpose robotics, it will trigger the largest economic transformation in human history, and the underlying infrastructure that captures that data will determine how fast we get there. We will achieve this by building the full-stack infrastructure for capturing, processing, transferring, and evaluating human experience into spatial, temporal, and semantic knowledge for machines. Despite all the research novelty behind ChatGPT, its success can be attributed to one foundational fact - the scaling law of transformers. We believe the same dynamics have made their way into robotics. Recent studies showed task completion rates jumping from 30% to 70% when human demonstration data scaled from 1,000 to 20,000 hours, a log-linear trend that mirrors exactly what we saw in language and vision. Seeing these emergent signs of scaling law curves in robotics, we believe we are entering the era of general-purpose robotics policies, which makes the next few years the most exciting time in the history of this field. But the library of physical interactions required to train general-purpose robot policies does not exist yet. Over the last 8 months, we've seen dozens of companies emerge in this space. We were really happy to see new companies pushing this space forward, but we also saw the same pattern repeat: every egocentric data company was making some tradeoffs between quality, scale, and diversity. We have built FPV labs on the core principle that high-quality data is orders of magnitude more valuable than sheer volume. Case in point, self-driving cars collect thousands of hours of data per day, but only a small fraction of that data is actually useful for training better models. Several studies, like RT-2, have shown that as little as 1% of data improves as much as 25% on task success. The quality and diversity of data matter a lot more than scale, so there is clearly a power law curve in the downstream impact of data. We've spent months obsessing over data quality by building our stack, discarding it, rebuilding it, and iterating until we found a formula that doesn't compromise downstream quality at scale. We believe the downstream impact here is far more profound than most people realize. Workers globally are paid around $60 trillion per year in aggregate, and a lion's share of that compensation goes to physical labor - tasks that require navigating real spaces, manipulating real objects, and negotiating the infinite variability of the physical world. Human-to-robot transfer will be one of the most important infrastructures that will shape our society in the near future, and if it works, the economic impact will dwarf every technology transition that came before it in an exponential manner and lead to the creation of goods and services we can’t imagine today. Our mission is to lay the groundwork for us to transition into this future - the future of abundance. We are deeply grateful to our earliest believers, Paras Chopra and Lossfunk, who played a critical role in shaping our thinking.

Abhishek Anand

82,785 просмотров • 5 месяцев назад

Mike is back and better than ever. Mike ( has undergone a major refresh packed with new features, design improvements, and quality-of-life enhancements. And as always, everything is open source. Here is a summary of the major updates: 1. Legal research Case law hallucinations are a massive problem. To address this, Mike's assistant can now produce grounded citations backed by case law from CourtListener. The integration is tight and highly intuitive. Every answer includes links to cited cases, along with supporting quotations from judicial opinions to substantiate each proposition made by the AI. The assistant also researches cases in much the same way a lawyer would. It retrieves potentially relevant authorities, verifies the accuracy of citations, and searches for keywords that align with the research topic. Support for additional jurisdictions will be added soon. 2. Improved user interface (UI) The UI has been redesigned around Apple's Liquid Glass design principles, giving the application a cleaner and more minimalist feel down to every modal and button. Mike is designed to be a place of calm where lawyers can focus amidst the chaos of legal practice. 3. Version control Lawyers are familiar with saving documents as new versions in document management systems such as iManage and NetDocuments. Mike now brings this workflow into projects by allowing users to save new versions on top of existing documents with a simple drag-and-drop action. This introduces order and clarity to matters where redlines and document versions can quickly become confusing, especially when there are five different files named "Final Execution Version." 4. Data control Users can now export their data and permanently delete their data directly from the Settings page. 5. Security Multi-factor authentication (MFA) is now available. Users can enable MFA to protect sensitive actions within Mike, such as data deletion, and can also optionally require MFA during login. There have also been numerous backend security improvements since the previous release. 6. MCP connectors This was one of the most requested features. Users can now add MCP connectors directly from Settings. All tokens and secrets are encrypted at rest. In the demo video, I showcase the Neimo MCP from k-ID, a global age verification and compliance platform. Mike can also be connected to other MCP servers, including email systems. The setup process is currently somewhat manual, but one-click authentication and onboarding will be added in a future release. Mike now has a strong open-source community behind it and is improving every single day. The summary above highlights just some of the work that has been shipped over the past month. I'll also be doing more demos to showcase Mike's capabilities. See you soon.

Will Chen

10,435 просмотров • 3 месяцев назад

Experiments in progress. The one on the right has been learning for ~3 hours, the one in the middle for ~1 hour, and the one on the left just started a few minutes ago. The initial motivation for making the physical Atari was just to commit ourselves to a subset of algorithms that can make progress in this setup. This commitment rules out algorithms that require billions of samples to learn (or worse, require multiple environments running in parallel). Atari games are simple enough that we should be able to show learning on them in a short amount of time with no prior knowledge. Since then, I've realized that this setup is also a good way to compare different paradigms in robotics in a principled way. These paradigms are sim2real, learning from tele-operated data, and learning directly on the robots. So far, I have observed that getting sim2real to work reliably is hard. It requires tweaks that don't scale. Policies that can play perfectly in simulation fall apart because of latencies and the messiness of the real world. These aspects could be modeled to improve the simulation, but not without sinking significant human engineering hours. I have higher hopes for learning from tele-operated data, but that requires a human to learn the task first. These experiments are on my to-do list. I have to learn to play some of the games well through the robot. I’m half-decent at playing Pong and Ms Pacman now. Learning directly on robots is looking like the most promising approach. This approach takes away pesky distribution shifts and makes it possible to have algorithms that continually improve with more data and time without any human intervention. It feels great to let experiments run overnight and wake up to find improved policies. With learning on robots, I should, in principle, be able to go on a long vacation and come back to find better policies for complex tasks beyond Atari games. Whether that is possible with current learning algorithms is a different question.

Khurram Javed

52,110 просмотров • 9 месяцев назад