Загрузка видео...

Не удалось загрузить видео

На главную

Submodular optimization for token/sentence selection from long contexts. Here's an interesting exp: first used jina-embeddings-v4's multi-vector feature to extract token-level embeddings from a passage, then applied submodular optimization to cherry-pick the tokens that provide the best coverage, finally call tokenizer and convert selections back to the strings at their...

13,192 просмотров • 1 год назад •via X (Twitter)

Комментарии: 3

Фото профиля Jina AI
Jina AI1 год назад

Try it on Google Colab's L4 GPU for free: This could be an interesting approach for extracting information from long documents, saving tokens for LLMs, etc. Check out our recent blog posts and learn more about submodular optimization.

Фото профиля Richard Collins, The Internet Foundation
Richard Collins, The Internet Foundation1 год назад

Can you scale to replace Google? Put your AI on it and put in the numbers. Back of the envelope or "in a spreadsheet" is better than "in your head somewhere as an idea only". It might be easier than you think now. If the whole Internet is coded as it goes in, not scraped and indexed and tokenized later - completely separated from the authors, without their permission or help. Check my writing on "global open tokens" where all tokens are linked to the real things in the world - not arbitrary strings of characters in one language. Using universal (global) tokens means "the sun", "the earth", "water" and those are independent of human language so ties things together. Yes, choose the things that matter, keep it lean and sufficient and sustainable, not shotgun or brute force, only for people with big computers. For all humans, not just a few. Richard Collins, The Internet Foundation

Фото профиля Franck Lebeau
Franck Lebeau1 год назад

interesting how "Late chucking" is condensed into "lateing" (tokens "late" + "##ing"). As I understand it, it means that the semantic of "chuncking" (tokens "chunck"+"#ing") is mainly supported by the contextualized embedding of the "#ing".

Похожие видео

So a core issues I see with a lot of crypto games is: No in game peer to peer economy. (ie. The game Devs make all the $$ off the items and you buy from them not other players) Couple of things I wanted to experiment with, in regards to on chain gaming is the peer to peer economy. So the way my team set it up, is you mine the token in game, but you sell the items peer to peer in SOL. (see video for a demo) This gives you two different ways to play to earn, you can farm items, from bosses and loot boxes, and sell them to players for solana:So11111111111111111111111111111111111111112 , then other players can buy them and use them to mine the token, and sell the token to traders who will speculate on the price. Or you can just buy the loot boxes, and use the items to farm the token. This should help with the sell pressure, because before the game flow was: Buy item (in token, From devs) --->farm a ton of token---> Sell token to recoup cost of item. New flow is : Player opens loot boxes (paid in token)/ Player Plays game (in game token sinks) ---> gets item ---Lists item--> other player buy Item (From players, In SOL) ----Farms token---> sells token. Now the equation is if the loot box-player spent more game tokens to get the item compared to what he gets for in sol, and it keeps more money in the in game economy, and the market will fluctuate more based on the price of the in game token. So players who play and are "locked in" can make money off arbing and this services the "trader" archetype of player, which I feel is the major player type in web3-games. (also the token name is just a test token so dont buy random shit on pump fun, thats just using the name)

Skely

15,064 просмотров • 2 месяцев назад