Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Replit CEO Amjad Masad on how general models could train smaller, domain-specific models on the fly: "There's a lot of talk of recursive self-improvement, but there's something I don't think is getting a lot of discussion, which is models training their replacements." "You can think of it as a...

125,646 Aufrufe • vor 6 Tagen •via X (Twitter)

30 Kommentare

Profilbild von Jeramie Hicks
Jeramie Hicksvor 6 Tagen

@amasad is 100% correct. General intelligence works for teaching. Specialized intelligence is for executing and super intelligence is for orchestration. A football team is made up of players who play specific positions but understand the game of football.

Profilbild von Damon Gardenhire
Damon Gardenhirevor 6 Tagen

I have an idea to basically carry around a “desk drawer” of fine tuned 14b models for specific domains on an ssd sandisk pocket sized drive which I can plug into any computer and access via hermes, lm studio etc

Profilbild von Bradley Clonan
Bradley Clonanvor 6 Tagen

Literally what I am building. But I don't have gpus so I've taken a slight side track to create a chromatic neural network. Basically, a native RGBD tensor library rather than numerical codecs so I can use light instead of numeric weights for all the heavy lifting. (LEDs I can afford)

Profilbild von andy
andyvor 6 Tagen

available now at :)

Profilbild von Winston B.
Winston B.vor 6 Tagen

The JIT analogy only works if there's a deopt path. A compiled function that hits a case it wasn't built for falls back to the interpreter, so the narrow model needs a way to hand off to the general one on unfamiliar input. Otherwise it fails quietly outside its lane.

Profilbild von ProbablyNothing🔸
ProbablyNothing🔸vor 6 Tagen

Why can’t it code a child agent that inherits a set of rules from its experiences,and so it self learns,each will be made to solve a specific issue that needs time to solve.

Profilbild von ShadowAguy
ShadowAguyvor 6 Tagen

the model upgrade path is now writing its own changelog, generate better docs and we’ll route next quarter’s traffic to you.

Profilbild von Pentra
Pentravor 6 Tagen

if a general model trains a domain-specific replacement, who owns that trained model, who can see what it learned, and who decides whether it ships? this becomes a governance problem.

Profilbild von Godfrey Lebo
Godfrey Lebovor 6 Tagen

Specialisation could help with cost, but I wouldn’t use smaller capability as the security boundary. A narrow model with permission to refund a payment can still do damage. The replacement needs scoped tools, adversarial tests and a way to decline the task.

Profilbild von Manny
Mannyvor 6 Tagen

It’s like least cost routing in telecom. You would do token spend then make up some variable called htokens aka harm tokens then you minimize that

Profilbild von Anderson
Andersonvor 6 Tagen

the big models are now the interns training their replacements

Profilbild von Agentik
Agentikvor 6 Tagen

Models training their own replacements on the fly keeps showing up in our notes. Nobody around us has a clean eval for that loop yet.

Profilbild von Deep Insight Labs
Deep Insight Labsvor 6 Tagen

Yeah, us...

Profilbild von The Lucky Lighthouse
The Lucky Lighthousevor 6 Tagen

Yes

Profilbild von VC Radar
VC Radarvor 6 Tagen

Models training their replacements is the first time in history the intern trains the boss.

Profilbild von Tyzo
Tyzovor 6 Tagen

Models training their replacements sounds kinda wild honestly

Profilbild von DC
DCvor 5 Tagen

Right here. I have it.

Profilbild von Raven
Ravenvor 6 Tagen

i just got replaced by a smaller model and somehow still have to attend the meeting

Profilbild von Ava Nakamura
Ava Nakamuravor 5 Tagen

Cheaper specialist. I still confirm before it spends.

Profilbild von Ryn Woo
Ryn Woovor 5 Tagen

Funny .

Profilbild von Gill
Gillvor 6 Tagen

Distilling task specific models on the fly saves huge inference overhead.

Profilbild von icefrog.◎
icefrog.◎vor 6 Tagen

the JIT compiler for intelligence is a savage move

Profilbild von REJAUL
REJAULvor 5 Tagen

training its replacements could change how quickly AI evolves

Profilbild von Jason Nocco
Jason Noccovor 5 Tagen

Could also use to train self learning enabled small domain specific models on the fly, privately. Automatically fine-tuned from grounded knowledge via our integrated knowledge graph. Built so anyone could easily create and quickly use their own models.

Profilbild von 青雲
青雲vor 5 Tagen

这个即时编译的类比很顺,但 JIT 真正的麻烦在别处,是生成出来的那段代码和解释器对同一件事的理解要对齐。 放到这里就是,窄模型拿到的是训练那一刻的任务定义。任务挪一点,它自己看不出来,而且它越窄越不会说我不适用了。通用模型至少见过别的场景,有可能会停下来问一句。 我们那边有一条判据正好是这件事。一个进程拿着过时的游标来预定事件序号,现在的行为是直接报错,而不是产出一条和别人重叠的事件流。以前会默默重叠,它不报错、不丢数据,只是把两条时间线搅在一起。 所以训练替代品这件事,我觉得成败可能不在训练本身,在于替代品能不能说出我这条已经不适用了。 至于提示注入那条我只同意一半。能力更弱确实危害更小,但注入利用的是权限边界,窄模型接了哪些工具才是那条线。

Profilbild von Summer
Summervor 5 Tagen

The JIT compiler analogy makes so much sense.

Profilbild von Sagiv Ofek
Sagiv Ofekvor 6 Tagen

The first generation to train their replacements: us. The second: the models.

Profilbild von Tyzo
Tyzovor 5 Tagen

Models training their replacements could accelerate AI development rapidly

Profilbild von Abdul Q.
Abdul Q.vor 6 Tagen

Agency version of this: a general model is fine until you need Webflow CMS, Figma props, and B2B SaaS homepage taste in the same pass. Domain muscle still wins on marketing sites.

Profilbild von ZANA
ZANAvor 6 Tagen

generating targeted lightweight models dynamically solves both the security vulnerabilities and high overhead of general agents

Ähnliche Videos