Loading video...

Video Failed to Load

Go Home

Replit CEO Amjad Masad on how general models could train smaller, domain-specific models on the fly: "There's a lot of talk of recursive self-improvement, but there's something I don't think is getting a lot of discussion, which is models training their replacements." "You can think of it as a...

123,960 views • 6 days ago •via X (Twitter)

30 Comments

Jeramie Hicks's profile picture
Jeramie Hicks6 days ago

@amasad is 100% correct. General intelligence works for teaching. Specialized intelligence is for executing and super intelligence is for orchestration. A football team is made up of players who play specific positions but understand the game of football.

Damon Gardenhire's profile picture
Damon Gardenhire6 days ago

I have an idea to basically carry around a “desk drawer” of fine tuned 14b models for specific domains on an ssd sandisk pocket sized drive which I can plug into any computer and access via hermes, lm studio etc

Bradley Clonan's profile picture
Bradley Clonan6 days ago

Literally what I am building. But I don't have gpus so I've taken a slight side track to create a chromatic neural network. Basically, a native RGBD tensor library rather than numerical codecs so I can use light instead of numeric weights for all the heavy lifting. (LEDs I can afford)

andy's profile picture
andy6 days ago

available now at :)

Winston B.'s profile picture
Winston B.6 days ago

The JIT analogy only works if there's a deopt path. A compiled function that hits a case it wasn't built for falls back to the interpreter, so the narrow model needs a way to hand off to the general one on unfamiliar input. Otherwise it fails quietly outside its lane.

ProbablyNothing🔸's profile picture
ProbablyNothing🔸5 days ago

Why can’t it code a child agent that inherits a set of rules from its experiences,and so it self learns,each will be made to solve a specific issue that needs time to solve.

ShadowAguy's profile picture
ShadowAguy6 days ago

the model upgrade path is now writing its own changelog, generate better docs and we’ll route next quarter’s traffic to you.

Pentra's profile picture
Pentra6 days ago

if a general model trains a domain-specific replacement, who owns that trained model, who can see what it learned, and who decides whether it ships? this becomes a governance problem.

Godfrey Lebo's profile picture
Godfrey Lebo5 days ago

Specialisation could help with cost, but I wouldn’t use smaller capability as the security boundary. A narrow model with permission to refund a payment can still do damage. The replacement needs scoped tools, adversarial tests and a way to decline the task.

Manny's profile picture
Manny6 days ago

It’s like least cost routing in telecom. You would do token spend then make up some variable called htokens aka harm tokens then you minimize that

Anderson's profile picture
Anderson6 days ago

the big models are now the interns training their replacements

Agentik's profile picture
Agentik6 days ago

Models training their own replacements on the fly keeps showing up in our notes. Nobody around us has a clean eval for that loop yet.

Deep Insight Labs's profile picture
Deep Insight Labs6 days ago

Yeah, us...

The Lucky Lighthouse's profile picture
The Lucky Lighthouse6 days ago

Yes

VC Radar's profile picture
VC Radar6 days ago

Models training their replacements is the first time in history the intern trains the boss.

Tyzo's profile picture
Tyzo5 days ago

Models training their replacements sounds kinda wild honestly

DC's profile picture
DC5 days ago

Right here. I have it.

Raven's profile picture
Raven6 days ago

i just got replaced by a smaller model and somehow still have to attend the meeting

Ava Nakamura's profile picture
Ava Nakamura5 days ago

Cheaper specialist. I still confirm before it spends.

Ryn Woo's profile picture
Ryn Woo5 days ago

Funny .

Gill's profile picture
Gill5 days ago

Distilling task specific models on the fly saves huge inference overhead.

icefrog.◎'s profile picture
icefrog.◎6 days ago

the JIT compiler for intelligence is a savage move

REJAUL's profile picture
REJAUL5 days ago

training its replacements could change how quickly AI evolves

Jason Nocco's profile picture
Jason Nocco5 days ago

Could also use to train self learning enabled small domain specific models on the fly, privately. Automatically fine-tuned from grounded knowledge via our integrated knowledge graph. Built so anyone could easily create and quickly use their own models.

青雲's profile picture
青雲5 days ago

这个即时编译的类比很顺,但 JIT 真正的麻烦在别处,是生成出来的那段代码和解释器对同一件事的理解要对齐。 放到这里就是,窄模型拿到的是训练那一刻的任务定义。任务挪一点,它自己看不出来,而且它越窄越不会说我不适用了。通用模型至少见过别的场景,有可能会停下来问一句。 我们那边有一条判据正好是这件事。一个进程拿着过时的游标来预定事件序号,现在的行为是直接报错,而不是产出一条和别人重叠的事件流。以前会默默重叠,它不报错、不丢数据,只是把两条时间线搅在一起。 所以训练替代品这件事,我觉得成败可能不在训练本身,在于替代品能不能说出我这条已经不适用了。 至于提示注入那条我只同意一半。能力更弱确实危害更小,但注入利用的是权限边界,窄模型接了哪些工具才是那条线。

Summer's profile picture
Summer5 days ago

The JIT compiler analogy makes so much sense.

Sagiv Ofek's profile picture
Sagiv Ofek6 days ago

The first generation to train their replacements: us. The second: the models.

Tyzo's profile picture
Tyzo5 days ago

Models training their replacements could accelerate AI development rapidly

Abdul Q.'s profile picture
Abdul Q.5 days ago

Agency version of this: a general model is fine until you need Webflow CMS, Figma props, and B2B SaaS homepage taste in the same pass. Domain muscle still wins on marketing sites.

ZANA's profile picture
ZANA6 days ago

generating targeted lightweight models dynamically solves both the security vulnerabilities and high overhead of general agents

Related Videos