Loading video...
Video Failed to Load
伙伴们在尝试让 AI 改网页的时候,有没有在这一步卡住: “把按钮挪到标题下面,去把这块删掉,还要旁边那张卡突出一点。” 看着页面觉得很清楚,但是写成文字,却得去解释到底是去搞哪个按钮、哪张卡。能不能直接在截图上圈出来,让 AI 照着改? 我用 Ant Ling Ling-3.0-flash-VL 做了个小测试。准备了一张活动报名页,同时把源码和标注截图一起交给它: 红圈里的横幅删掉,按钮移到蓝箭头指向的位置,黄框里的票种设为唯一推荐项。 同一个任务测了三次,三次首版都把按钮放错了。 接着,我把前两次实际渲染出来的页面截图回传给它,让它继续修改。两次都在一轮反馈后通过了全部 17 项验收。第三次停在首版,保留为失败;这组最终完成了 2/3。 这里模型负责看图、修改代码,本地工具负责运行页面、截图和验收。 如果你也在用 AI 改落地页、活动页,这次测试有个值得试的方法: 先在截图上标明要改哪里;首版不对,就把实际页面再截给它看,指出没完成的地方。修改前也写清哪些内容要保留,避免挪好了按钮,却改掉价格或弄坏交互。 这也是我为什么设了 17 项检查:除了位置和推荐卡,还要核对原有文字,点开、关闭报名弹窗,展开 FAQ,再看手机端有没有遮挡、溢出和控制台报错。页面看着改好了,还得点一遍才知道能不能用。 不过,换个目标,它仍然会出错。我把黄框改到“观察席”,它却继续推荐“共创席”。这个变体只测了一次,失败了。 所以,在这组样本里,我看到了模型收到页面截图后修正代码的效果;但换一个位置、换一张卡,仍然需要重新检查。还不能把整页修改放心交给它后就不管了。 57 秒视频里放了标注输入、反馈后的变化、成功和失败结果,标注为已生成结果回放。 你可以照这个过程试自己的页面,也能看到哪些地方需要留给人验收。 - 蚂蚁数科国际站 : - 蚂蚁数科 maas 国内站: - Ling studio: 这几个都可以,目前可用的体验地址
14,318 views • 22 days ago •via X (Twitter)
46 Comments

@AntLingAGI The visual feedback loop approach is brilliant! Feeding rendered screenshots back into the model to refine layout precision is definitely the way forward for AI UI design. 🖼️💻

@AntLingAGI This is exactly the kind of video workflow breakthrough people have been waiting for. The model giving precise timestamps plus full editing suggestions including transitions and BGM turns hours of manual scrubbing into minutes. Really practical for both pros and side creators.

Relying on text to fix visual bugs is almost always slower because spatial instructions get messy fast. Direct interaction tools where you can just circle an element on a canvas and tell the model to "move this here" or "change this color" completely eliminate that translation layer. Sending back a screenshot of the corrected visual state also builds a perfect error handling chain, letting the agent see immediately where its code output didn't quite hit the target so it can iterate and fix the geometry, padding, or CSS rules without you writing another confusing paragraph.

@AntLingAGI Although... but actually, the essence of UI design is spatial topology, so it should use spatial language for dialogue, which also makes it convenient for large models to perform error correction.

@AntLingAGI 对的,再怎么不如看的直观

@AntLingAGI Sending back the actual rendered images before making changes is a pretty practical approach; only after passing all 17 inspections can it be considered truly fixed.

@AntLingAGI Annotated screenshots are way clearer than pure text descriptions—I've fallen into that pit of misplacing buttons too. Send back the actual rendered image for another round of tweaks, and the success rate shoots up; this closed loop is super practical.

@AntLingAGI 哈哈挺好的

@AntLingAGI 截图标注 + 实际渲染反馈,这个工作流很实用。 尤其是让 AI 自己看最终页面再修,比单纯描述问题更容易定位偏差。

@AntLingAGI Pointing out layout changes visually is much easier than typing, modifying website code through simple screenshots feels like real magic

@AntLingAGI This is super amazing

@AntLingAGI First pass missing the button three times is the data. Screenshot-back fixing 2/3 is the method. 17 checks including the modal and mobile overflow is why this isn’t just “it looks closer.”

@AntLingAGI When revising the landing page, I often get stuck on descriptions like "move that button down a bit" or "make that card next to it stand out a little more," and end up going back and forth with the AI, explaining for ages.

@AntLingAGI This is actually pretty useful. Being able to point at the exact part of a page and let AI fix it saves a lot of back and forth.

@AntLingAGI 是的。省事

What stood out to me is how quickly the model corrected the button placement once it saw the actual output. That suggests the issue is less about understanding intent and more about grounding the code in the current visual state. Still, the failure when you changed the recommended card shows the model is still brittle on spatial generalisation.

@AntLingAGI 模型能看懂截图以后,改页面的沟通确实方便了。

@AntLingAGI 是的,这波长眼睛了

@AntLingAGI Seventeen checks including the modal, FAQ, and mobile overflow is the right bar. Looking fixed is not the same as the form still opening.

@AntLingAGI 这个不要太真实!😂有时就是明明一眼就知道该改哪儿,但跟AI描述半天反而说不清!现在这种直接圈出来,再把改完的截图丢回去让它自己纠错,比单纯靠文字指挥靠谱太多了...

@AntLingAGI 哈哈没错的,让他懂意思多少需要功夫

@AntLingAGI This is really useful

@AntLingAGI 有用就好

@AntLingAGI Visual feedback makes AI web editing far more practical

@AntLingAGI 太懂这个卡点了!“把那个按钮挪到下面,把这块删掉” 用文字根本说不清,AI每次都改错地方。直接在截图上画红圈、蓝箭头、黄框让Ling-3.0-flash-VL照着改,这个视觉标注的思路太对了。源码+标注截图一起给,再配合本地截图反馈闭环,2/3通过17项验收已经很强了。

@AntLingAGI 改网页最烦的就是解释位置,直接截图圈出来省事太多。Ling-3.0-flash-VL再配合渲染截图回传继续修,这套流程挺贴近实际使用

@AntLingAGI 哈哈主打实用

@AntLingAGI I’ve struggled with this exact problem — knowing exactly what I want changed on a page but struggling to explain it clearly in text. Your approach of marking it directly on the screenshot and then showing the live render back to the model is a much smarter way.

@AntLingAGI 57 秒视频里放了标注输入、反馈后的变化、成功和失败结果,标注为已生成结果回放。

@AntLingAGI You nailed a very common pain point: the gap between what we see clearly on screen and how hard it is to describe it in words. Using Ling-3.0-flash-VL with direct visual markup and rendered feedback is a smart solution.

@AntLingAGI That first task "move button to yellow box" failing 3 times then passing after 2 refinements is so real. Did you notice if it got better at understanding the red circle annotation after the first round?

@AntLingAGI Marks on a signup page, then code. That is the hard version of a layout ask. Most models only remember that a button had to move.

@AntLingAGI 这个流程可以直接拿去改活动页。我这边也试过类似做法:先标位置,再把实际渲染图回传。最大坑不是挪按钮,是一改布局就把原有文案或交互状态弄丢。

@AntLingAGI 很强了

@AntLingAGI 带 vision 的模型就是爽,指哪改哪

@AntLingAGI 哈哈确实了

@AntLingAGI 圈图改页比打字清楚,一轮反馈就能过验收,但换位置还容易错,最后还是得人把关。

@AntLingAGI 是的呢

The visual feedback loop you showed is exactly the practical method many of us need when asking AI to edit web pages. Circling elements directly on the screenshot and then feeding the actual rendered page back for correction feels like a real breakthrough compared to describing changes in text.

@AntLingAGI 这个测试方法很专业啊!模型负责看图改代码,本地工具负责跑页面截图验收,渲染反馈再改,完整闭环了。三次首版都把按钮放错很真实,但能通过一轮反馈就修正通过,说明Ling-3.0-flash-VL的视觉定位+代码修改能力是真能用在落地页场景了,值得试试。

@AntLingAGI Receipt with 30 yuan error caught - 2092.10 total, 12 items. Most VL fails on small text OCR, Ling didn't

@AntLingAGI Have you experimented with giving the model a short checklist of "must preserve" elements (price, interactions, mobile constraints) right in the annotated screenshot prompt, or does that still get ignored until the feedback round?

@AntLingAGI This visual feedback loop for web edits is super smart. Drawing circles and arrows on the screenshot, then showing the model its own rendered output, actually gets Ling-3.0-flash-VL to fix the mistakes reliably.

@AntLingAGI 虽然……但是其实UI设计的本质是空间拓扑,就该用空间语言去对话,这样也方便大模型来进行纠错

@AntLingAGI 哈哈确实如此

@AntLingAGI I would steal the mark language. Red delete, blue move, yellow recommend. Talking about that card always picks the wrong one.
