雪踏乌云's banner
雪踏乌云's profile picture

雪踏乌云

@Pluvio9yte • 48,907 subscribers

💭 云安全→独立开发 | AI博主 📌 聚焦:AI & Vibe coding & Prompt & Agent 🌏 在做:创业 | Web出海 | Web3 🌊 持续分享AI编程、AI工具、出海干货和创业经验 ✉️ 合作私信

Shorts

说个暴论,蒸馏就是AI时代的合法抢劫。 使用这个开源项目,任何爆款视频都能快速复制了 现在有了hypit,任何视频都能够很快速的复刻。 这里也给大家一个视频成本忽略不计的秘诀,租卡本地部署MiniMax H3,一个小时一块五到三块钱,视频生成成本几乎忽略不计了。 详细教程和开源地址在引用文章

说个暴论,蒸馏就是AI时代的合法抢劫。 使用这个开源项目,任何爆款视频都能快速复制了 现在有了hypit,任何视频都能够很快速的复刻。 这里也给大家一个视频成本忽略不计的秘诀,租卡本地部署MiniMax H3,一个小时一块五到三块钱,视频生成成本几乎忽略不计了。 详细教程和开源地址在引用文章

301,513 просмотров

Anthropic 在 2026 年 9 月 17 日公布了一组进展:在两名技术人员监督下,Claude 用了不到四周时间,协助优化了 30 多个开源生物学模型,其中包括 AlphaFold3 和 Boltz-2,相关优化代码已经开源。 这件事情吸引人的地方,在于它处理实际科研代码的具体过程,也就是从进入工程仓库、定位问题、修改代码,到最后一步步拿指标做检查。 顺着这种科研工作流的思路,我最近试用了 ScienceBuddy。它是 PhAI Labs 推出的生物医学 Agent 工作台,我先用它跑了一个具体的文献梳理任务。 我给出的提示词是整理近几年用机器学习预测 CRISPR-Cas9 脱靶活性的相关研究,主要看基因编辑在目标区域之外产生误切的情况。 为了方便后续比对,我直接指定了输出结构,要求把研究任务、数据规模、主要发现、研究局限和 DOI 整理成规整的列。 ScienceBuddy 在当前任务窗口里生成了一张论文信息表格。不同研究使用的数据集、得出的结论以及原文入口被归纳在同一套字段下。这样的交互挺省事:输入的提示词、生成的表格以及后续的追问都停留在同一个工作界面,左侧列表也保留着历史任务记录,随时可以点回来继续补充条件。 拿到了整理好的表格,我抽查了其中一行,点击 CRISMER 对应的 DOI 链接,跳转到了 bioRxiv 上的原始预印本摘要。 表格里抓取的 CHANGE-seq 和 SITE-seq,确实是原文摘要里明确写出的训练数据;论文自身报告的 F1 值为 0.728、PR-AUC 值为 0.818,也逐字出现在摘要数据中,页面同时标有预印本提示。这里需要说明,这些准确率数值是 CRISMER 那篇论文自己的实验结果,并不是软件的性能数据。这次抽查做的事情很简单,就是顺着生成结果里的来源链接,确认关键字段能不能在原文里找到出处。 在另一个关于 GEO 公开组学数据的任务中,我查看了它的运行轨迹(Trajectory)。界面上展示了 query_geo、fetch_source、query_pubmed 等工具调用节点,清晰记录了系统在推进任务时尝试发起的查询动作。 关于底层的实现,项目方提到 ScienceIDE 是 ScienceBuddy 的底层技术框架,主要作用是把真实的科研代码封装成 Agent 可以执行、练习和验证的环境,并尝试通过科学标准来检查结果。面向普通用户日常使用的产品是 ScienceBuddy。 整套体验下来,最直观的感受是ScienceBuddy 把文献线索集中整理成表,并在同一个任务流里继续追问,遇到关键数据也能够直接从引用中找到 DOI 回原文做核对。 产品访问地址是 PhAI Labs 开发。我体验的时候 Preview 版本处于免费开放状态。如果平时需要梳理生物医学领域的论文脉络或收集开题证据,可以找一个具体课题试一试。

Anthropic 在 2026 年 9 月 17 日公布了一组进展:在两名技术人员监督下,Claude 用了不到四周时间,协助优化了 30 多个开源生物学模型,其中包括 AlphaFold3 和 Boltz-2,相关优化代码已经开源。 这件事情吸引人的地方,在于它处理实际科研代码的具体过程,也就是从进入工程仓库、定位问题、修改代码,到最后一步步拿指标做检查。 顺着这种科研工作流的思路,我最近试用了 ScienceBuddy。它是 PhAI Labs 推出的生物医学 Agent 工作台,我先用它跑了一个具体的文献梳理任务。 我给出的提示词是整理近几年用机器学习预测 CRISPR-Cas9 脱靶活性的相关研究,主要看基因编辑在目标区域之外产生误切的情况。 为了方便后续比对,我直接指定了输出结构,要求把研究任务、数据规模、主要发现、研究局限和 DOI 整理成规整的列。 ScienceBuddy 在当前任务窗口里生成了一张论文信息表格。不同研究使用的数据集、得出的结论以及原文入口被归纳在同一套字段下。这样的交互挺省事:输入的提示词、生成的表格以及后续的追问都停留在同一个工作界面,左侧列表也保留着历史任务记录,随时可以点回来继续补充条件。 拿到了整理好的表格,我抽查了其中一行,点击 CRISMER 对应的 DOI 链接,跳转到了 bioRxiv 上的原始预印本摘要。 表格里抓取的 CHANGE-seq 和 SITE-seq,确实是原文摘要里明确写出的训练数据;论文自身报告的 F1 值为 0.728、PR-AUC 值为 0.818,也逐字出现在摘要数据中,页面同时标有预印本提示。这里需要说明,这些准确率数值是 CRISMER 那篇论文自己的实验结果,并不是软件的性能数据。这次抽查做的事情很简单,就是顺着生成结果里的来源链接,确认关键字段能不能在原文里找到出处。 在另一个关于 GEO 公开组学数据的任务中,我查看了它的运行轨迹(Trajectory)。界面上展示了 query_geo、fetch_source、query_pubmed 等工具调用节点,清晰记录了系统在推进任务时尝试发起的查询动作。 关于底层的实现,项目方提到 ScienceIDE 是 ScienceBuddy 的底层技术框架,主要作用是把真实的科研代码封装成 Agent 可以执行、练习和验证的环境,并尝试通过科学标准来检查结果。面向普通用户日常使用的产品是 ScienceBuddy。 整套体验下来,最直观的感受是ScienceBuddy 把文献线索集中整理成表,并在同一个任务流里继续追问,遇到关键数据也能够直接从引用中找到 DOI 回原文做核对。 产品访问地址是 PhAI Labs 开发。我体验的时候 Preview 版本处于免费开放状态。如果平时需要梳理生物医学领域的论文脉络或收集开题证据,可以找一个具体课题试一试。

60,621 просмотров

Codex 大更新 —— Chrome 内置最完全功能应用解析 1. 内置浏览器 + Chrome Cookie 一键导入 现在导入了 Cookie 之后,能够把抖音、小红书、推特的账号 Cookie 都导入。现在打开,直接就是登录状态。 2. 开发者模式 可以查 DOM、Console、Network 请求、性能等。 比如说,原来我开发网站如果遇到性能问题,让它优化就会麻烦一些。现在直接用内置的浏览器,速度非常快,这个优化速度。 下面是一些应用场景 内容创作与研究 让 Codex 带着你的账号去小红书、微博、知乎等平台“逛”,找选题、分析爆款、抓用户反馈。 开发者调试 开启开发者模式后,Codex 能实时 inspect 页面: 快速定位 console error、网络请求异常。 分析 DOM 结构、调试已登录态的内部工具。 提示词示例:“开启开发者模式,打开这个页面,检查 Network 请求里和用户数据相关的接口,并把 console 里的 error 列出来。” 测试需要真实登录态的 Web 应用等

Codex 大更新 —— Chrome 内置最完全功能应用解析 1. 内置浏览器 + Chrome Cookie 一键导入 现在导入了 Cookie 之后,能够把抖音、小红书、推特的账号 Cookie 都导入。现在打开,直接就是登录状态。 2. 开发者模式 可以查 DOM、Console、Network 请求、性能等。 比如说,原来我开发网站如果遇到性能问题,让它优化就会麻烦一些。现在直接用内置的浏览器,速度非常快,这个优化速度。 下面是一些应用场景 内容创作与研究 让 Codex 带着你的账号去小红书、微博、知乎等平台“逛”,找选题、分析爆款、抓用户反馈。 开发者调试 开启开发者模式后,Codex 能实时 inspect 页面: 快速定位 console error、网络请求异常。 分析 DOM 结构、调试已登录态的内部工具。 提示词示例:“开启开发者模式,打开这个页面,检查 Network 请求里和用户数据相关的接口,并把 console 里的 error 列出来。” 测试需要真实登录态的 Web 应用等

160,257 просмотров

AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:

AI Is Moving Beyond “Generating Videos” — Toward “Generating Worlds” Over the past two years, AI video models have advanced at an astonishing pace. From Runway and Pika to Sora and Veo, AI-generated videos have become increasingly realistic and more consistent with the physical laws of the real world. Many people believe the next objective is simply to generate videos that are longer, sharper, and more lifelike. But if we take a step back, we can see that the real transformation is not happening in video itself. It is happening in world models. What Is a World Model? In 1943, psychologist Kenneth Craik proposed an idea that would influence artificial intelligence research for decades. He argued that the human brain does not merely react to the outside world. Instead, it maintains an internal model of how the world works. Because we have this internal model, we can predict the outcome of an action before we actually take it. Before crossing a road, we estimate whether a car will pass by. Before catching a ball, we predict its trajectory. These abilities come from continuously simulating the world in our minds, rather than relying entirely on trial and error. This idea later became known by a more formal term: World Model. A world model does not describe a single image or a fixed video clip. It is an internal representation capable of continuously simulating the rules and dynamics of the real world. Why Is AI Research Turning Toward World Models? Because predicting “what comes next” is becoming increasingly central to how AI systems work. Language models predict the next token. Image models predict the next step in the denoising process. Video models predict the next frame. A world model, however, attempts to predict something broader: What should the world look like in the next moment? In 2018, David Ha and Jürgen Schmidhuber proposed in their paper World Models that an intelligent agent could first learn a model of the world, and then use that internal model to plan its actions. The Dreamer series later demonstrated that many complex tasks could be learned by training agents inside an “imagined world.” At the same time, the development of video models such as Sora and Veo led researchers to another realization: A model capable of continuously generating video has already learned, at least implicitly, many of the rules governing the real world. As a result, these two research directions have gradually begun to converge. But Video Is Not Yet a World This is where the distinction is often misunderstood. For a world model to support meaningful real-time interaction, it must solve several critical problems. Most video models today are essentially answering one question: What should the next frame look like? A true world model needs to answer much more: What happens if I take one step forward? If I walk behind a building and then return, will the building still be there? If I suddenly change the camera angle, will the entire space remain consistent? If I enter a command such as: “Summon a dragon.” Will the world respond immediately? In other words, a world model must do more than generate content. It must understand space. It must understand time. It must understand causality. And it must understand interaction. Moving from watching to participating is where the real difficulty of world models begins. World Models Are Entering the Interactive Era One of the latest attempts in this direction is Alaya World, recently open-sourced by Alaya World, or Alaya Lab. Instead of generating a fixed video clip, it generates a world that users can explore in real time. Users can begin with text, an image, or a video, enter the generated scene, move freely through it, and introduce new prompts at any moment during generation. The world responds immediately. According to the publicly released information, Alaya World provides: Real-time streaming generation at 720p and 24 FPS Stable continuous exploration for more than one minute The ability to switch prompts and trigger skills or events during generation Model weights and inference code released under the Apache 2.0 License Training code and datasets planned for future release What makes these capabilities important is not simply the technical specifications. It is that the generated “world” can now support continuous interaction. The official demo shows that users can genuinely control, transform, and explore the generated environment. AI Is Evolving From a Tool Into an Environment Over the past few years, most discussions around AI have focused on content generation. Generating text. Generating images. Generating videos. But world models raise a fundamentally different question: Can AI generate an environment that people can inhabit, explore, and continuously evolve? If the answer is yes, the impact will extend far beyond video generation. Game development, robotics training, embodied intelligence, digital twins, virtual production, and many other fields could be transformed by the development of world models. World models are still at a very early stage. Yet from Craik’s proposal of an internal mental model more than eighty years ago to the emergence of today’s interactive world-generation systems, a clear evolutionary path is beginning to take shape. Perhaps what AI is ultimately learning has never been limited to images, videos, or language. Perhaps it is learning the world itself. References GitHub: Technical Report:

113,347 просмотров

接入anysearch后,我的agent搜索效率提高了 Parallel、Perplexity、Tavily 这些搜索工具给 Agent 用,有个一直没解决好的问题——返回的是链接加摘要,Agent 拿到之后还要自己点开网页、筛选内容、判断哪条有用。金融、学术、代码这些垂直领域的搜索质量更差。输出没有结构,光解析内容就要烧掉大量 token。 分享一个我一直在用的为 AI Agent 设计的搜索基础设施 —AnySearch,一个为 Agent 量身打造的「搜索 Skill」: - 实时网页搜索 - 支持金融、学术、代码、社会媒体等垂直领域搜索 - 直接把网页转成干净的 Markdown(extract) - 输出结构化,Agent 能直接拿来用 安装 AnySearch SKILL 超级简单:复制下方提示词并发送给你的 Agent,即可自动完成安装:

接入anysearch后,我的agent搜索效率提高了 Parallel、Perplexity、Tavily 这些搜索工具给 Agent 用,有个一直没解决好的问题——返回的是链接加摘要,Agent 拿到之后还要自己点开网页、筛选内容、判断哪条有用。金融、学术、代码这些垂直领域的搜索质量更差。输出没有结构,光解析内容就要烧掉大量 token。 分享一个我一直在用的为 AI Agent 设计的搜索基础设施 —AnySearch,一个为 Agent 量身打造的「搜索 Skill」: - 实时网页搜索 - 支持金融、学术、代码、社会媒体等垂直领域搜索 - 直接把网页转成干净的 Markdown(extract) - 输出结构化,Agent 能直接拿来用 安装 AnySearch SKILL 超级简单:复制下方提示词并发送给你的 Agent,即可自动完成安装:

95,088 просмотров

抖音起号做出爆款诀窍:首帧生成 I2V。 如果你的目标是做出爆款AIGC短视频起号,而不是出长期稳定的动画剧集,那么我更推荐使用视频生成模型的首帧生成模式,该模式对于seedance系列与minimax-h3都是适用的。 一开始我在抖音上调研时,发现了一个爆款的 AIGC 博主。他的视频风格非常华丽,角色动作干脆利落,特效炫酷。而且他非常慷慨,把提示词都分享了出来(是真的自用提示词,甚至连他的本地缓存图片路径都露出来了) 结果我看他的提示词,虽然动作很华丽,但编排其实特别简单,与我们想象中长篇大论的那种“导演视角”做出来的效果完全不一样。 他分享的真正诀窍在于:使用视频模型首帧生成。只要规定好首帧的场景、人物等约束,并通过首帧画面定住风格,后续画面就会非常统一,哪怕只用简单的提示词,都能做出非常震撼的效果。 最后从技术层面解释一下,为什么首帧生成比上传参考图,设定图更具有优势。 Seedance 官方有写说明文档:若要严格保证开场画面和指定图片一致,优先走真正的图生视频首帧。 官方论文提到,I2V 主流做法是把给定图像当作时间轴上的第 0 帧: 图像 latent 与噪声拼接、首帧特征进时空注意力、甚至用图像噪声先验做残差预测。 后续帧是在「已经钉死的像素布局」上长出来的运动。 参考图常见路径更偏语义嵌入或特征引导,能够保证「像不像这个人 / 这种风」,不保证第 1 帧构图、光影、材质、镜头焦段和你上传的那张图像素级对齐。 不过首帧参考图的生成也是有很多门道,这点是被很多人忽略,也是那些爆款视频不说的秘诀(是的,那位抖音博主分享了所有提示词,唯独没有分享首帧怎么生成的,属于是真正的技术壁垒) 我自己深入扒了两天,把里面的门道也探索了个七七八八,下期给大家分享我自己实操总结的首帧生成技巧。 最后依然是展示视频效果,搭配bgm食用更佳

抖音起号做出爆款诀窍:首帧生成 I2V。 如果你的目标是做出爆款AIGC短视频起号,而不是出长期稳定的动画剧集,那么我更推荐使用视频生成模型的首帧生成模式,该模式对于seedance系列与minimax-h3都是适用的。 一开始我在抖音上调研时,发现了一个爆款的 AIGC 博主。他的视频风格非常华丽,角色动作干脆利落,特效炫酷。而且他非常慷慨,把提示词都分享了出来(是真的自用提示词,甚至连他的本地缓存图片路径都露出来了) 结果我看他的提示词,虽然动作很华丽,但编排其实特别简单,与我们想象中长篇大论的那种“导演视角”做出来的效果完全不一样。 他分享的真正诀窍在于:使用视频模型首帧生成。只要规定好首帧的场景、人物等约束,并通过首帧画面定住风格,后续画面就会非常统一,哪怕只用简单的提示词,都能做出非常震撼的效果。 最后从技术层面解释一下,为什么首帧生成比上传参考图,设定图更具有优势。 Seedance 官方有写说明文档:若要严格保证开场画面和指定图片一致,优先走真正的图生视频首帧。 官方论文提到,I2V 主流做法是把给定图像当作时间轴上的第 0 帧: 图像 latent 与噪声拼接、首帧特征进时空注意力、甚至用图像噪声先验做残差预测。 后续帧是在「已经钉死的像素布局」上长出来的运动。 参考图常见路径更偏语义嵌入或特征引导,能够保证「像不像这个人 / 这种风」,不保证第 1 帧构图、光影、材质、镜头焦段和你上传的那张图像素级对齐。 不过首帧参考图的生成也是有很多门道,这点是被很多人忽略,也是那些爆款视频不说的秘诀(是的,那位抖音博主分享了所有提示词,唯独没有分享首帧怎么生成的,属于是真正的技术壁垒) 我自己深入扒了两天,把里面的门道也探索了个七七八八,下期给大家分享我自己实操总结的首帧生成技巧。 最后依然是展示视频效果,搭配bgm食用更佳

39,840 просмотров

让Codex调用Grok之后,我的工作流被大大优化了。 Grok Build是Grok的CLI,能够被目前的Codex最强的SOL模型这一最强大脑调用 Sol做规划或者思考大脑,然后调用Grok去X上能够搜索很多技巧或者帖子 调用方式直接用自然语言告诉Codex调用本地的Grok Build或者用下面这个skill

让Codex调用Grok之后,我的工作流被大大优化了。 Grok Build是Grok的CLI,能够被目前的Codex最强的SOL模型这一最强大脑调用 Sol做规划或者思考大脑,然后调用Grok去X上能够搜索很多技巧或者帖子 调用方式直接用自然语言告诉Codex调用本地的Grok Build或者用下面这个skill

56,782 просмотров

又给大家送福利了! 经常有做设计、做自媒体或者搞开发的朋友问我:现在做批量出图和生图应用,调用商业 API 成本太高怎么办? 今天必须给大家安利一个超良心的平台——Flatrouter。 原因特别实在,他们把生图价格打到了地板价:最新上架的 GPT-Image-2.5,单张只要 0.05 元! 而且这次 OpenAI 刚发布的 GPT Image 2.5,可以说是整体提升非常暴力,下面是模型概况: Flare:延迟比上代直接减半,响应飞快。适合快速出图或做概念图验证。 Sunburst(画质天花板·精修档模型):最强基座,主打极致精度。光影漫反射、皮肤毛孔、布料纹理、金属质感表现细腻,连细微指令(比如左右手朝向)都能精准执行,专攻广告级成片与高要求物料。 根据社区评价,主要特性有下面 4 点: 1.多轮改图不崩、不换脸 2.指哪改哪,极度听话:说“只换背景”就绝不动前景主体,甚至能精准修改局部细节,不会冲掉画面原有的品牌特征与结构。 3.中文排版不再乱码:海报、包装盒、杂志排版中的中文字符极其稳定,直接能拿来做商业版面。 4.原生透明背景 + 4K 超清:最高支持 4K 级别分辨率 最新最强的大模型,配上 0.05 元一张 的神仙价格: 🔗

又给大家送福利了! 经常有做设计、做自媒体或者搞开发的朋友问我:现在做批量出图和生图应用,调用商业 API 成本太高怎么办? 今天必须给大家安利一个超良心的平台——Flatrouter。 原因特别实在,他们把生图价格打到了地板价:最新上架的 GPT-Image-2.5,单张只要 0.05 元! 而且这次 OpenAI 刚发布的 GPT Image 2.5,可以说是整体提升非常暴力,下面是模型概况: Flare:延迟比上代直接减半,响应飞快。适合快速出图或做概念图验证。 Sunburst(画质天花板·精修档模型):最强基座,主打极致精度。光影漫反射、皮肤毛孔、布料纹理、金属质感表现细腻,连细微指令(比如左右手朝向)都能精准执行,专攻广告级成片与高要求物料。 根据社区评价,主要特性有下面 4 点: 1.多轮改图不崩、不换脸 2.指哪改哪,极度听话:说“只换背景”就绝不动前景主体,甚至能精准修改局部细节,不会冲掉画面原有的品牌特征与结构。 3.中文排版不再乱码:海报、包装盒、杂志排版中的中文字符极其稳定,直接能拿来做商业版面。 4.原生透明背景 + 4K 超清:最高支持 4K 级别分辨率 最新最强的大模型,配上 0.05 元一张 的神仙价格: 🔗

25,549 просмотров

将最近全网爆火的分手片段复刻了一下。 同样的提示词,来看 Minimax H3 和 Seedance 2.5 的实测对比,差异这么大啊 Seedance2.5明显人物对微表情控制的更好。

将最近全网爆火的分手片段复刻了一下。 同样的提示词,来看 Minimax H3 和 Seedance 2.5 的实测对比,差异这么大啊 Seedance2.5明显人物对微表情控制的更好。

39,545 просмотров

Claude Fable 5 经过1个月紧急下架后再回归,我这两天也重新去测试了一下。 正好一直刷到赛博斗蛐蛐的视频,正好拿来测试一下,看看能不能做个同款出来。 给大伙看看fable 5 的效果:视频1 再来对比一下GPT-5.5的效果,老实说有点拉了,为什么把restart按钮直接盖住了屏幕啊?不是很理解:视频2 最后再来看看国产Deepseek V4 Pro的效果,物理效果感觉还是不行,有点迟钝的感觉:视频3 Fable 5的整体效果是最接近我刷到的小视频的效果的。 虽然效果很强,但是现在fable 5最大的问题在于,官方在6月22日以后,订阅套餐不能用Fable,只能走API。 但是价格实在是遭不住:每 M 输出 50 美元的定价让社区吐槽声不断!我推荐用Zenmux,这个省一些。目前有个加赠20%的活动,每人最高可享 300刀 ! 官网在这里: RPM、不限流。可调用全平台 200+ 模型(含 Claude Fable 5),非常适合做这种评测、并排对比和长期开发。

Claude Fable 5 经过1个月紧急下架后再回归,我这两天也重新去测试了一下。 正好一直刷到赛博斗蛐蛐的视频,正好拿来测试一下,看看能不能做个同款出来。 给大伙看看fable 5 的效果:视频1 再来对比一下GPT-5.5的效果,老实说有点拉了,为什么把restart按钮直接盖住了屏幕啊?不是很理解:视频2 最后再来看看国产Deepseek V4 Pro的效果,物理效果感觉还是不行,有点迟钝的感觉:视频3 Fable 5的整体效果是最接近我刷到的小视频的效果的。 虽然效果很强,但是现在fable 5最大的问题在于,官方在6月22日以后,订阅套餐不能用Fable,只能走API。 但是价格实在是遭不住:每 M 输出 50 美元的定价让社区吐槽声不断!我推荐用Zenmux,这个省一些。目前有个加赠20%的活动,每人最高可享 300刀 ! 官网在这里: RPM、不限流。可调用全平台 200+ 模型(含 Claude Fable 5),非常适合做这种评测、并排对比和长期开发。

48,569 просмотров

I went a little overboard with Codex last week and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.

I went a little overboard with Codex last week and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.

23,107 просмотров

我的视频复刻Skill在Sol模型的加持下,现在非常牛逼,只要你觉得某个视频是Remotion或者Hyperframes做的,都能复刻。 评论区还有视频样例和开源地址

我的视频复刻Skill在Sol模型的加持下,现在非常牛逼,只要你觉得某个视频是Remotion或者Hyperframes做的,都能复刻。 评论区还有视频样例和开源地址

24,078 просмотров

Videos

Pluvio9yte's profile picture

这个产品有点🐂❗ 一条已经生成完成的 AI 视频,竟然还能如此细致地修改,做AI视频再也不用担心“不可编辑”、反复抽卡耗费高昂成本了。 为了证明我没有吹牛,可以先看下面的展示视频 我们以前用AI模型做视频,通常只能生成几十秒的单个片段,一旦有个细节错了,就得全部推翻重新生成。 凑齐画面后,还得把这些“零件”导入到 PR 或剪映里繁琐地手动拼接、加音频和字幕,频繁切换多个软件,一条视频做下来费时费力。 但我这次用 Fotor Video Agent 做了一条 NovaQuiet 降噪耳机产品视频,体验完全不同。 第一轮,我只需提供产品 Brief、晚高峰地铁场景、核心文案,以及需要的画面要求。 Fotor Video Agent 依托长时序规划,直接生成了视频初稿,并自动将素材、文字、字幕和声音全部编排到了多轨道时间轴上。在一个系统内就能完成所有制作,根本无需切换其他软件。 初稿出来后,我觉得人物和耳机主体不够突出。以前遇到这种情况只能重头再来,但 Fotor Video Agent 生成的是支持二次修改的工程文件,所有元素都能继续调整。 于是我直接在聊天里进行第二轮编辑:“给人物加一层透明度约 20% 的冷白轮廓光,让耳罩旁的声场数值从 72 dB 降到 18 dB,再让声波、颗粒和字幕主动避开人物。” 这些要求原本需要拆成好几个复杂的剪辑动作:跟踪人物轮廓、处理描边透明度、制作数字变化、逐个检查字幕挡脸等。 但在 Agent 里面,我只需要把想要的效果讲清楚。 Agent 接着就在多轨时间轴里修改了原来的项目。人物从拥挤的地铁画面里被凸显出来,声场数值精准变化,周围的视觉元素也开始智能避让人物。 这正是 Fotor Video Agent 这次最值得介绍的能力。它通过自动规划和编排,能够让你在单一系统内高效出片。 原本需要 3-5 天的繁琐拼接修改工作,现在差不多 40 分钟就能搞定,并且还能生成 4K MG动画,元素、数据、文案、图标位置等都可以精细化编辑,综合成本大幅降低。 如果你也有做AI视频反复抽卡的痛点,不妨去试一下Fotor 体验地址

雪踏乌云

209,754 просмотров • 29 дней назад

Pluvio9yte's profile picture

过去做一次深入的竞品调研,你往往需要买一堆每个月上百刀的订阅,然后在 Similarweb、Ahrefs、Reddit 和 X 之间来回切网页,最后填到表里的数据还经常对不上。 但现在,你不需要单独买任何一个月费订阅,也不用再搞一堆复杂的 API 密钥。 你只需要把产品网址丢给已经接好 AIsa 的 Claude Code 或 Codex,用一句最实在的业务目标,把它打发去干活。Agent 会自己去查产品事实、找竞品、比流量渠道、翻用户原话,最后交给你一份能直接拿去开会讨论的竞品报告。 我们可以直接看一个最典型的场景。这是我扔给 Agent 的原话: “请分析这个产品 Youmind,识别 5 个主要竞品。比较它们的定位、核心功能、网站流量来源、用户评价和内容渠道,最后输出一份竞品分析报告,并给出 3 个值得切入的差异化方向。” 注意,我没有指定它该调哪个接口。规划任务、筛选数据源、交叉核对,全权交由 Agent 自己完成。 真正接通了外部商业数据的 Agent,在这个过程中会走出一条非常清晰的思考链: 先读官网,搞清楚这产品卖给谁、主打哪一句; 接着用 Tavily 或 Exa 把公开事实和候选竞品补全; 随后调用 Similarweb,把几家网站的访问规模、渠道结构和搜索表现直接拉到同一张桌子上; 同时,Reddit 和 X 负责另一半——去看看真实用户在夸什么、骂什么、在什么场景下反复提及; 最后收口,一份包含定位、功能、流量来源、评价、内容渠道的报告生成完毕,并顺手给你指出了 3 个还能挤进去的差异化方向。 这里面最核心的一层,必须单拎出来说一下,那就是底层的 Similarweb 数据。 AIsa 是目前极少数能把 Similarweb 官方 API V5 直接交给 Agent 调用的平台。走的是正规授权接口,和市面上那些爬虫、镜像站压根不是一路东西,而且那些镜像站最多只能让你在网页上查查数据,根本不提供API。 你要知道,这种官方服务主要面向企业客户,个人开发者和小团队自己去谈,一年从 5 万美元起步很常见。哪怕是 Ahrefs、Semrush,各开一份进阶订阅,一个月也得大几百刀。 而 AIsa 的做法是:把这些顶级情报源全部打包到一个账户里,共享余额,按次扣费。这意味着,Agent 每次帮你做竞品调研,你不用再同时订阅多份服务。 一份真正敢拿来做商业决策的报告,必须留下证据。 在这份报告里,哪段数据来自 Similarweb,哪段是 Reddit 的用户原话,哪段来自搜索结果,全都能精准指回去。模型自己补的判断,和接口返回的硬数据,被分得清清楚楚。 以后你拿这份报告做决策,至少知道哪一行经得起推敲复查。 顺着这套打法,你能做的事还有很多: 产品 SEO 检测:丢一个网站 URL,它能出 SEO 诊断,告诉你关键词覆盖率、SERP 表现,以及最该先修哪几个页面。 GEO / AI 可见度检测:丢品牌名加几个竞品,它能直接告诉你,在 Perplexity 这类 AI 答案页里谁被提得多、引用从哪来、你的内容该往哪补。 产品地址: 查看 AIsa 支持的完整 Premium API 目录: 探索完整的 Go-to-Market 解决方案:

雪踏乌云

39,001 просмотров • 25 дней назад

Pluvio9yte's profile picture

告诉大家一个反直觉的现象,用自媒体一个月盈利七位数的人,他们往往不只是内容做的好。 最近两条打法大家应该都看见了。 天策:做号,后端分成,挂中转、开通、社群,一个月百万营收 dontbesilent栋哥:元自媒体,教人如何做自媒体,一年千万级别 想要赚得多,往往是要做好商业化。最近体验了一个产品,能够帮我们验证商业化的可行性。我把这个问题丢给了 atypica.AI,接下来带着大家走一下商业分析全链路。 正如视频中所示,他先详细拆解了我的需求。拆解完需求之后,又对我进行了几轮问话,最终将需求澄清。 澄清好了之后,他开始进行调研,并最终模拟真人进行高强度的问答,给出了最终的商业化建议方案 这个调研过程不是让 ChatGPT 装用户,它先找了 5 个会掏钱的人:写代码的、做设计的、职场小白、县城开店的、带孩子想再赚钱的。一个一个问,这些人设是基于真实的社交媒体内容和用户访谈语料,经过产品自研的模型清洗、多维度分析后生成的。这些商业调研问题上,和真人回答大约八到九成接近。 问完就一件事:以上三种服务,先做哪一种? 结论竟然很准确,走的最直接的变现路径。先挂中转。社群往后放。帮人开通 ChatGPT,现在先别上。 写代码的人说:钱要砸在工具稳不稳上,别买心理按摩。凌晨接口挂了,社群再暖也没用。 开店的人说:万一号封了,钱打水漂。光开个账号不会用,那就是空壳。 五个人里,四个不买代开通。怕的是封一次,把你整个号的信任一起带走。 中转他们能接受,因为用了就知道有没有用。价就在 99 到 199 一个月。说到「秒开、吊打官方」,设计师那种人直接划走。 社群不是不能卖。要等工具先卖出去,人觉得你靠谱了再卖。一上来卖圈子、灌鸡汤、晒收款,前面提到县城开店这类人会看一眼就走。 综合下来,我用下来的感受是,这些很接近于真实的人设,你会感受到背后确实是基于真实的人的需求来进行的。而且往往不同的人格,他们的需求很精准,对应的就能细化出不同的商业方向。这和笼统的让 ChatGPT 去模仿一个用户相比,是 ChatGPT 所达不到的水准。 通过和模拟的用户进行交流,更加能够确定我们的商业化方向。 俗话说“选品决定生死”,其实某种程度上,商业化的路径也会决定生死。而选品离不开人的需求。 现在新注册的用户会赠送 100w tokens,大概能跑 1 到 2 次调研。如果你目前有任何不清楚的想法,或者想落地什么商业项目,都可以去聊一下 地址:

雪踏乌云

49,604 просмотров • 1 месяц назад

Pluvio9yte's profile picture

用了两天的Fable-5,得到了一个与大部分人截然相反的结论,实测体验如下: 速度:非常慢,尤其是开始分析和规划的时候能感受到像乌龟慢爬 价格:我是Max5订阅用户,虽然输入输出定价为Opus的两倍,但是实际消耗量并没有像流水一样,消耗速度确实变快了,但命中缓存后消耗没有特别高,没有人传人的消耗那么离谱的快 所以铺天盖地说消耗非常快的只有三种可能: 第一种是Pro订阅账号容量本来就小, 第二种是没自己账号纯造谣, 第三种是本来对Opus的消耗没有感知,看到别人说fable消耗快自己也一用发现消耗快于是觉得是真的消耗快了。本身Opus的消耗就非常快嘛,现在只不过是变为了原来的大概1.5x 对于我来说,我通常只用Opus,所以这个1.x倍Opus的消耗量能够接受。实际使用中建议把缓存重建修改为1小时。 能力:能力强体现在思维边界更广,架构能力更强。 对于前端能力测试确实是有很大提高,尤其是创意能力。真正落地到工程项目中,我的体感是能够真正发现GPT-5.5的问题了,不会再像之前基本一直认为GPT的分析是对的。 实际能力类似于Claude Opus4.6++与GPT-5.5++的结合版,目前没有到非常惊艳的程度 结合价格和6月22之后就没了,我也没有升级Max20的打算,Max5+GPT-5.5乃至即将出现的5.6我觉得完全够用。也建议大家理性消费。如果想走按量体验可以看下评论区。 最后给大家看一下昨晚用Fable简单构建的前端的效果,关于为什么我吐槽过还要做前端,是因为我也需要在抖音小红书炸裂搞流量,人作为视觉动物,前端是最开始抓住注意力的开头的。 视频如下,网站部署在verel地址放在评论区了,大家可以自行查看

雪踏乌云

24,406 просмотров • 3 месяцев назад