Han Xiao's banner
Han Xiao's profile picture

Han Xiao

@hxiao20,794 subscribers

VP, AI @Elastic prev: founder & ceo @JinaAI_

Shorts

I love Clawdbot, but most parts can be just Claude Code --dangerously-skip-permissions + pipe via Telegram. Made a simple version using cloudflare tunnel + tmux + StopHook.

I love Clawdbot, but most parts can be just Claude Code --dangerously-skip-permissions + pipe via Telegram. Made a simple version using cloudflare tunnel + tmux + StopHook.

317,139 görüntüleme

Saw a poster at ICLR showing contrastive learning (InfoNCE included!) ≈ a closed-form spectral decomposition in RKHS. Got curious whether a map could adapt embedding spaces across model families, while preserving NDCG retrieval perf. Did some experiments on my flights back to sf. Video shows diff maps on simple swiss roll data.

Saw a poster at ICLR showing contrastive learning (InfoNCE included!) ≈ a closed-form spectral decomposition in RKHS. Got curious whether a map could adapt embedding spaces across model families, while preserving NDCG retrieval perf. Did some experiments on my flights back to sf. Video shows diff maps on simple swiss roll data.

14,349 görüntüleme

Pure MLX of UMAP, t-SNE, PaCMAP, TriMap, DREAMS, and CNE for Apple Silicon. Beyond the algorithms, I also built a circle-splatting renderer in MLX - scatter-add alpha blending on Metal, piped to H264 hardware encoding. So 70K points, 784 dimensions, from raw data to rendered video < 5s on M3 Ultra. pip install mlx-vis Paper: Code:

Pure MLX of UMAP, t-SNE, PaCMAP, TriMap, DREAMS, and CNE for Apple Silicon. Beyond the algorithms, I also built a circle-splatting renderer in MLX - scatter-add alpha blending on Metal, piped to H264 hardware encoding. So 70K points, 784 dimensions, from raw data to rendered video < 5s on M3 Ultra. pip install mlx-vis Paper: Code:

15,950 görüntüleme

Videos

hxiao's profile picture

在《金庸群侠传》crpg-bench上跑了下 Fable 5.1 感觉并没有很惊艳,在玩游戏上智商水平和fable5/opus5/sonnet5没有显著差异。金庸群侠传bench是我用来测试大模型在long-horizon task能力的一个沙盒。评测方法很简单:通过截取游戏画面,视觉反馈,执行动作,不断循环,最终找齐14本天书完成游戏。玩过《金庸群侠传》的都懂: - 游戏是开放世界,通过升级招人打怪,最终集齐14本天书结束游戏,所以是彻头彻尾的long-horizon task。 - 游戏采用2.5D斜45度视角,加上1996年的320x200分辨率,绝对是VLM的噩梦:怪视角+渣画质+分布外。 - Caveat:虽然是开放式,但开局必先去找南贤北丑,不然寸步难行而且很容易误闯故事线直接被打死。 至少我一开始是这么设计的:完全根据游戏的进度来设计评测指标。但实际做了一段测试后发现,很多frontier大模型连走出始发地都非常困难,更别提集齐14本天书。我也特地设计了相应的SKILL,让模型能对游戏内容任务和操作有基本的了解。为了避免消耗无意义的token,评测就局限在给定20分钟的时间下,谁能探索的最多就算谁厉害些。也就基本沦为了2.5D视角下视觉迷宫问题。 经过测试了一些国产和国外的模型后,20分钟内能顺利走出出生地的模型是: - Fable 5.1:4分钟 - Fable 5:3分钟 - Opus 5:5分钟 - Sonnet 5:6分钟 - Gemini-3.7-flash:12分钟 并且上述这些模型没有一个成功找到南贤居。

Han Xiao

40,742 görüntüleme • 14 gün önce

Daha fazla içerik yok.