Okay this is kinda wild 👀 NVIDIA is basically... handing out FREE API keys to 100+ top AI models - GLM 5.2, DeepSeek V4, Kimi K2.6, MiniMax M3, their own Nemotron and a ton more. it's called NVIDIA NIM and I had no idea it existed. it's rate-limited, so not something you'd run in production - but for personal use it's honestly great. you can poke a dozen frontier models, and some you can even grab and self-host. dropped a key into OpenCode and it just worked. the hyped ones like GLM 5.2 are jammed with queue right now, but the lighter models fly. all of it free.show more

Stefan 3D AI
26,272 次观看 • 2 个月前
You can now use GPT 5.5, Gemini 3.7 Flash,... Kimi K3 and 47 other AI models completely free😱 No subscription. No credit card. Even the API usage costs $0. AIHubMix just opened a free catalog with 50 AI models. Some of the available models: • Ox Alpha • Gemini 3.7 Flash • GLM 5.2 • Kimi K3 • MiniMax M3 • GPT 5.5 • 40+ more And you don’t need separate API keys for each model. Setup takes 2 minutes: > Step 1: Go to > Create an account using your email or OAuth. No card needed. Step 2: Create one API key > The same key works with every free and paid model. Step 3: Add it to any OpenAI-compatible tool Base URL: Then choose any model ending in -free, such as: coding-glm-5.2-free gpt-5.5-free That’s it. One API key. 50 AI models. $0 for both input and output. Save this. You might need a free multi-model setup later.show more

CDG
15,388 次观看 • 1 个月前
Whoa, GLM-5.2 is INSANE for UI/UX with the right... prompting. Open models have finally closed the gap. You have to push it, but it's super creative. I just asked for a beautiful personal site; this was all the model's idea (multiple turns):show more

Anshu
178,213 次观看 • 3 个月前
You can now use GLM-5.3 Flash completely FREE 😱... GLM-5.3 is now officially out and free on No bank card needed. You can also use Qwen3.8, HY3, MiMo-V2.5 and other models for $0. Setup: 2 mins: > Go to > Create a free account > Pick any model showing 0 credits > Use it in chat or create an API key No official end date yet. Grab it before the free access disappears.show more

CDG
13,405 次观看 • 29 天前
GLM 5.3 is FREE right now and you can... actually use it. you can access it through ZenMux with a free API key, and Z. ai is also offering free access rn. you get: - GLM 5.3 - 1M context window - Up to 128K output - Tool calling + MCP - Free API access ZenMux currently lists GLM 5.3 at $0 for both input and output. to get started: > Go to > Create an account > Generate an API key > Use as the base URL > Select z-ai/glm-5.3-free you can also test GLM 5.3 directly at free access can change, so this is worth testing while it’s availableshow more

MIKE
42,231 次观看 • 1 个月前
Big win for open-source LLMs! DeepSeek V4 Pro holds... the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.show more

Avi Chawla
44,124 次观看 • 2 个月前
DeepSeek R1 is *the* best model available right now.... It's at the level of o1, but you can use it for free, and it's much faster. A huge leap forward that nobody saw coming. No wonder so many people are throwing tantrums online trying to discredit the Chinese students who built this. You can use DeepSeek in Visual Studio Code right now: 1. Install the Qodo Gen AI extension 2. Select DeepSeek R1 from their list of models The Qodo team is hosting DeepSeek on their servers, so none of your data will go to China. I've been building a Tetris game using DeepSeek, and this is the most impressive model I've seen so far.show more

Santiago
1,224,409 次观看 • 1 年前
deepseek v4 pro is basically free right now 😳... teamorouter is offering deepseek v4 pro at almost no cost you can also use deepseek v4 flash for free what you get: - deepseek v4 pro free - deepseek v4 flash free - 1M context - openai compatible api - no subscription - credits never expire why this is worth checking: > v4 pro is currently listed at $0 input and output > flash is also available at $0 > works with claude code, codex and other tools > one api gives you access to multiple models getting started: 1. go to 2. create your account 3. open the dashboard 4. create your api key base url: 5. add it to your ai coding tool 6. select deepseek-v4-pro-free worth testing while it’s availableshow more

K2S
38,998 次观看 • 1 个月前
Codex can run Qwen-3.8-max now as well!! Alibaba most... capable model, dropped today and it's already in my codex picker. It's a token plan subscription, not metered api billing. You take the key from your Qwen plan, drop it into Codex Router, and it spends down the plan instead of your card. There's a catch though. Qwen's official setup switches your whole codex over to them, so your ChatGPT models stop showing up at all. That's exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, Grok, kimi, Deepseek and now Qwen 3.8 max all sit in the same picker and it can grab whichever one suits the job. Router's open source, setup's in the video 👇show more

Ziwen
419,031 次观看 • 1 个月前
We are entering an extremely exciting era for open-weight... models. Kimi K2.6 now feels like a top agentic model. I took it for a spin via Fireworks AI fast inference APIs. Kimi K2.6 has impressive agentic capabilities, design skills, and the ability to synthesize large amounts of information. I built a little Skill that produces survey papers on any AI research topic you want. (see example in the clip) You can use the skill to tell your agent to generate a survey on whatever topic and watch it go to work. The artifact was fully generated by Kimi.ai's Kimi K2.6. It's cheap and fast. Next step for me is to explore ways to continue integrating the capabilities of these models on use cases like automating my LLM knowledge bases and augmenting my agent memory capabilities. Stay tuned for more.show more

elvis
47,678 次观看 • 5 个月前
Muse Spark 1.3 is now in Codex!! OpenCode's free... tier runs it for nothing, and $10 on OpenCode Go gets you over 45,300 requests. It's basically Opus 5 tier too. There's a catch though. Meta may use your inputs and outputs for training, which is the whole reason the tier is called contributor and the whole reason it's free. Muse 1.3 just sits in the picker next to my paid models now. I'll just use Astra + Muse Spark for subagents. Gonna keep everything in stock in case we need it.show more

Ziwen
92,093 次观看 • 23 天前
OpenCode Go is now wired into Codex!! The pricing... is insane. $10 gets you 10,000 DeepSeek requests every 5 hours (no weekly limits). I converted that into DeepSeek API dollars because I thought I was reading it wrong, and the same 5 hours of usage would run somewhere between $10 and $30 depending on how big your context gets. So one afternoon of use already covers the whole sub. It comes with Kimi K3 as well. Both sit in the picker next to my other models now. Next time we hit a limit in the middle of a loop, we can grab it and keep going.show more

Ziwen
434,241 次观看 • 1 个月前
I explored a further possibility with local models: Qwen3.6... 35B A3B + NVIDIA LocateAnything-3B as a local Computer Use agent (proof of concept). In the demo, I asked it to switch my Mac to light mode. It did. Then back to dark. Did that too — finding the right toggle in System Settings, clicking it, and verifying the change itself. It's fully screenshot-based, so no Accessibility API needed. If it's on screen, the agent can see it and act on it. This runs entirely on your own hardware — private, local, built from two small open models.show more

stevibe
44,248 次观看 • 3 个月前
✨ 3d models are now LIVE on Photo AI... 😊 You can now turn any AI photo you make into a 3d model by pressing [ 📦 Make 3d model ] And then you can view it inside Photo AI or download it as a .GLB 3d model file It's still very early in AI generated 3d model world but it's nice to have this feature working already As always, the models will keep improving, so this feature will keep getting better (like it did with video, it sucked before, now it's getting passable) Next would be nice to switch to .USDZ so you can load it straight into your iPhone with ARKit and put it in your room Available now for everyone on the Premium and Ultra planshow more

@levelsio
112,930 次观看 • 1 年前
Union Alpha is now in Codex!! The speed is... insane over 300 to 400 tokens a second, 262k context, images in, and it's free for a week. It's a good timing too, most of the Codex usage is basically gone right now. Zai did this exact thing in August: Ox Alpha showed up unnamed, free for a week, built for agentic coding, and a week later it was GLM-5.3-Flash. So what model do you guys think it is?show more

Ziwen
98,712 次观看 • 9 天前
This free model just beat every closed-source AI on... coding benchmarks. Open weights. 6x cheaper than Opus. The labs don't want you to know it exists. > GLM-5.2 from Zai just topped Code Arena - the first open-weights model to ever hold the #1 coding spot. Not a leaked weight, not a fine-tune. A fully open model beating GPT and Claude on their own turf. It's live on Hugging Face Inference API right now. 5 providers: Novita, Together AI, Fireworks, Deepinfra, Zai. OpenAI-compatible client. → Go to huggingface(.)co → grab HF_TOKEN from account settings → plug into any OpenAI-compatible client 6x cheaper than Opus. Companies bleeding on AI bills are already routing to this for orchestration, caching, and token optimization. The smart money moved before the headline dropped. > Now you know. huggingface(.)co/zai-org/GLM-5.2 Bookmark this before everyone else figures it out.show more

Atenov int.
12,395 次观看 • 3 个月前
Nvidia just put a $250,000 cloud workload on your... desk for $2,999 - and killed your $1,900/month AWS bill in the process You don't rent it, you don't manage it, you don't pay a single cloud bill - you just plug it in and let it eat the workloads you used to wire to AWS every month It looks like a small Mac mini, it's actually a full GB10 Grace Blackwell stack with 128GB of unified memory running models up to 200B parameters It's called DGX Spark, the consumer version of the rack Nvidia ships to OpenAI The reason Nvidia did this is simple Cloud GPU pricing is a tax on every developer building AI right now $1,900/month per seat, billions in margin flowing to AWS, Lambda, and CoreWeave Nvidia just cut themselves in by removing the cloud entirely Their solution is to skip the middleman, ship the rack to your desk, and let you keep every dollar of margin you used to wire to a hyperscaler This is much cheaper, faster, and you own the asset at the end But there is still a question nobody is answering yet, what happens to AWS, GCP, and Lambda when 500,000 developers move their inference back to a $2,999 box on their desk Also, technically you can stack four of these and run a 1.6 trillion parameter model locally for under $12,000 Even a single Spark out-performs the cloud subscription Anthropic engineers were running two years ago bookmark this, it pays back in 60 days 👇show more

ZEUS⚡️
85,803 次观看 • 4 个月前
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 次观看 • 1 个月前
made a prototype to see how it feels, and... i kinda like it not promising this is gonna be out anytime soon, but curious if it's something you'd be interested in? if so, it'd need to go through design and engineering so they can think about it holistically. some things look simpler on the surface than they actually are!show more

Pedro Duarte
34,966 次观看 • 4 个月前