AI, actually tested · weekly

I test AI tools so you don't have to read the docs.

Recent updates

updated 4 days agoread all →
Jul 24 · Models · the big one

Anthropic ships Claude Opus 5, its new frontier model

Near-Fable-5 frontier intelligence at half the price ($5/$25 per Mtok) and state of the art on coding evals, live today across the API, Claude.ai, and Claude Code. If you build on frontier models, benchmark Opus 5 on your own tasks now that the price-per-capability has shifted.

read the source →

GitHub trending

view all →
1
stablyai/orca
An ADE to run a fleet of coding agents in isolated git worktrees.
+5.7k/wk
2
iOfficeAI/OfficeCLI
A CLI for agents to edit Office documents with no Office install.
+5.3k/wk
3
ogulcancelik/herdr
Terminal multiplexer for many AI agents, with detachable sessions.
~17.4k

Fresh releases

view all →
Jul 19 · Alibaba · High
Qwen3.8-Max (2.4T preview)
Jul 16 · Moonshot · High
Kimi K3 (2.8T open MoE)
Jul 10 · Cursor · High
Cursor 3.11 side chats + hooks
Jul 9 · Meta · High
Muse Spark 1.1 / Meta Model API

Changing AI Trends

read all →
// the big one

Agent safety and data-privacy went front-page.

Copilot's MCP trust layer, Codex's wider rm detection, Claude Code's runaway caps, and the Grok Build repo-exfiltration scare all landed in the same two weeks. Safety is now a shipping feature, not an afterthought.

02
Local inference keeps getting faster. Multi-Token Prediction speculative decoding merged into llama.cpp (1.4 to 2.2x speedups on consumer GPUs), Ollama v0.32.0 added an agentic "Chat, Code" mode, and local MoEs dominate r/LocalLLaMA.
03
Verification is the bottleneck now. Code-gen speed is no longer the constraint; human and pipeline review capacity is. "Agents don't work for us" now means "our verification pipeline can't absorb the volume."
04
Cost arbitrage graduated to core infra. Custom model routers are everywhere: Warp routers, Vercel AI Gateway, and Product Hunt winners. Cross-provider cost routing is now a funded category, not a hack.

The AI that actually mattered this week, in five minutes.

No hype. No 12-part funnels. Just the tools I've actually run.

one email a week · unsubscribe in one tap

From the channel

watch all on YouTube →
Grok 4.5 vs GPT-5.6 vs Claude: Who Won?1:08Sol Pro vs Fable 5: Don't choose WRONG0:57GPT-5.6 Just Beat Claude & Gemini (You Can't Use It)1:08WTF is AI Poisioning ?0:54You're Using Claude WRONG - 5 Free Skills1:10Anthropic's Fable 5 is Back : But with a Catch1:04The Future of AI work : Loop Engineering0:49
~/subscribe $