Fable 5 Nerf, Kimi K3 Sold Out, Hermes v0.19 — AI Daily Jul 20
693 messages · 87 active members
Overview
Topics
Widespread complaints that Fable 5 is silently auto-switching to Opus 4.8, flagging benign prompts as cybersecurity violations (insurance forms, analytics dashboards), and burning usage 2–3x faster than last week. Anthropic also emailed users about new weekly caps up to 50%. Workarounds include /config to disable auto-switch and Cowork's double limits, but most power users are consolidating on GLM 5.2, Kimi K3, GPT-5.6, and Grok 4.5 — a Codex $200 + Grok $99 + Kimi $200 stack delivers ~4–5x more usage than multi-seat Fable at $800/mo.
Kimi K3 is the community's new favorite for agentic coding — @arielletolome one-shotted a corporate website, advertorials, and app fixes. Plans sold out for most users; legacy Kimi accounts and US VPN are the only reliable access paths. Signup hack: start on $99, upgrade to $199 after maxing to get a refund and full quota reset. The $199 tier delivers 30x CLI / 10x desktop usage, but K3 is token-hungry, so keep K2.7 Highspeed as a speed layer to preserve K3 quota.
Hermes v0.19.0 shipped with ~80% faster startup, GPT-5.6 support, better background agent observability, and default-on reasoning streams. @jarvisballer made the case that Hermes' real value is as an ops layer — memory, cron, Kanban queues, skills, and multi-profile routing — not Slack integration. @samb69 shared a ~20-profile setup with an orchestrator router, Hindsight memory, and per-task model selection (Grok 4.5 for speed, GPT Terra Medium for reasoning). Users with modified files should stage updates rather than run `hermes update` in place.
@jcartu detailed a 4x RTX 6000 Pro Blackwell (96GB each) rig running hybrid NF3/NVFP4 GLM 5.2 at 90–110 tps, 3500 tps prefill, ~500k context, drawing 3kW under load at ~60–70k EUR build cost. @momodxb pushed back: hybrid cloud + Zai subscription beats local ROI unless running heavy concurrent workloads, especially with RAM/NVMe prices ~2.5x higher YoY. OMP harness also gained traction with OAuth to Cursor, Venice, and multi-model routing — useful for accessing Grok 4.5 in restricted regions.
@justingacina and @tounano shared a workflow: generate 20–30 top-fold mockups with minimal prompting, convert winners to code (Daisy UI recommended), then LLM-code the bottom fold. Same for logos — broad prompts + LLM-as-judge against criteria like 'looks great on a Chrome tab.' Save working patterns as skills/SOPs in markdown, not session memory. On skills vs MCP: @thewildzeno and @filiuser explained MCP tool descriptions load into every session, while API-backed skills only trigger when needed — prefer skills for context efficiency.
Key Takeaways
- Fable 5 is silently auto-downgrading to Opus and burning usage 2–3x faster — disable via /config or migrate to Kimi K3, GPT-5.6, or GLM 5.2.
- Kimi K3 signup hack: start on $99, upgrade to $199 after maxing, refund for a full quota reset from 0%.
- Stack Codex $200 + Grok $99 + Kimi $200 for ~4–5x more usage than multi-seat Fable at similar spend.
- MCP tool descriptions load every session; API-driven skills only trigger when needed — use skills for integrations to save context.
- For creative AI output, prompt broadly and generate volume — detailed prompts collapse variance; save winning patterns as skills/SOPs, not memory.
Hot Threads
Fable 5 degradation, auto-downgrades to Opus, and weekly caps
What's the real Hermes use case beyond Slack/Kanban?
Local GLM 5.2 rig specs, quantization tradeoffs, and ROI