Claude Degradation, Parallel Agents, Kimi K3 — AI Daily Jul 22

432 messages · 76 active members

432
messages
76
active members
@Kieran, @arielletolome, @mb29266
top contributors

Overview

Anthropic continues to lose builder trust as Claude, Fable, Sonnet, and Opus ship degraded outputs, burn tokens on unrequested testing, and expand scope beyond user instructions (e.g., adding DM automation to a LinkedIn connection task). Builders are hedging with Codex 5.6 Sol High, GLM 5.2, Kimi K3, Gemini, and Grok, with several members downloading open-weight models to local SSDs as insurance. Rumors even circulated about Anthropic self-sabotaging their IPO. Field reports from GeekOut and ATO showed 1-2 person teams scaling massive ad campaigns via fully automated creative pipelines. That set up the day's deepest tactical thread: multi-agent orchestration, with Navuud running 20+ concurrent Claude sessions via git worktrees and machine gates, and c_1media detailing GPT-5.6 SOL Max supervising Fable low-reasoning inside cmux with hyper-detailed dependency-chained plans. Paseo.sh, Nebula, and Buzz surfaced as candidate orchestration layers, with cron scheduling and reviewable artifact outputs still the unsolved UX problems. Kimi K3 subscriptions briefly opened before flipping back to waitlist, with a free rate-limited endpoint appearing on zenmux. Dragodimitrov burning through $100 of Fable credits sparked a discussion on how subsidized token pricing is, with consensus that Chinese open-weight models create a competitive floor US labs can't breach. Smaller threads covered long-form AI video pipelines (Seedance, Kling, Hermes), Perplexity as agent web search backend, and Mobbin MCP plus Matt Pocock's skills repo as fixes for design and engineering context.

Topics

Fable, Sonnet, and Opus are producing D/F-tier outputs, over-testing with Ultra, contradicting themselves, and expanding scope beyond user instructions. Builders are shifting to Codex 5.6 Sol High, GLM 5.2, Kimi K3, and Gemini, with members storing open-weight models on SSD as insurance. Speculation surfaced that someone may be trying to tank Anthropic's IPO.

Navuud runs ~20 concurrent Claude sessions via isolated worktrees, serialized merges, rebase, and machine gates — just upgraded to 128GB RAM for more tmux sessions. Rocaboca recommended paseo.sh; Leewardbound and Samb69 debated Nebula.gg and Buzz.xyz as Paperclip killers. Unsolved UX: cron-scheduled tasks with reviewable artifact outputs and team collaborative chat.

C_1media detailed the current meta: GPT-5.6 SOL Max as supervisor keeping Fable low-reasoning on a tight leash in cmux, with dependency-chained plans iterated 20x to remove vagueness. Jarvisballer uses Fable as architect with Grok as fast implementer. GeekOut/ATO showed 1-2 person teams scaling ad campaigns via fully automated Claude Code pipelines. Consensus: precise instructions unlock even basic models.

Kimi subscription briefly went live before flipping back to waitlist; LionOnX shared a free rate-limited K3 endpoint via zenmux. Dragodimitrov burning through 20x Fable plus $100 credits sparked debate on how subsidized token pricing is. Thewildzeno argued US frontier labs can't raise prices because cost-efficient Chinese open-weight models create a competitive floor.

Rstmaur and Calequiram orchestrate Seedance via Hermes with a Claude skill layer for captions; consensus is reference-image-to-video plus locked voice/tonality for multi-clip continuity. Perplexity replaces Google as the web-search backend for agents on cheap $20 plans. Mobbin MCP, Refero styles, and Matt Pocock's /wayfinder skills repo emerged as fixes for design slop and real engineering context.

Key Takeaways

  • Anthropic output quality is cratering — route production through Codex 5.6 Sol High, Gemini, or Grok as fallback, and keep GLM 5.2 and Kimi K3 on local storage.
  • Multi-session parallel orchestration (20+ concurrent agents) works when you isolate worktrees, serialize merges, and gate against drift — 128GB RAM helps.
  • Best model stack right now: GPT-5.6 SOL Max supervisor directing Fable low-reasoning implementer in cmux with hyper-detailed dependency-chained plans iterated 20x.
  • State conversation purpose (brainstorm vs. build) up front — recent RL-tuned models will otherwise expand scope and burn tokens on unrequested work.
  • US frontier pricing is structurally capped by Chinese open-weight competition — expect efficiency gains rather than price hikes, and Kimi K3 API access is worth grabbing via zenmux.

Hot Threads

@c_1mediastarted

GPT-5.6 SOL Max supervising Fable low-reasoning in cmux with dependency-chained plans

12 replies4 participants
@leewardboundstarted

Nebula.gg and Buzz.xyz as potential Paperclip killers for agent orchestration

14 replies4 participants
@scalingfrogstarted

Claude eating tokens and shipping D-tier outputs — who else is seeing it?

12 replies8 participants

Linked Items