Sonnet 5.5, Deepseek Flash Coder, $1 AI VSLs — AI Daily Sep 28

620 messages · 77 active members

620
messages
77
active members
@anonymous, @expadz, @thewildzeno
top contributors

Overview

Anthropic dropped Sonnet 5.5 hours after Opus 5.5, prompting the community to re-run benchmarks and re-evaluate stacks. The consensus workflow is crystallizing: use high-reasoning models (Opus 5.5, Fable 5.1) for planning and specs, then delegate execution to fast/cheap workers like Luna 6, Deepseek 4.1 Flash, or Grok 4.6. @thewildzeno's benchmarks crowned Deepseek 4.1 Flash on low reasoning as the new grunt coder king — beating Luna, Grok 4.7, and Muse on cost, TPS, and cache hits — while @expadz showed Luna 6 fixes bugs in ~1 min vs Opus 5.5's ~14 min at ~100x lower cost. Creative workflows hit new highs: builders shipped playable 2D games, kawaii onboarding videos, and full 3-minute VSLs for ~$1 using Opus 5.5 + Codex sub-agents + Remotion/Hyperframes + ElevenLabs — no human editors needed. @expadz broke down the economics: Opus 5.5 via oauth costs ~$1 for tasks that run $15 on API, keeping the $200 sub as the cheapest path to frontier models. Meanwhile OpenAI reportedly killed the 20x plan and is reinstating the $200 tier at half usage, signaling API/subscription pricing convergence. On infrastructure, builders are fleeing local machines after WSL memory leaks and dead motherboards killed agent workflows — Hetzner, Vultr, and EC2 with Ubuntu are now the default. Multi-account Codex CLI OAuth remains painful (tokens revoke on new-machine login; isolated terminal sessions per account is the workaround), while Oh My Pi is emerging as the cleaner alternative. Kling 4.0 Flash, ElevenLabs v4, and rumors of Astra 6.1 and Gemini 4 Pro have everyone frontrunning OpenAI dev day tomorrow.

Topics

Anthropic released Sonnet 5.5 same day as Opus 5.5 momentum peaked, with early benchmarks close to Opus but mixed reports on speed after a few hours. Kling 4.0 Flash, ElevenLabs v4 (10-sec voice cloning), and rumors of Astra 6.1, Gemini 4 Pro, and OpenAI's Grok-bot competitor all landed as builders frontrun OpenAI dev day. Community read: OpenAI looks more compute-constrained than Anthropic right now.

@thewildzeno benchmarked coding models and crowned Deepseek Flash on low reasoning for best cost/performance — beating Luna (weak below medium reasoning), Grok 4.7 (2x cost, 2x time, TSC errors), and Muse with top cache hits. @expadz's parallel tests showed Luna 6 fixing bugs 10x faster than Opus 5.5 at ~100x lower cost. @Anonymoushat hit 30k-70k TPS sustained on Baseten. Emerging pattern: plan with Opus/Fable, execute with Deepseek Flash or Luna.

Builders are generating full 3-minute VSLs for ~$1 using Opus 5.5 in Claude Code with SVG/React + Remotion/Hyperframes — no editors needed. @thewildzeno built a 90-sec kawaii onboarding video first-try on ~1% of the $20 quota; @iammikevineyard shipped playable 2D games in 30 min via Opus 5.5 + Codex sub-agent + Hyperframes + ElevenLabs. @yangthegoat: 'A one-person 8-figure ecom brand might be very feasible soon.'

After motherboard deaths and WSL memory leaks killing agent runs, builders are moving to Hetzner, Vultr, and EC2. @thewildzeno recommends 4 vCPU VPS to start, forcing portable setups with git-committed secrets. @geilt runs per-dev EC2 boxes with dynamic domain routing. Ubuntu on a $20/mo VPS consistently outperforms WSL and beefy local PCs for parallel agent orchestration.

Codex CLI revokes OAuth tokens the moment you log in from another machine, making multi-account VPS setups brutal. Workaround: run isolated terminal sessions per account so tokens persist for weeks. Oh My Pi and Claude are emerging as smoother alternatives for managing 10+ OAuth sessions cleanly. New claudex-loop skill also landed for four-phase plan hardening with adversarial cross-model review.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Deepseek 4.1 Flash on low reasoning is the new grunt coder king — best cost/performance, TPS, and cache hits, beating Luna, Grok 4.7, and Muse.
  • Luna 6 fixes bugs ~10x faster than Opus 5.5 at ~100x lower cost — pair with a high-reasoning planner like Opus 5.5 or Fable 5.1.
  • Full 3-minute VSLs can be generated for ~$1 using Opus 5.5 + Remotion/Hyperframes with SVG/React — no video editors required.
  • Opus 5.5 via oauth ($200 sub) costs ~$1 for tasks that run $15 on API; OpenAI killing the 20x plan and reinstating $200 at half usage signals pricing convergence.
  • Go remote-first for agent work: WSL leaks memory under parallel agents; Ubuntu on a $20/mo Hetzner/Vultr/EC2 box handles it far better and keeps you portable.

Hot Threads

@thewildzenostarted

Deepseek Flash benchmark results and grunt coder testing

15 replies6 participants
@expadzstarted

Luna 6 benchmarks vs Opus and Deepseek Flash for bug fixing

30 replies5 participants
@Kieranstarted

Video editor agency pricing and SOP-driven creative workflows

25 replies4 participants

Linked Items