Multi-Model Review, Harnesses, AI Video — AI Daily Sep 13
351 messages · 61 active members
Overview
Topics
@jasonakatiff formalized a three-model audit pipeline: Fable orchestrates, Opus codes, then Fable x-high + Grok + Astra x-high all review before deploy. @samb69 called it the 'holy trinity.' Rationale: different models catch different classes of bugs and single-model review misses too much.
Builders paired Codex, Opus, and Grok — cheaper models as grunt workers, stronger models as orchestrators. Consensus formed that the harness (Nvidia's, OMP) drives outcomes more than the base model, with one member reporting a jump from 30% to 100% task success. Thewildzeno's system prompt cycler surfaced tradeoffs between brute speed and orchestrate depth, while Kieran explored wave-based parallelism and file-write conflicts.
Sentiment shifted hard against Opus with @jasonakatiff finding 'a ton of issues' in Opus-produced code and @samarh90 doing their first rollback because of it. Kieran declared it time to scale down Claude and scale up GPT, citing Fable's constant refusals. @jcartu confirmed Astra has not been nerfed based on shipped work.
Members compared Omni, Veo, Seedance, Hyperframes and Google Flow for AI video work. Kieran shared that 8-10s clips remain the cost-efficient sweet spot with AI editing workflows stitching scenes together. Higgsfield came up as a top-tier tool for creative automation, with debate on whether it truly beats Replicate beyond UI polish.
Members joked about M4 Pro anonymous MacBooks sounding like microwaves under heavy coding loads, with cooling and longevity concerns. kie.ai's 72-90% off GPT-6 Astra API pricing raised eyebrows — @thewildzeno is building an open-source validator to verify these routes actually serve the real model. @samb69 argued he could drop 70% of his infra and move faster since models are now smart enough that over-engineering is the bottleneck.
Words worth knowing
Three terms from the glossary.
Key Takeaways
- Multi-model review (Fable + Grok + Astra) is becoming standard — single-model audits miss too much before deploy.
- The harness matters more than the model — Nvidia's harness reportedly took Opus from 30% to 100% task success.
- Opus is producing enough bad code that experienced builders are rolling back and questioning it as a primary coder.
- Split roles: use cheaper models (Grok, Sol) as grunt coders and stronger models (Opus) as orchestrators; match system prompt mode to task type.
- Discounted API resellers like kie.ai are tempting at 72-90% off but likely run on farmed accounts — treat credentials with caution.
Hot Threads
Opus producing bad code and the three-model review stack
System prompt cycler testing: brute vs normal vs orchestrate modes
'I'm the biggest bottleneck' — dropping infra and letting models cook