Multi-Model Review, Harnesses, AI Video — AI Daily Sep 13

351 messages · 61 active members

351
messages
61
active members
@anonymous, @Kieran, @jasonakatiff
top contributors

Overview

September 13 was dominated by how builders are stacking, reviewing, and orchestrating coding models. Sentiment shifted hard against Opus — @jasonakatiff found 'a ton of issues' auditing Opus code, @samarh90 did their first rollback because of it, and Kieran flipped from calling Fable and Opus 'goated' to declaring it time to scale down Claude and scale up GPT after Fable's constant refusals blocked work. @jcartu confirmed Astra hasn'anonymous been nerfed based on shipped work. The multi-model review pattern is hardening into standard practice: @jasonakatiff formalized a pipeline of Fable orchestrating, Opus coding, then Fable x-high + Grok + Astra x-high reviewing before deploy — what @samb69 called the 'holy trinity.' Parallel to that, builders emphasized the harness matters more than the model itself, with one report of Nvidia's harness taking Opus from 30% to 100% task success. Thewildzeno tested a system prompt cycler swapping brute, normal, and orchestrate modes, while Kieran discussed wave-based agent orchestration and file-write conflicts. Side threads covered AI video tooling (Omni, Veo, Seedance, Hyperframes, Higgsfield), MacBooks melting under heavy coding loads, SF Tech Week meetups, and kie.ai's 72-90% off GPT-6 Astra API pricing raising suspicions of farmed accounts. @samb69 captured the mood: 'I'm the biggest bottleneck in all my work' — models are smart enough that over-engineering is now the drag.

Topics

@jasonakatiff formalized a three-model audit pipeline: Fable orchestrates, Opus codes, then Fable x-high + Grok + Astra x-high all review before deploy. @samb69 called it the 'holy trinity.' Rationale: different models catch different classes of bugs and single-model review misses too much.

Builders paired Codex, Opus, and Grok — cheaper models as grunt workers, stronger models as orchestrators. Consensus formed that the harness (Nvidia's, OMP) drives outcomes more than the base model, with one member reporting a jump from 30% to 100% task success. Thewildzeno's system prompt cycler surfaced tradeoffs between brute speed and orchestrate depth, while Kieran explored wave-based parallelism and file-write conflicts.

Sentiment shifted hard against Opus with @jasonakatiff finding 'a ton of issues' in Opus-produced code and @samarh90 doing their first rollback because of it. Kieran declared it time to scale down Claude and scale up GPT, citing Fable's constant refusals. @jcartu confirmed Astra has not been nerfed based on shipped work.

Members compared Omni, Veo, Seedance, Hyperframes and Google Flow for AI video work. Kieran shared that 8-10s clips remain the cost-efficient sweet spot with AI editing workflows stitching scenes together. Higgsfield came up as a top-tier tool for creative automation, with debate on whether it truly beats Replicate beyond UI polish.

Members joked about M4 Pro anonymous MacBooks sounding like microwaves under heavy coding loads, with cooling and longevity concerns. kie.ai's 72-90% off GPT-6 Astra API pricing raised eyebrows — @thewildzeno is building an open-source validator to verify these routes actually serve the real model. @samb69 argued he could drop 70% of his infra and move faster since models are now smart enough that over-engineering is the bottleneck.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Multi-model review (Fable + Grok + Astra) is becoming standard — single-model audits miss too much before deploy.
  • The harness matters more than the model — Nvidia's harness reportedly took Opus from 30% to 100% task success.
  • Opus is producing enough bad code that experienced builders are rolling back and questioning it as a primary coder.
  • Split roles: use cheaper models (Grok, Sol) as grunt coders and stronger models (Opus) as orchestrators; match system prompt mode to task type.
  • Discounted API resellers like kie.ai are tempting at 72-90% off but likely run on farmed accounts — treat credentials with caution.

Hot Threads

@jasonakatiffstarted

Opus producing bad code and the three-model review stack

14 replies6 participants
@thewildzenostarted

System prompt cycler testing: brute vs normal vs orchestrate modes

8 replies5 participants
@samb69started

'I'm the biggest bottleneck' — dropping infra and letting models cook

9 replies5 participants

Linked Items