Model Tiering, GPT-Live Orchestration, pi-vcc — AI Daily Sep 14

621 messages · 79 active members

621
messages
79
active members
@anonymous, @thewildzeno, @Kieran
top contributors

Overview

September 14 was dominated by harness economics and orchestration architecture. @thewildzeno laid out a CEO/CTO/labourer model tiering framework — frontier models (Astra, Fable) for planning, mid-tier (Gemini Flash, Grok) for orchestration, and cheap fast models (Luna, Muse 1.3 Spark, Qwen) for actual code — while @Anonymoushat's 6B-tokens-in-6-days report sparked debate on cache-inflated counters, aggressive auto-compact limits, and Luna delegation. pi-vcc emerged as the consensus replacement for OMP's default snapcompact, and @thewildzeno shipped a GitHub repo of OMP extensions (throughput visualizer, session-mode cycler, brute/orchestrate templates). Voice-first orchestration was the other big thread: @leewardbound detailed Codex desktop + mobile with GPT-Live driving dozens of Orca remote sessions via OMP, with @julhi123 surfacing getpatter.com as an MIT-licensed alternative to Vapi/Retell. On pricing, @expadz and @leewardbound argued Claude API is unworkable for SaaS ($30–60 per video edit), with Luna at $0.20/M input tokens now replacing 4o and Sonnet 3.5 for grunt work. @anonymous is running 6 concurrent CC agents with sessions up to 97 days, reporting emergent workflows. On the creative and ops side, @thewildzeno shared Gemini-scored TikTok hook rules (early payoff, movement, immediate VO, location cues) and vented on creators demanding recurring usage fees regardless of performance. @bartekadamczyk wired a Cloudflare webhook around Whop's no-custom-pixel limit across ~100 conversions, another builder is assembling a Samsung S10 PCB phone farm for agent-driven TikTok, and Kieran/@Guto_gouvea are migrating Convertri landers to raw HTML on Cloudflare.

Topics

@thewildzeno's CEO/CTO/labourer framework — frontier for planning, mid-tier for orchestration, cheap/fast for code — resonated with builders burning 1B+ tokens/day. @Anonymoushat's 6B-in-6-days report prompted discussion of cache-inflated counters, tighter auto-compact limits (~200k), idle compaction, and delegating grunt work to Luna, Muse 1.3 Spark, or Qwen.

@leewardbound runs Codex desktop + mobile with GPT-Live orchestrating dozens of Orca remote sessions through OMP, with a VPS bridging Orca and desktop. @thewildzeno is building a Windows/VPS variant, and @julhi123 shared getpatter.com, an MIT-licensed SDK wiring GPT-Live, Gemini Live, Ultravox and ElevenLabs ConvAI directly as a Vapi/Retell alternative.

Builders adopted pi-vcc as a superior replacement for OMP's default snapcompact, overriding /compact and remote compact calls with better token savings, info retention, and speed. @thewildzeno also published a GitHub repo with a throughput visualizer (TPS for main and subagents), an Alt+O session-mode cycler, and brute/orchestrate system.md templates — warning that cycling modes purges caching.

@expadz and @leewardbound argued Anthropic API pricing is prohibitive for SaaS, with one video edit costing $30–60 in tokens. @leewardbound migrated bots, summarizers and analytics to Luna at $0.20/M input tokens, arguing it now beats 4o/Sonnet 3.5 when paired with a stronger planner. @thewildzeno flagged Grok 4.6 behaves differently via direct API vs OpenRouter, raising routing questions.

@thewildzeno shared Gemini-scored hook rules from Thai influencer videos: early payoff, movement over stills, immediate VO framing, and visible location cues. @julhi123 is building a rule-checker that fails builds when rules lack tests. Builders also vented on creators demanding recurring 30–60 day royalties or per-video usage fees decoupled from actual conversion performance.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Right-size models per task: frontier (Astra, Fable) for planning, cheap/fast (Luna, Qwen, Muse 1.3 Spark) for code — sending expensive models to write code is wasteful.
  • Luna at $0.20/M input tokens now replaces 4o and Sonnet 3.5 for API grunt work; Claude API can hit $30–60 per video edit and is pushing SaaS builders off Anthropic.
  • pi-vcc beats OMP's default snapcompact on token savings, info retention, and speed — set it as your /compact interceptor and pair with ~200k auto-compact limits.
  • GPT-Live at $0.05/min plus an oauth sub is enabling hands-free phone-driven orchestration of dozens of Codex/Orca remote agents.
  • Video hook performance correlates with early payoff, movement over stills, immediate VO framing, and visible location cues — testable rules for automated scoring.

Hot Threads

@thewildzenostarted

Model tiering: CEO/CTO/labourer analogy for harness design

25 replies6 participants
@Anonymoushatstarted

Burning 1B tokens/day — how to optimize multi-agent spend

22 replies5 participants
@leewardboundstarted

Bridging GPT-Live into Codex + Orca remote orchestration

20 replies3 participants

Linked Items