Model Tiering, GPT-Live Orchestration, pi-vcc — AI Daily Sep 14
621 messages · 79 active members
Overview
Topics
@thewildzeno's CEO/CTO/labourer framework — frontier for planning, mid-tier for orchestration, cheap/fast for code — resonated with builders burning 1B+ tokens/day. @Anonymoushat's 6B-in-6-days report prompted discussion of cache-inflated counters, tighter auto-compact limits (~200k), idle compaction, and delegating grunt work to Luna, Muse 1.3 Spark, or Qwen.
@leewardbound runs Codex desktop + mobile with GPT-Live orchestrating dozens of Orca remote sessions through OMP, with a VPS bridging Orca and desktop. @thewildzeno is building a Windows/VPS variant, and @julhi123 shared getpatter.com, an MIT-licensed SDK wiring GPT-Live, Gemini Live, Ultravox and ElevenLabs ConvAI directly as a Vapi/Retell alternative.
Builders adopted pi-vcc as a superior replacement for OMP's default snapcompact, overriding /compact and remote compact calls with better token savings, info retention, and speed. @thewildzeno also published a GitHub repo with a throughput visualizer (TPS for main and subagents), an Alt+O session-mode cycler, and brute/orchestrate system.md templates — warning that cycling modes purges caching.
@expadz and @leewardbound argued Anthropic API pricing is prohibitive for SaaS, with one video edit costing $30–60 in tokens. @leewardbound migrated bots, summarizers and analytics to Luna at $0.20/M input tokens, arguing it now beats 4o/Sonnet 3.5 when paired with a stronger planner. @thewildzeno flagged Grok 4.6 behaves differently via direct API vs OpenRouter, raising routing questions.
@thewildzeno shared Gemini-scored hook rules from Thai influencer videos: early payoff, movement over stills, immediate VO framing, and visible location cues. @julhi123 is building a rule-checker that fails builds when rules lack tests. Builders also vented on creators demanding recurring 30–60 day royalties or per-video usage fees decoupled from actual conversion performance.
Words worth knowing
Three terms from the glossary.
Key Takeaways
- Right-size models per task: frontier (Astra, Fable) for planning, cheap/fast (Luna, Qwen, Muse 1.3 Spark) for code — sending expensive models to write code is wasteful.
- Luna at $0.20/M input tokens now replaces 4o and Sonnet 3.5 for API grunt work; Claude API can hit $30–60 per video edit and is pushing SaaS builders off Anthropic.
- pi-vcc beats OMP's default snapcompact on token savings, info retention, and speed — set it as your /compact interceptor and pair with ~200k auto-compact limits.
- GPT-Live at $0.05/min plus an oauth sub is enabling hands-free phone-driven orchestration of dozens of Codex/Orca remote agents.
- Video hook performance correlates with early payoff, movement over stills, immediate VO framing, and visible location cues — testable rules for automated scoring.
Hot Threads
Model tiering: CEO/CTO/labourer analogy for harness design
Burning 1B tokens/day — how to optimize multi-agent spend
Bridging GPT-Live into Codex + Orca remote orchestration