Claude Code Token Bloat, Hermes Orchestration, GLM 4.6 — AI Daily Jul 30
464 messages · 73 active members
Overview
Topics
Kieran shared research showing Claude Code sends ~33k tokens of harness overhead before the user prompt on the same model where OpenCode sends only ~7k. Builders speculated frontier labs may be incentivized to burn tokens via their own harnesses, strengthening the case for OpenCode or OMP.
Members shared elaborate orchestration stacks running Hermes, Fable, OpenCode, and Buzz together over Telegram, with agents negotiating tool choice and concurrency between themselves. Discussion centered on 2-3 layers of delegation, with @sibunting running 2 Opus dev/QA pairs plus Hermes and OC subagents under Fable as main orchestrator. @rmktg's Hermes-based media buyer is nearing $300k in managed spend.
Several builders migrated orchestrators to GLM 4.6 for cost and looser refusals compared to Claude, especially for marketing workflows. MiamiMega reported only a few dollars per week for non-stop cron use, and Ariel confirmed it's powering her end-to-end organic content pipeline with Ideogram 4 image generation.
@expadz shared workflow details for running ~300-1000 Veo clip generations per week through useapi with fresh Gmail accounts and no recaptcha issues, noting a ~1000/day soft limit and useapi's support for 50 rotating accounts. @Andrenator8 flagged a reseller offering 25K credits for $87 vs $250 retail, pending testing.
@bartekadamczyk got blocked using a personal profile token; Kieran recommended system user tokens via a dev app in testing mode. @amster93 detailed that server-to-server app review took only 2 hours and grants higher rate limits, but requires a verified BM with roughly $1k in spend to share the app with partner BMs.
Key Takeaways
- Claude Code sends ~33k tokens of harness overhead vs ~7k for OpenCode on the same model — switching harnesses can meaningfully cut token spend.
- Multi-agent Telegram orchestration is maturing — agents negotiate tool selection, concurrency, and model choice among themselves before returning plans for human sign-off.
- GLM 4.6 is emerging as a cost-effective orchestrator alternative to Claude at a few dollars per week, especially for marketing tasks Claude refuses.
- useapi supports ~1000 Veo gens/day per Google account with auto-rotation across up to 50 accounts; use fresh Gmails to isolate risk from main ad accounts.
- For Meta API, use a system user token from a dev app in testing mode; server-to-server review is optional but only takes ~2 hours if you need higher rate limits.
Hot Threads
Multi-layer agent orchestration with Fable, Hermes, OC
Meta API token blocking and app review process
useapi Veo/Flow account rotation and rate limits