Smart Orchestrator Stacks, Jev Reality Check, Gemini 3.8 Flash — AI Daily Sep 20

402 messages · 59 active members

402
messages
59
active members
@anonymous, @jasonakatiff, @jcartu
top contributors

Overview

Sunday's chat was dominated by orchestration philosophy: @jcartu, @tounano, and @thewildzeno pushed the 'smart orchestrator, dumb workers' pattern hard, arguing that Gemini 3.8 Flash (or GLM Flash) as the main orch — with Fable/Astra reserved for ADR planning and critical gates — dramatically improves speed and usage economics. @seekersight's stack (Astra Medium orch, Fable+Astra High plan, Luna xHigh build, Grok review, Haiku scout) became a reference config, and the key trick is telling the planner exactly which dumb model will execute so it builds rubrics and 'kill the boss' checkpoints accordingly. Builders stress-tested Jev on real workflows — funnel lead analysis, transcript risk, campaign name interpretation — and largely concluded it's a classifier that needs heavy pre-mapping, examples, and the official skill installed. Kieran called it 'more hype than anything' for general use, while @jasonakatiff is finding narrower wins in transcript risk scoring, E2E testing, and B2B lead scoring. Meanwhile @scalingfrog wired Grok into ad accounts to kill losers and iterate winners, reporting first sales and sparking debate on micromanaging spend vs. trusting Meta CBO. On infrastructure, @watchmedropship hit ~300 tok/s on Qwen 3.8 Next Flash via tensor parallelism (with a clean explainer from @samb69), and PrismML's Ternary Bonsai 2 27B (5.9GB, 98.2% of Qwen3.8 27B) plus Qwen3.8-LiveTranslate dropped. Usage anxiety continued as Anthropic paused new $200 accounts with a banked reset promised Tuesday, and @jasonakatiff shared a hard-won lesson on defining strict platform vocabularies — ambiguous terms like 'workflow' silently degrade agent output.

Topics

@jcartu, @tounano, and @thewildzeno converged on using a fast/cheap model (Gemini 3.8 Flash, GLM Flash) as the main orchestrator while reserving Fable/Astra for ADR planning and critical gates. The key discipline: tell the planner exactly which dumb model will execute so it builds rubrics, gates, and must-pass checkpoints into the plan. @seekersight and Kieran shared full role configs mapping models to orchestrate/plan/build/review/scout.

Multiple builders tested Jev on funnel analysis, transcripts, and campaign interpretation. Consensus: it's a classifier that needs the official skill installed and heavy pre-mapping, and often loses to tuned local models like Llama or GLM 5.3 Flash. Real wins are narrow — transcript risk scoring, E2E testing, and B2B lead scoring — not general workflow intelligence.

@samb69 clocked Gemini 3.8 Flash at 350 tok/s and @jcartu called its vision tower 'sick,' making it the current favorite for main orchestration. Detractors like @seekersight found it wasn'anonymous passing full context to Fable during planning handoffs. Gemini 4 is reportedly being tested for release this week.

@scalingfrog wired Grok to ad accounts to kill underperformers and iterate winners, reporting first sales and sparking a CBO-vs-micromanage debate. @jcartu shared the Grok API playbook: $1k spend + 30 days to reach Tier 3, where it becomes usable and excels at browser automation. Anthropic paused new $200 accounts but old ones can reactivate, with a banked reset expected Tuesday.

@watchmedropship hit ~300 tok/s on Qwen 3.8 Next Flash via tensor parallelism, with @samb69 explaining the tradeoff vs. pipeline parallel. Qwen3.8-LiveTranslate (2.3s lag, 60 languages) and PrismML's Ternary Bonsai 2 27B (5.9GB, 98.2% of Qwen3.8 27B) dropped. Separately, @jasonakatiff warned that overloaded terms like 'workflow' silently degrade agent output — treat naming as first-class architecture.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Use Gemini 3.8 Flash or GLM Flash as your main orchestrator and reserve Fable/Astra for ADR planning and critical gates — speed and cost improve dramatically without sacrificing plan quality.
  • When planning with a frontier model, explicitly tell it which dumb model will execute the work so it builds rubrics, gates, and must-pass checkpoints into the plan.
  • Jev is a classifier, not an intelligence layer — install the official skill, pre-map examples, and scope to transcript risk, E2E testing, or lead scoring.
  • Tensor parallelism unlocks ~300 tok/s on Qwen 3.8 Next Flash, and PrismML's Ternary Bonsai 2 27B retains 98.2% of Qwen3.8 27B performance at just 5.9GB.
  • Define a strict platform vocabulary early; overloaded words like 'workflow' will silently degrade agent output as your codebase grows.

Hot Threads

@jcartustarted

Smart orchestrator with dumb workers — why Fable/Astra for orch is backwards

25 replies8 participants
@Kieranstarted

Testing Jev on overnight data analysis for lead reverse-engineering

12 replies5 participants
@scalingfrogstarted

Grok bot managing ad accounts — kill losers, iterate winners

6 replies4 participants

Linked Items