Grok 4.5 Workhorse, Hermes Stack, Opus 5 Prompting — AI Daily Aug 1
392 messages · 63 active members
Overview
Topics
Grok 4.5 fast won praise as the most token-efficient workhorse with strong initial TPS burst, while GPT-5.6 Luna Max delivered better quality but took 80 minutes on a pipeline Grok completed in 25. Fable on low reasoning became the preferred orchestrator, with users layering Opus 5 as advisor. @thewildzeno modded OMP to force fast mode on specific agents like commit/small.
@jcartu shared a detailed setup: Gemini 3.6 Flash on max as Hermes orchestrator, Sol/Kimi K3 for coding in OMP CLI, ragflow with Brave/Tavily/Sonar for research, browserbase over Playwright for residential proxies, and hindsight for cross-session memory shared between Hermes and OMP. The payoff: a personalized system that gets smarter daily and replaces Manus.
@iggot shared a workflow where he fed Anthropic's official Opus 5 prompting docs to Claude and instructed it to rewrite CLAUDE.md files, stripping legacy behavioral instructions from older models. Result: faster responses, shorter replies, and better QA. He also switched to Wispr Flow for voice input and longer, batched prompts. Frustration with Anthropic refusals also drove users toward Codex and K3 for tasks Claude wouldn't touch.
@bartekadamczyk observed unusually high CPMs on fully AI-generated videos (built with Claude Code + ffmpeg) vs CapCut-edited AI assets, with @Anonymoushat suspecting a Meta detector. @startropics reported hired actors still outperform pure AI on the same script, though 5-second AI b-roll hooks perform reliably. Meanwhile Snap announced it's blocking AI-generated content from distribution, sparking debate on whether Meta or TikTok will follow.
OverSkill launched a $10k vibe-coding contest requiring $39 in credits, with community flagging ToS letting sponsors rip participants' code and a likely cheap-model backend. @ferchonaso shared a UI/UX Pro Max skill repo plus Nutlope's Hallmark anti-slop design skill for Claude Code, Cursor, and Codex. @Vaibhavcste surfaced AirLLM claiming 70B inference on a single 4GB GPU, and a Fraser tweet about a $21k MRR AI video app sparked saturation discussion.
Key Takeaways
- Grok 4.5 fast is the most token-efficient model right now — great as a coding driver but bad for orchestration; needs more structured prompts.
- Full Hermes stack recipe: Gemini 3.6 Flash orchestrator + ragflow (Brave/Tavily/Sonar) + browserbase for residential proxies + hindsight for shared memory across OMP and Hermes sessions.
- Rewrite your CLAUDE.md for Opus 5 using Anthropic's official prompting docs — strip old-model behavioral hacks for faster, shorter outputs, and batch longer prompts via voice input.
- OMP is open-source and moddable — @thewildzeno added fast-mode forcing per subagent, and turning off Todo reminders in /settings prevents the harness from auto-continuing scope.
- Fully AI-generated Meta video ads are seeing abnormally high CPMs vs CapCut-edited AI assets; hired actors still beat pure AI on the same script, but 5-second AI hooks perform reliably.
Hot Threads
Full Hermes + OMP + hindsight orchestration setup walkthrough
Grok 4.5 vs Luna Max vs Fable — speed and quality comparison
Fully AI-made Meta video ads driving unusually high CPMs