Grok 4.5 Workhorse, Hermes Stack, Opus 5 Prompting — AI Daily Aug 1

392 messages · 63 active members

392
messages
63
active members
@jcartu, @tounano, @Kieran
top contributors

Overview

Model wars dominated the day, with Grok 4.5 fast emerging as the token-efficient workhorse favorite while GPT-5.6 Luna Max impressed on quality but disappointed on speed (80 min vs Grok's 25 min on the same pipeline). Fable on low reasoning became the go-to orchestrator, and @thewildzeno shared TPS benchmarking methodology using pi-tps-meter. Meanwhile K3 on Fireworks burned through $50 in two hours but delivered wild results. @jcartu dropped a masterclass on building a full Hermes + OMP stack: ragflow for research (Brave, Tavily, Sonar), browserbase over Playwright for residential proxies, Gemini watching YouTube for summaries, and hindsight for shared memory banks across sessions. The consensus: after a few months of use, this beats paying for Manus. Separately, @iggot shared a workflow feeding Anthropic's official Opus 5 prompting docs to Claude to rewrite CLAUDE.md files, stripping legacy behavioral instructions for faster, shorter, higher-QA responses. On the distribution side, Snap officially moved to block AI content, raising questions about whether Meta or TikTok will follow, while @bartekadamczyk flagged unusually high CPMs on fully AI-generated Meta videos vs CapCut-edited AI assets. Builders also dissected OverSkill's new $10k vibe-coding contest (concerning ToS, likely cheap-model backend), Nutlope's Hallmark anti-slop design skill, a UI/UX Pro Max skill repo, and AirLLM claiming 70B inference on a single 4GB GPU.

Topics

Grok 4.5 fast won praise as the most token-efficient workhorse with strong initial TPS burst, while GPT-5.6 Luna Max delivered better quality but took 80 minutes on a pipeline Grok completed in 25. Fable on low reasoning became the preferred orchestrator, with users layering Opus 5 as advisor. @thewildzeno modded OMP to force fast mode on specific agents like commit/small.

@jcartu shared a detailed setup: Gemini 3.6 Flash on max as Hermes orchestrator, Sol/Kimi K3 for coding in OMP CLI, ragflow with Brave/Tavily/Sonar for research, browserbase over Playwright for residential proxies, and hindsight for cross-session memory shared between Hermes and OMP. The payoff: a personalized system that gets smarter daily and replaces Manus.

@iggot shared a workflow where he fed Anthropic's official Opus 5 prompting docs to Claude and instructed it to rewrite CLAUDE.md files, stripping legacy behavioral instructions from older models. Result: faster responses, shorter replies, and better QA. He also switched to Wispr Flow for voice input and longer, batched prompts. Frustration with Anthropic refusals also drove users toward Codex and K3 for tasks Claude wouldn't touch.

@bartekadamczyk observed unusually high CPMs on fully AI-generated videos (built with Claude Code + ffmpeg) vs CapCut-edited AI assets, with @Anonymoushat suspecting a Meta detector. @startropics reported hired actors still outperform pure AI on the same script, though 5-second AI b-roll hooks perform reliably. Meanwhile Snap announced it's blocking AI-generated content from distribution, sparking debate on whether Meta or TikTok will follow.

OverSkill launched a $10k vibe-coding contest requiring $39 in credits, with community flagging ToS letting sponsors rip participants' code and a likely cheap-model backend. @ferchonaso shared a UI/UX Pro Max skill repo plus Nutlope's Hallmark anti-slop design skill for Claude Code, Cursor, and Codex. @Vaibhavcste surfaced AirLLM claiming 70B inference on a single 4GB GPU, and a Fraser tweet about a $21k MRR AI video app sparked saturation discussion.

Key Takeaways

  • Grok 4.5 fast is the most token-efficient model right now — great as a coding driver but bad for orchestration; needs more structured prompts.
  • Full Hermes stack recipe: Gemini 3.6 Flash orchestrator + ragflow (Brave/Tavily/Sonar) + browserbase for residential proxies + hindsight for shared memory across OMP and Hermes sessions.
  • Rewrite your CLAUDE.md for Opus 5 using Anthropic's official prompting docs — strip old-model behavioral hacks for faster, shorter outputs, and batch longer prompts via voice input.
  • OMP is open-source and moddable — @thewildzeno added fast-mode forcing per subagent, and turning off Todo reminders in /settings prevents the harness from auto-continuing scope.
  • Fully AI-generated Meta video ads are seeing abnormally high CPMs vs CapCut-edited AI assets; hired actors still beat pure AI on the same script, but 5-second AI hooks perform reliably.

Hot Threads

@jcartustarted

Full Hermes + OMP + hindsight orchestration setup walkthrough

14 replies3 participants
@tounanostarted

Grok 4.5 vs Luna Max vs Fable — speed and quality comparison

22 replies4 participants
@bartekadamczykstarted

Fully AI-made Meta video ads driving unusually high CPMs

12 replies5 participants

Linked Items