DeepSeek V4 Pro, Grok Bot, Seedance UGC — AI Daily Aug 12

651 messages · 81 active members

651
messages
81
active members
@jcartu, @arielletolome, @jonmacofficial
top contributors

Overview

Model drops dominated the day. DeepSeek V4 Pro (0813) launched with ~65 tok/sec throughput, near-free cached tokens, and aggressive pricing — but hands-on testing quickly soured, with @jcartu calling it 'hot garbage' after it choked on a Rubik's cube build, and the community flagging classic benchmaxxing symptoms plus missing vision support. GLM, Kimi K3, and Qwen were highlighted as more honest alternatives worth defaulting to right now. Grok 4.6 tested well as a fast orchestrator (though still behind Sol and Fable on Terminal Bench v3.0), and Grok Bot's per-agent virtual desktops with 16GB RAM — capable of signing into apps and continuing work offline — were framed as a genuine consumer-grade leap, essentially a mainstream Hermes. On the coding harness side, builders converged on multi-model routing: Fable or Opus 5 for planning, Codex/Sol or Grok for execution, and Gemini for quick no-nonsense tasks. GPT-5 sol-xhigh consistently over-engineers from scratch but polishes existing code well. OMP users reported bloated outputs traced to imported Claude Code and Codex agent profiles — stripping those and keeping only native OMP agents restored quality. An unexpected Codex weekly usage reset landed late in the day, and Google promoted Koray Kavukcuoglu to head DeepMind, signaling a shipping-focused era for Gemini. Creative and infra threads rounded out the day: @arielletolome scaled Seedance 2.5 UGC ads from $200 to $500/day using full unedited videos (prompt away from on-screen text or switch to Kling), @swh800 shipped a Hermes-driven image batch pipeline running 250 images overnight on a flat Codex plan, and RTX 6000 Pro Blackwell cards climbed to ~$15k CAD as @jcartu chased a full-precision GLM 5.2 rig with 1M context. @jcartu also made the case for Hindsight as SOTA memory (top of LongMemEval and BEAM), with Mastra and Mem0 as close-enough alternatives.

Topics

DS4Pro-0813 launched fast (~65 tok/sec, ~2x K3) and dirt cheap with SSD-backed caching, but real-world testing exposed benchmaxxing — @jcartu's Rubik's cube build failed after 90 minutes, and there's no vision support. GLM, Kimi K3, and Qwen emerged as the community's preferred 'honest' picks, with only the DeepSeek flash variant getting a pass for its size.

Grok 4.6 tested well as a fast, competent orchestrator, though Terminal Bench v3.0 still puts it behind Sol and Fable. Grok Bot's headline feature — persistent per-agent virtual desktops with 16GB RAM that sign into apps and continue work offline — was called a normie Hermes and a real consumer-grade UX leap. @jarvisballer already prefers Grok over Fable for execution work.

@arielletolome pushed daily UGC spend from $200 to $500 by running full unedited Seedance 2.5 videos with b-rolls and POV shots. Community consensus: prompt away from on-screen text (Seedance mangles it, Kling handles it better), and Seedance 2 mini is the higher-volume, lower-cost play. Offer strength matters far more than production polish.

Builders converged on splitting roles: Fable or Opus 5 for planning, Codex/Sol or Grok for execution, Gemini for quick tasks. GPT-5 sol-xhigh over-engineers from scratch but polishes well. OMP users fixed bloated outputs by stripping imported Claude Code and Codex agent profiles (@Tz1888), and Kieran flagged auto-resume of to-do lists as worth disabling. Codex hit an unexpected weekly usage reset.

@jcartu positioned Hindsight as SOTA memory (leads LongMemEval and BEAM), with Mastra and Mem0 close enough that switching rarely pays off. @swh800 shipped a Hermes-powered image pipeline running 250 images overnight on a flat Codex plan instead of paying Kie API per call. RTX 6000 Pro Blackwell hit ~$15k CAD; four cards unlock full-precision GLM 5.2 with 1M context locally.

Key Takeaways

  • DeepSeek V4 Pro looks great on paper but fails real workloads — default to GLM, Kimi K3, or Qwen until independent evals catch up.
  • Grok Bot's per-agent VMs (16GB RAM, persistent, app-authenticated) are a genuine consumer UX shift, even if the underlying model is derivative.
  • Best coding results come from role splitting: Fable/Opus plans, Codex/Sol or Grok executes, Gemini handles quick tasks — never let GPT-5 sol-xhigh start from zero.
  • If OMP outputs feel bloated, purge agent profiles imported from Claude Code and Codex; keep only native OMP agents.
  • Full unedited Seedance 2.5 videos scale UGC ad spend past $500/day — prompt away from on-screen text or hand text-heavy scenes to Kling.

Hot Threads

@jcartustarted

DeepSeek V4 Pro real-world testing, benchmaxxing, and honest-model alternatives

22 replies7 participants
@realcrischicostarted

Fable vs Opus 5 vs Codex — multi-LLM routing and OMP harness cleanup

32 replies11 participants
@arielletolomestarted

Seedance 2.5 scaling UGC ads to $500/day with unedited videos

20 replies7 participants

Linked Items