MiniMax H3 Local Video, Claude Subagents, RTX 6000 — AI Daily Aug 4
496 messages · 71 active members
Overview
Topics
Builders benchmarked H3 across the full GPU range: 90s per 5s clip on a 5090, 15s generation in under 5 min on RTX Pro 6000 via ComfyUI-Spectrum + Kijai sage attention + SeedVR2, ~30 min for 9-step on a 3090, and 40 min on a 3070. A pruned build fits in 5GB VRAM, with quality approaching Seedance 2.
@tounano shared that the general-purpose subagent inherits the orchestrator's reasoning effort — the caller can only override the model. Solution: define explicit high/low-effort agent variants (e.g. general-purpose-high on Opus) so a cheap orchestrator can escalate to Opus-high selectively, cutting cost without losing quality.
@jcartu is hunting six RTX 6000s at €13k each between Moscow and Dubai for a local inference build. Discussion covered export availability, 12-24 month financing making ownership competitive with rental, and the case for keeping business inference fully off cloud to prevent alpha leakage.
Multiple members reported Fable feeling 'lobotomized' lately and have switched to GPT Sol or Opus 5 for daily coding work. Fable still holds ground for system design and less-technical vibe-coders, but consistency complaints dominated the thread.
Kieran, @tounano and @rmktg debated autonomy limits: ad uploads and batch script testing are fine to fully automate, but subjective data interpretation and high-volume budget allocation still need human-in-the-loop. @rmktg's 140k-LOC Hermes buildout hit a bloat wall — he's now considering an MCP webapp with lighter-weight agents.
Key Takeaways
- MiniMax H3 15s video generation is down to ~5 min on RTX Pro 6000 using ComfyUI-Spectrum fork + Kijai sage attention + SeedVR2 upscaling; pruned builds fit in 5GB VRAM.
- Claude Code's general-purpose subagent inherits orchestrator reasoning effort — define explicit high/low agent variants to let cheap orchestrators escalate to Opus-high selectively.
- RTX 6000s hitting €13k in Moscow/Dubai; 12-24 month financing now makes ownership competitive with cloud rental for full-local business inference.
- GPT Sol is winning daily coding workflows over Fable on reliability, though Fable retains fans for system design and vibe-coding.
- Never run tax or sensitive business strategy through cloud LLMs — assume sessions are logged; @jarvisballer saved ~$7k prepping scenarios locally before a human advisor.
Hot Threads
MiniMax H3 local setup, benchmarks, and Spectrum fork optimization across 3070/3090/5090/RTX 6000
Limits of full autonomous creative and media buying agents
Fable frustrations vs GPT Sol consistency for coding