MiniMax H3 Local Video, Qwen 3.8-Max, Kimi K3 — AI Daily Aug 3
455 messages · 71 active members
Overview
Topics
MiniMax H3 dropped on Hugging Face with WanGP-compatible pruned INT8 and a Heretic uncensored ComfyUI build. Builders are generating 5s clips in ~3 min and 15s 768p scenes with dialogue in ~10-11 min on local 5090s, killing Seedance/Fal.AI bills. Verdict: matches SD2, beats LTX, handles character likenesses aggressively well; 30s generations are weak so users stitch two 15s clips via first/last frame.
Qwen 3.8-Max was announced at 2.4T parameters with open weights next week, signaling a ~2.8T industry baseline. @jcartu noted local inference now realistically needs 8-12 cards even at 3-bit, and he's pushing GLM 5.2 toward 4-bit with 3.34-bit currently near-lossless. Practical floor is 3.4-3.5 bit, ideal is 4-bit.
Kimi K3 is outperforming Claude for UI/design work, while Anthropic models still lead general frontend. Recommended stack: Opus/Fable to set theme and rules, then Luna(max) fast mode or Deepseek V4 Flash for implementation. Opus 5 works better as advisor than Sol/Terra; Lavish Axi and Plannotator are new UI artifact tools worth adopting.
@bartekadamczyk shared A/B data on the same account: ffmpeg-edited AI ads ran $158 CPM vs $67 for CapCut edits of identical clips, and $270 vs $91 on a second test. Next iteration: using Grok to compare the two versions and reverse-engineer what CapCut is doing to the footage.
Rumors of M5 Ultra (768GB) and M7 Ultra (1.5TB) Mac Studios have builders eyeing Apple silicon for local inference of 2.8T-class models. Meanwhile M5 Max MacBook Pros are overheating with $895 repair bills, and MacBook Air M4 users report noticeable slowdown on terminal-heavy AI work — @jarvisballer recommends SSH-ing into a dedicated PC with Orca handling remote sessions well.
Key Takeaways
- MiniMax H3 runs on a single 5090 — 15s 768p clip in ~10 min at zero per-generation cost, ending Seedance/Fal.AI dependency for many builders.
- Qwen 3.8-Max (2.4T params) signals a new ~2.8T baseline; serious local inference now needs 8-12 cards even at 3-bit, with GLM 5.2 near-lossless at 3.34-bit.
- For design work Kimi K3 beats Claude; best stack is Opus/Fable for theme + Luna(max) fast mode or Deepseek V4 Flash for implementation.
- Same AI video clips edited in CapCut vs ffmpeg showed 2-3x lower Meta CPMs — editing tool choice materially affects ad performance.
- Luna(max) on fast mode is cheap enough to leave on as a main driver — one user burned 1B tokens in 48 hours.
Hot Threads
MiniMax H3 local generation, uncensored build, and Fal.AI cost escape
Model selection for UI/design + MacBook Air bottlenecks vs SSH to desktop
ffmpeg vs CapCut CPM test results on identical AI clips