MiniMax H3 Local Video, Qwen 3.8-Max, Kimi K3 — AI Daily Aug 3

455 messages · 71 active members

455
messages
71
active members
@arielletolome, @fmill1, @jonmacofficial
top contributors

Overview

MiniMax H3 dominated the day as builders got a strong open-weight video model running locally on a single 5090 — generating 5s cinematic clips in ~3 minutes and 15s 768p scenes with dialogue in ~10-11 minutes. A WanGP-compatible pruned INT8 checkpoint and a Heretic uncensored ComfyUI build both circulated, and multiple members are already queuing 50 overnight generations, effectively ending Seedance and Fal.AI bills. Comparisons put H3 above LTX and on par with SD2, with the main limitation being weak 30s generations (workaround: stitch two 15s clips via first/last frame). On models, Qwen 3.8-Max landed at 2.4T parameters with open weights next week, cementing a new ~2.8T baseline that pushes serious local inference to 8-12 cards even at 3-bit — @jcartu is finding GLM 5.2 near-lossless at 3.34-bit. For coding and design, consensus formed that Kimi K3 beats Claude for UI/design work, with the recommended stack being Opus/Fable to set theme and rules, then Luna(max) fast mode or Deepseek V4 Flash for implementation (one user burned 1B tokens in 48 hours). Lavish Axi and Plannotator are emerging as go-to UI artifact tools. On ops, @bartekadamczyk shared A/B data showing ffmpeg-edited AI ads ran $158-270 CPM vs $67-91 for CapCut edits of the same clips — editing tool choice materially affects Meta performance. Hardware debates covered rumored M5 Ultra (768GB) and M7 Ultra (1.5TB) Mac Studios against reports of M5 Max MacBook Pros overheating with $895 repair bills, plus workflow discussion on SSH-ing from MacBook Air into a desktop for AI-heavy dev.

Topics

MiniMax H3 dropped on Hugging Face with WanGP-compatible pruned INT8 and a Heretic uncensored ComfyUI build. Builders are generating 5s clips in ~3 min and 15s 768p scenes with dialogue in ~10-11 min on local 5090s, killing Seedance/Fal.AI bills. Verdict: matches SD2, beats LTX, handles character likenesses aggressively well; 30s generations are weak so users stitch two 15s clips via first/last frame.

Qwen 3.8-Max was announced at 2.4T parameters with open weights next week, signaling a ~2.8T industry baseline. @jcartu noted local inference now realistically needs 8-12 cards even at 3-bit, and he's pushing GLM 5.2 toward 4-bit with 3.34-bit currently near-lossless. Practical floor is 3.4-3.5 bit, ideal is 4-bit.

Kimi K3 is outperforming Claude for UI/design work, while Anthropic models still lead general frontend. Recommended stack: Opus/Fable to set theme and rules, then Luna(max) fast mode or Deepseek V4 Flash for implementation. Opus 5 works better as advisor than Sol/Terra; Lavish Axi and Plannotator are new UI artifact tools worth adopting.

@bartekadamczyk shared A/B data on the same account: ffmpeg-edited AI ads ran $158 CPM vs $67 for CapCut edits of identical clips, and $270 vs $91 on a second test. Next iteration: using Grok to compare the two versions and reverse-engineer what CapCut is doing to the footage.

Rumors of M5 Ultra (768GB) and M7 Ultra (1.5TB) Mac Studios have builders eyeing Apple silicon for local inference of 2.8T-class models. Meanwhile M5 Max MacBook Pros are overheating with $895 repair bills, and MacBook Air M4 users report noticeable slowdown on terminal-heavy AI work — @jarvisballer recommends SSH-ing into a dedicated PC with Orca handling remote sessions well.

Key Takeaways

  • MiniMax H3 runs on a single 5090 — 15s 768p clip in ~10 min at zero per-generation cost, ending Seedance/Fal.AI dependency for many builders.
  • Qwen 3.8-Max (2.4T params) signals a new ~2.8T baseline; serious local inference now needs 8-12 cards even at 3-bit, with GLM 5.2 near-lossless at 3.34-bit.
  • For design work Kimi K3 beats Claude; best stack is Opus/Fable for theme + Luna(max) fast mode or Deepseek V4 Flash for implementation.
  • Same AI video clips edited in CapCut vs ffmpeg showed 2-3x lower Meta CPMs — editing tool choice materially affects ad performance.
  • Luna(max) on fast mode is cheap enough to leave on as a main driver — one user burned 1B tokens in 48 hours.

Hot Threads

@arielletolomestarted

MiniMax H3 local generation, uncensored build, and Fal.AI cost escape

25 replies9 participants
@thewildzenostarted

Model selection for UI/design + MacBook Air bottlenecks vs SSH to desktop

15 replies5 participants
@bartekadamczykstarted

ffmpeg vs CapCut CPM test results on identical AI clips

10 replies5 participants

Linked Items