Grok 4.7, Gemini 4 Pro Leak, Agentic Commerce — AI Daily Sep 21

532 messages · 77 active members

532
messages
77
active members
@anonymous, @samb69, @leewardbound
top contributors

Overview

Sunday opened with @robinroy sharing a leaked benchmark claim that Gemini 4 Pro beats Astra on every benchmark at 6x lower cost with a 2M context window. Skepticism ran high (@samb69 called it 'hopium', @jcartu declared 'there is no AGI'), but builders agreed cutthroat competition benefits everyone. Hours later SpaceXAI dropped Grok 4.7 with a 2x fast mode at half the cost of Sol — capped at 500k context. Early testers reported it prefers one-shot execution over todo planning and often lands correct outputs first try, though benchmark scores are mixed. @william_hodges ran it autonomously all day on a 4-feature PRD, burning 35% of his weekly Supergrok Plus quota. Agentic commerce dominated the middle of the day after Amazon reportedly blocked Muse's agentic shopping. @leewardbound pushed back on the '90% of traffic will be agentic' framing, arguing $spend unlocked and non-shopper onboarding (like parents) are the real metrics, projecting 3-5 years for meaningful cultural shift. Builders shared real workflows — booking Airbnb via Instinct AI in 2 minutes — while @samb69's consulting story about ecom staff shipping raw Claude output crystallized concerns about 'slop grenades' replacing real QA. On tooling, Claude went down and throttled hard, with 20x accounts burning ~4x faster than usual, pushing people toward Fable, Opus 4.6 med, and Baseten's DeepSeek 4.1 flash. @julhi123 shared HarnessRouter (OpenRouter-style unified API across Codex, Claude Code, Hermes, Jev), while @samb69 is now orchestrating 5 remote machines from a single Herdr window. A standout thread explored AI for football film analysis — a coach's 196-videos-per-game workflow — with consensus that a classification/tagging layer is needed before LLM ingestion.

Topics

SpaceXAI released Grok 4.7 with a fast mode at half the cost of Sol. Early testers report it one-shots tasks instead of building todo lists and often lands correct outputs, though benchmarks are mixed and the 500k context cap disappoints vs Fable/GLM's 1M. @william_hodges burned 35% of weekly Supergrok Plus quota running it autonomously on a 4-feature PRD.

@robinroy shared a leak from Google contacts claiming Gemini 4 Pro beats Astra on every benchmark at one-sixth the cost with a 2M context window. Reactions split between excitement and skepticism, with @greglavidaloca noting nobody has run the benchmarks themselves. No release date confirmed.

@kryzk1 flagged Amazon blocking Muse for agentic commerce. @leewardbound pushed back on the '90% of web traffic will be agentic' framing, arguing the meaningful metrics are $spend unlocked and onboarding non-shoppers like parents, projecting 3-5 years for cultural shift. Real examples included booking Airbnb via Instinct AI in 2 minutes.

Claude went down and throttled hard, with 20x accounts lasting 9 hours instead of 36. Members shifted to Fable high, Opus 4.6 med, and Baseten's DeepSeek 4.1 flash. @julhi123 shared HarnessRouter — a unified API across Codex, Claude Code, Hermes, and Jev — while @samb69 is orchestrating 5 remote machines from a single Herdr window with parallel agent panes.

@thewildzeno asked for Gemini 3.8 Flash alternatives for analyzing thousands of videos; @anonymous recommended TwelveLabs for bulk search/index/embed. Separately, a coach with 3000+ game videos (196 clips/game) can'anonymous get Claude to grasp football fundamentals — @leewardbound and @iannagy suggested a classification layer (à la Jev) to tag events before LLM ingestion, noting soccer tooling is more mature.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Grok 4.7 prefers one-shot execution over todo planning and often lands correct outputs first try — but 500k context cap limits long orchestration vs Fable/GLM's 1M.
  • Gemini 4 Pro leak (2M ctx, 6x cheaper than Astra) is unverified — treat as hopium until benchmarks reproduce, but competitive pressure is real.
  • Agentic commerce metrics should be $spend unlocked and non-shopper onboarding, not '% of web traffic' — expect 3-5 years for meaningful normie adoption.
  • Claude 20x accounts burning ~4x faster today; Fable high, Opus 4.6 med, and Baseten DeepSeek 4.1 flash are viable fallbacks for bounded tasks.
  • Video/film analysis needs a classification/tagging layer before LLM ingestion — raw footage into Claude won'anonymous grasp domain fundamentals like football down or block assignments.

Hot Threads

@leewardboundstarted

Debunking '90% of traffic will be agentic' and real agentic commerce metrics

20 replies6 participants
@anonymousstarted

AI for football film analysis and untapped sports app market

20 replies5 participants
@robinroystarted

Gemini 4 Pro leak beats Astra at 6x cheaper cost

15 replies10 participants

Linked Items