Grok 4.7, Gemini 4 Pro Leak, Agentic Commerce — AI Daily Sep 21
532 messages · 77 active members
Overview
Topics
SpaceXAI released Grok 4.7 with a fast mode at half the cost of Sol. Early testers report it one-shots tasks instead of building todo lists and often lands correct outputs, though benchmarks are mixed and the 500k context cap disappoints vs Fable/GLM's 1M. @william_hodges burned 35% of weekly Supergrok Plus quota running it autonomously on a 4-feature PRD.
@robinroy shared a leak from Google contacts claiming Gemini 4 Pro beats Astra on every benchmark at one-sixth the cost with a 2M context window. Reactions split between excitement and skepticism, with @greglavidaloca noting nobody has run the benchmarks themselves. No release date confirmed.
@kryzk1 flagged Amazon blocking Muse for agentic commerce. @leewardbound pushed back on the '90% of web traffic will be agentic' framing, arguing the meaningful metrics are $spend unlocked and onboarding non-shoppers like parents, projecting 3-5 years for cultural shift. Real examples included booking Airbnb via Instinct AI in 2 minutes.
Claude went down and throttled hard, with 20x accounts lasting 9 hours instead of 36. Members shifted to Fable high, Opus 4.6 med, and Baseten's DeepSeek 4.1 flash. @julhi123 shared HarnessRouter — a unified API across Codex, Claude Code, Hermes, and Jev — while @samb69 is orchestrating 5 remote machines from a single Herdr window with parallel agent panes.
@thewildzeno asked for Gemini 3.8 Flash alternatives for analyzing thousands of videos; @anonymous recommended TwelveLabs for bulk search/index/embed. Separately, a coach with 3000+ game videos (196 clips/game) can'anonymous get Claude to grasp football fundamentals — @leewardbound and @iannagy suggested a classification layer (à la Jev) to tag events before LLM ingestion, noting soccer tooling is more mature.
Words worth knowing
Three terms from the glossary.
- Ad CopyThe written text in an advertisement, often generated or refined with AI models.
- Wispr FlowA voice-to-text tool that fixes transcription errors intelligently, used throughout workflows.
- Deep ResearchUsing an AI model with extended prompts to gather, synthesize, and analyze information thoroughly.
Key Takeaways
- Grok 4.7 prefers one-shot execution over todo planning and often lands correct outputs first try — but 500k context cap limits long orchestration vs Fable/GLM's 1M.
- Gemini 4 Pro leak (2M ctx, 6x cheaper than Astra) is unverified — treat as hopium until benchmarks reproduce, but competitive pressure is real.
- Agentic commerce metrics should be $spend unlocked and non-shopper onboarding, not '% of web traffic' — expect 3-5 years for meaningful normie adoption.
- Claude 20x accounts burning ~4x faster today; Fable high, Opus 4.6 med, and Baseten DeepSeek 4.1 flash are viable fallbacks for bounded tasks.
- Video/film analysis needs a classification/tagging layer before LLM ingestion — raw footage into Claude won'anonymous grasp domain fundamentals like football down or block assignments.
Hot Threads
Debunking '90% of traffic will be agentic' and real agentic commerce metrics
AI for football film analysis and untapped sports app market
Gemini 4 Pro leak beats Astra at 6x cheaper cost