TypeSafe Jev, GLM 5.3 & OMP Delegation — AI Daily Sep 16

656 messages · 69 active members

656
messages
69
active members
@anonymous, @samb69, @leewardbound
top contributors

Overview

September 16 was dominated by the launch of TypeSafe's Jev, a small, fast decision engine that returns confidence scores for classification, ranking, and boolean choices. Builders scrambled for waitlist access, tested it on ad copy ranking, Meta Ads screenshot analysis, and agent orchestration, and debated its scope versus OpenAI-style Structured Outputs. Consensus emerged around a clean rubric from @geilt: classify with TypeSafe, reason with LLMs. @leewardbound framed the bigger thesis — small models will handle 95%+ of inference as structured calls while big LLMs focus on the dev loop — though the open-source HuggingFace release is reportedly a dud versus the hosted version. @geilt also shipped TypeSafe CLI, and Kieran and others are holding off pending clearer use cases. Model selection and orchestration filled the rest of the day. Builders tiered their stacks: GLM 5.3 flash as the daily workhorse on established codebases with huge quotas, Fable 5.1 for mid-tier work, and Astra 6 reserved for the hardest problems — capped at 2–3 hour autonomous runs with forced realignment to prevent drift. Grok 4.7 began rolling out and slotted into coding workflows alongside OMP's new `^[Model Name]` syntax for delegating subagent tasks. JC shared his stack (Astra/Fable for ADRs, GLM for big builds, Gemini 3.8 Flash for Hermes), and the Grok Bot Galaxy livestream helped non-devs grok orchestration, though skeptics noted a properly configured Hermes is model-agnostic and matches Grok Bot. Go-to-market threads covered Franky Shaw-style 6–8 minute storytelling video ads (~$25–50/gen, modular scripts yielding 30+ variants tested for hook rate and CVR), Maria Wendt's tightly scoped micro-courses and WarriorBabe's $30M+/year fitness model as low-ticket funnel templates, real-time personalized landing pages (pre-render 100 sections, pick a winner), and Leewardbound's verdict that chart-only crypto trading bots lose to commodities and equities because of richer external signals.

Topics

TypeSafe launched Jev — a fast, cheap decision model for structured outputs — and @geilt shipped TypeSafe CLI with confidence scores so agents can auto-resolve ambiguity. Builders tested it for hook ranking, Meta Ads analysis, and session-turn validation, concluding it excels at choices, ranking, and booleans but falls short of nested schemas needed for full tool calling. Rule of thumb: classify with TypeSafe, reason with LLMs.

GLM 5.3 flash emerged as the daily workhorse for established codebases with huge quotas, Fable 5.1 for important-but-not-hard work, and Astra 6 reserved for the hardest tasks via multi-terminal orchestration with Herdr and Orca. Best practice: cap autonomous runs at 2–3 hours with forced realignment to prevent drift into unrequested features.

OMP shipped `^[Model Name]` syntax to delegate subagent tasks to specific models like Fable 5.1 or Grok 4.6, and Grok 4.7 began rolling out into coding workflows. JC shared his stack (Astra/Fable for planning, GLM for big builds, Gemini 3.8 Flash for Hermes), while the Grok Bot Galaxy livestream helped non-devs grasp orchestration — though skeptics argued a properly configured Hermes is model-agnostic and matches Grok Bot.

Franky Shaw-style 6–8 minute drama ads dominated the ads chatter at ~$25–50 per generation plus editing. Modular scripting with multiple characters lets teams generate 30+ variants to test hook rate and CVR, and government attention on some creatives signals the format is landing.

Discussion on Maria Wendt's tightly scoped micro-courses ($2–3 price point, few-hundred-dollar LTV via cross-sells) and WarriorBabe's $30M+/year fitness model as templates. Samb69 also pitched real-time personalized landing pages, refined to pre-rendering 100 section options and picking a winner rather than full runtime generation.

Words worth knowing

Three terms from the glossary.

Key Takeaways

  • Classify with lightweight decision engines like Jev; reserve LLMs for genuine multi-step reasoning — Jev shines on choices, ranking, and booleans but falls short on nested tool-call schemas.
  • Tier your models: GLM 5.3 flash as workhorse, Fable 5.1 for important work, Astra 6 for hardest problems — and cap autonomous runs at 2–3 hours with forced realignment.
  • OMP now supports `^[Model]` syntax for delegating subagent work; a properly configured Hermes is model-agnostic and matches Grok Bot for most orchestration needs.
  • Storytelling video ads work as modular scripts — one script, 30+ character/edit variants, tested for hook rate and CVR at ~$25–50/gen.
  • For trading bots, external context signals (news, earnings, tweets) beat chart data — commodities and equities win over crypto outside the top 5 assets.

Hot Threads

@anonymousstarted

Getting Jev access and figuring out what it actually does

45 replies12 participants
@geiltstarted

TypeSafe CLI launch and decision-engine use cases

9 replies3 participants
Kieranstarted

Low-ticket high-TAM info offers and removing key-person risk

14 replies4 participants

Linked Items