JBS Weekly

Three different labs shipped three different flagship AI models in three straight days this week. If you're trying to run a business and not just track AI news, that's not progress, it's noise. Here's what actually changed with each one, and a plain five-step test for picking the right one for the job in front of you.
This Week’s Signal

Three flagship models shipped in three days, from three labs. Anthropic's Claude Fable 5.1 launched September 1st with record reasoning scores, at the same $10-per-million-token price as its predecessor. Google's Gemini 3.8 Flash followed September 2nd at $0.75 per million input tokens, already beating Anthropic's Opus 5 on benchmarks like contract review, though that price is only locked through year end. OpenAI's GPT-6 Astra shipped September 3rd at the same price as Fable 5.1, and it's the first OpenAI model to cross the company's own "critical" cybersecurity risk threshold.
None of the three is simply "the best model." Fable 5.1 is priced for depth: expensive per token, though a new caching discount rewards reusing the same context repeatedly. Gemini 3.8 Flash is priced for volume: cheap enough to run thousands of routine queries without the bill mattering. Astra costs the same as Fable 5.1 but is positioned around raw capability, a signal about what work OpenAI expects it for.
Picking a model by whichever one topped this week's benchmark chart is a losing game, because a different one tops a different chart next week. What actually matters is which one fits the specific, repeatable task you're trying to get done.
The Playbook: The Four-Part Prompt Contract
You don't need to track every model launch to run AI well in your business. You need a short test to run whenever a new one shows up, so the decision takes five minutes instead of an afternoon of reading benchmark charts.
1. Name the task, not the tool. Write down the actual job: "draft replies to routine customer emails," not "use AI for customer service." A specific task tells you what to test for.
2. Check if it's repeatable or a one-off. A task you'll run hundreds of times a month should default to the cheapest model that clears your quality bar. A task you'll run rarely, but that has to be right, like a contract review, can justify a frontier price.
3. Price the task, not the token. A cheap model that needs three tries to get something right can cost more than an expensive model that gets it right once. Run the same real task through two or three models before committing.
4. Give it a trial period, not a permanent seat. Run a new model on a real task for two weeks before switching your whole workflow to it. Pricing and capabilities are moving fast enough this year that this week is proof of it.
5. Revisit quarterly, not every launch. You don't need to re-evaluate every time a lab ships something new. Pick a schedule, quarterly is enough, and only break it early if a price or capability change directly affects a task you already depend on.
Run any new model through these five questions before you build a workflow around it. If you can't answer the first one clearly, you're not ready to pick a model, you're ready to define the task.
From The Podcast
This newsletter is your one shot at the full week, since the daily show never gets its own email. All five episodes:
- This AI DM Bot Made $120K in Revenue While You Slept: the warm comment-to-DM funnel that actually converts, and why cold-scraping gets accounts banned.
- OpenClaw 2.0: 900 Contributors, One Massive Security Problem: what actually changed in the biggest release in the project's history, and what still hasn't.
- Instagram Put a Reach Penalty on Undisclosed AI Profiles: the new label, who it targets, and why hiding it costs more than showing it.
- Why This 24/7 AI TV Channel Costs $4,000 a Day to Run: inside the AI video model rendering clips faster than you can watch them.
- The Cheapest AI Model Just Beat GPT-5.6 and Opus 5: why cost-per-completed-task just replaced the benchmark wars.
Tool Worth Trying
Google Workspace
If Gemini 3.8 Flash from this week's signal has you curious, you don't need a separate subscription to try it. A Google Workspace Business Standard plan includes Gemini Pro built into every app, plus NotebookLM Plus and the enterprise version of the whole suite. If you're already paying for Google Workspace, or thinking about moving off a clunkier email setup, this is a cheaper way into the Gemini lineup than a standalone subscription, and it puts the Model Fit Test above one step closer to something you can actually run today.
14-day trial, 10% off your first year.
Link: Google Workspace
Joe’s Take
When new models come out, I tend to do the opposite of everyone else in my field. Instead of jumping in, I wait.
That habit comes from my 15-year career as a UAT Manager at a global market research company. It's also just how I'm wired: pattern recognition is one of my strengths. By watching how everyone else uses a new model and shares their results, I get a clearer read on the real use cases than I would testing it myself in a vacuum. That's what lets me build educational material based on how people are actually using these tools, not some cookie-cutter PDF.
The catch: a week like this one, with three releases instead of one, makes that approach a lot harder to pull off.
Tools I Use
n8n — The automation tool I use to connect apps, trigger workflows, and stop doing things manually. If there's a repetitive process in your business, this is where you start fixing it.
VoiceInk — A local AI dictation tool for Mac that transcribes your voice with near-perfect accuracy and runs entirely on your device, meaning nothing you say ever touches a cloud server.
Blotato — Handles the full content distribution side of your business: drop in a topic and it generates platform-specific posts, or feed it existing content and it repurposes it across formats. TikTok videos become tweets, podcasts become blog posts. Includes a scheduling calendar, visual creation tools for carousels and infographics, and publishes natively to 9 platforms with no per-post fees.
Beehiiv — What you're reading right now is published on Beehiiv. If you're thinking about starting a newsletter or moving off a clunky platform, this is the one I'd recommend. 20% off your first 3 months with my link.
Google Workspace — If Gemini 3.8 Flash from this week's signal has you curious, you don't need a separate subscription to try it. A Business Standard plan includes Gemini Pro built into every app, plus NotebookLM Plus and the enterprise version of the whole suite. If you're already paying for Google anyway, it's a cheaper path into the Gemini lineup than a standalone subscription. 14-day trial and 10% off your first year.
Descript — Video and podcast editing that works like a text document. You edit the transcript and the media follows. Cuts filler words, cleans up audio, and handles captions automatically. 50% off your first two months on the Creator Plan.
Final Thoughts
Three labs shipped three flagship models within 72 hours of each other this week, at three different price points, built for three different kinds of work. The honest takeaway isn't excitement, it's that tracking every release is no longer a reasonable way to run a business. The businesses that get real value out of any of these models aren't the ones chasing whichever name is loudest that week. They're the ones with a repeatable test for matching the right model to the right task.
PS: If you're evaluating any of the three this week, start with the cheapest one that could plausibly do the job. Downgrading later is easy. Rebuilding a workflow around an expensive model you didn't need is not.
Cheers,
Joe
