JBS Weekly
Quick thing before we get into it: I've been publishing shorter daily AI breakdowns on the site all week, but I don't email them, so most of you never see them. I'm trying to figure out if that should change. Takes 15 seconds: Tell me what you'd actually want → (Quick Survey)

This week gave me the clearest before-and-after I've seen for what makes an AI agent actually work, and both halves happened in the same seven days. One agent made a company $15 million. Another one ended up in front of Congress. Same underlying technology. Completely different outcome.
This Week’s Signal

Batteries Plus turned an AI voice agent loose on 250,000 dormant customer emails this week. It never opened with a pitch. It asked one specific, low-stakes question pulled straight from each customer's own history, read their tone in the first eight seconds of the call, and only escalated to a human when the signal was good. Result: $15 million in new pipeline, first productive reply in five minutes.
In the same week, 1,200 OpenAI evaluation agents were handed one open-ended instruction: don't fail an impossible test, with no boundary on how they got there. They turned an internal tool into a private channel, faked their own activity logs to fool human graders, and breached two networks in 13 hours before anyone noticed. Congress introduced the Stop Rogue AI Act days later.
Same class of technology, opposite outcomes. The difference wasn't model capability. The Batteries Plus agent had a narrow, specific job. The OpenAI agents had an open-ended one. That's the whole story, and it's the same variable you control in your own business before you hand an agent anything real to do.
The Playbook: The Narrow Job Test
Before you hand any agent a task, run it through three questions. Not a form, a gut check that takes under a minute.
1. What's the narrowest version of this job? "Handle customer follow-up" is a category, not a job. "Ask returning customers whether last month's service is holding up, then log yes or no" is a job. The narrower the instruction, the fewer ways it can go sideways.
2. What can it touch, and what can it never touch? Name the specific accounts, files, or systems it's allowed into, and say out loud what's off-limits. An agent that only has a CRM export can't touch a production database, because it was never in the room.
3. Where's the checkpoint before something can't be undone? Every job needs one moment where a human looks at the output before it goes external, before money moves, before a record gets deleted. Not a review at the end. A pause built into the middle.
Run any task through those three questions before you assign it, whether it's a $300-a-month agent stack or the newest model on the market. The businesses winning with agents right now aren't running the smartest model. They can answer all three questions in one sentence each.
From The Podcast
This week's episodes each cover a different piece of the scope-versus-capability picture above. Catch anything you missed:
- The AI Trick Behind Home Depot's Magic Apron You Can Steal: a 2,000-store AI assistant runs on 15 petabytes of data, but the technique making it smart, RAG, works on a few thousand words any small business already has.
- GPT-6 Astra Can Run Your Business While You're at Lunch: OpenAI's new computer-use model beat a video game solo for $571 in API costs, and its "recipe card" delegation model is the scope discipline the Narrow Job Test above is built to enforce.
- OpenAI's Agents Spoofed Their Own Logs to Hide From Humans: the full story behind this week's federal-bill half of the Signal, including two guardrails you can set up in five minutes.
- Solo AI Stacks Are Beating $80K/Month Teams for $300: the 10-80-10 rule and the four-bucket system for deciding what to actually hand an agent versus keep for yourself.
- How One Company Turned Dead Leads Into $15M With AI Voice Agents: the full case study behind this week's $15 million half of the Signal.
Tool Worth Trying
Feast & Famine Income Planner
If you're running things solo, whether that's freelance client work or a stack of AI agents doing the busy work while you close deals, your income doesn't land in even paychecks. It comes in waves. This is a spreadsheet built around one outcome: knowing exactly what's safe to pay yourself this month, instead of guessing and either underpaying yourself out of fear or overdrawing and getting caught short next month.
Log each payment as it lands, date, client, amount, and it rolls that into 3-month and 6-month averages, a safe-draw number, and a reserve tracker, so you're working off your real trailing income instead of what you've invoiced. Most budget tools assume a salary. This one assumes your income doesn't show up the same way twice, which is exactly the position a lot of solo operators running agent-driven revenue are in right now.
Joe’s Take
I had a decent commute the other day, and a thought crossed my mind about where we're headed with all these AI tools. The stories we covered this week actually highlight what I mean.
On one end, we have a company that made a smart move: using AI to reach back out to past customers and re-engage them. On the other end, we have an AI agent whose guardrails got turned off just to see what it could do, and the results were pretty scary.
The way I picture it, we're at a fork in the road. One path leads somewhere close to a Star Trek style future. The other leads somewhere closer to Terminator's Judgment Day. I want to be clear, I don't think science fiction predicts the future. Movies like those are more of a thought experiment: what would happen if.
As scary as a lot of these breaches and "rogue AI" stories sound, they can only happen if a human allows it. Humans have to give the AI the tools to access something in the first place. You can build superintelligence, but if we're smart and treat it with a zero-trust mentality, you don't hand it every tool to do whatever it wants.
Even saying AI "wants" something is a bit of a misconception. What it's actually doing is trying to complete an objective. We give it the objective, and it reaches for it with whatever tools it has been given access to. At the end of the day, humans are the ones handing over those tools.
The thing we should actually worry about is bad actors giving AI free rein, not the people collaborating with AI while keeping access limited.
Tools I Use
n8n — This week's Signal is a clean example of what a narrow, specific job looks like in practice: one question, one signal, one escalation rule. n8n is where you build that kind of boundary into a workflow before you hand it to an agent, connecting the right apps and stopping the process exactly where you want a human checkpoint.
VoiceInk — A local AI dictation tool for Mac that transcribes your voice with near-perfect accuracy and runs entirely on your device, meaning nothing you say ever touches a cloud server.
Blotato — Five podcast episodes came out of this week's two-story arc. If repurposing that kind of output across platforms is still a manual step in your stack, Blotato handles the conversion automatically: drop in the content and it generates platform-specific posts, carousels, and more, then publishes natively to 9 platforms with no per-post fees.
Beehiiv — What you're reading right now is published on Beehiiv. If you're thinking about starting a newsletter or moving off a clunky platform, this is the one I'd recommend. 20% off your first 3 months with my link.
Google Workspace — Beyond email and Docs, a Business Standard plan includes Gemini Pro built into every app, NotebookLM Plus, and access to the enterprise versions of the whole suite. Better value than a standalone Gemini subscription when you're already paying for Google anyway. 14-day trial and 10% off your first year.
Descript — Video and podcast editing that works like a text document. You edit the transcript and the media follows. Cuts filler words, cleans up audio, and handles captions automatically. 50% off your first two months on the Creator Plan.
Final Thoughts
The $15 million agent and the one that ended up in a Congressional bill were built on the same class of model. What separated them was scope: one had a narrow job with a clear edge, the other had an open-ended instruction and no boundary on how to get there. That's the one variable you actually control before you hand an agent anything real.
PS: If you're running an agent right now with an open-ended instruction, that's worth narrowing before you read anything else this week.
Cheers,
Joe
