Hey — and welcome to Sunday Signal.
This week I read twenty emails my AI had written for me. All twenty ended with the same sentence.
Not similar. The same. The middles came in three families too: eleven near-identical, then four, then five.
That is not a writing problem. That is a copying machine doing exactly what it was shown.
I had rewritten the rules for that system twice. Careful rules, specific bans, the lot. And underneath them sat six sample emails I had pasted in months earlier and forgotten about. The rules said one thing. The samples showed another.
The samples won. Every single time.
📡 The Signal
1. A 2.4 trillion parameter model landed on Monday, and the price floor keeps falling
Alibaba launched Qwen3.8-Max on August 3. It edges out Claude Opus 4.8 on one agentic benchmark, runs a million-token context, and costs $2 per million tokens in. Open weights are on the way. Meanwhile the cheap end is getting silly: DeepSeek's latest coding model runs near the frontier at roughly one eighty-ninth the output price of a top-tier model.
Why it matters for you: The model stopped being your bottleneck a while ago. When near-frontier work costs pennies, the gap between your output and someone else's is not the tool. It is the quality of what you put in front of it.
2. Google reshuffled the top of its AI org
On August 6, Demis Hassabis stepped down as DeepMind's CEO to become chair and Alphabet's chief scientist, saying he wanted room to focus on the big picture. Koray Kavukcuoglu, already leading Gemini, took over daily operations. The same day, Jeff Dean left after 27 years to start something new.
Why it matters for you: Note what Hassabis kept. Not the output, the direction. Deciding what good looks like is a real job, and it is the same job you are doing every time you show a tool an example rather than another instruction.
3. The big labs are heading to the White House on safety testing
OpenAI, Anthropic and Google are joining a White House meeting on a voluntary framework for safety-testing AI models, stemming from a June executive order.
Why it matters for you: Nobody is being asked to describe how safe their model is. They are being asked to demonstrate it on a test everyone can see. Claims are getting cheaper and demonstrations are getting more valuable, and that shift will reach small businesses long before it reaches a law.
Indianapolis Business Journal — https://www.ibj.com/articles/openai-anthropic-google-to-join-white-house-ai-safety-meeting
🌶️ Hot Take
Everyone is busy writing better prompts. Almost nobody checks the examples sitting underneath them.
If you have ever pasted "here's one I wrote earlier" into a chat, that sample is now doing more work than every instruction you typed after it. Your AI is not really reading your rules. It is studying your samples and handing you more of the same. A mediocre example quietly cancels a brilliant prompt, and you will spend a week blaming the model.
✦ The Build Log
The fix was not cleverer rules. It was six new examples.
I threw out all six and wrote fresh ones, deliberately different from each other, because three genuinely varied examples beat ten that rhyme. The identical closes stopped immediately.
There was a second thing hiding in the same mess, and it is the more embarrassing one. In the previous version I had banned a whole category of language from those emails. The ban worked perfectly. What I had not noticed was that the banned paragraph was carrying the only proof that I knew what I was talking about, and I never replaced it. So for weeks the system produced clean, rule-abiding, completely unconvincing emails.
Take something out, leave a hole. Every ban needs a replacement or you have just deleted a reason to trust you.
Elsewhere it was a week for looking at real data instead of guessing. My trading bots had an ugly stretch, so I sorted every losing trade by cause rather than theorising, and the pattern was obvious within an hour. Nineteen fresh emails are now queued. Whether the new formula works is genuinely unknown until someone replies.
Week 11 of building in the open. What it actually looked like:
🔧 The fleet teardown · ~7h / ten fixes across the agents, one found dead since July, and every scheduled job now logs its own failures instead of dying quietly
✉️ The email rebuild · ~4h / read all twenty drafts end to end before changing anything, then replaced the examples
📉 The loss audit · ~3h / sorted every losing trade by cause rather than by feeling, then shipped seven guards
🎬 Two videos, three carousels · ~12h / plus the long-form finally in front of a camera
✦ Try This
Audit what you are showing your AI, not what you are telling it. Ten minutes.
First, find your examples. Open whatever you use to get AI to write for you: a saved prompt, a custom assistant, a project, a document you paste in every time. Look for any sample you pasted, any "match this style," any old piece you told it to copy.
If you found some, audit them:
Read them as a stranger would, back to back. Ignore your instructions entirely.
Name what they share. Same opener, same closing line, same rhythm, same length. Whatever they share is what you will keep getting, regardless of what your rules say.
Replace them. Two or three genuinely different good ones beat ten that sound alike. Write them yourself. Rough is fine, samey is not.
If you found none, that is the finding, and it is good news. It means the tool has been guessing at your taste from nothing. Five minutes fixes it:
Go find two things you have written that you were actually happy with. Real ones, from your own work.
Make sure they differ from each other. Two samples that sound the same teach a narrower lesson than one.
Paste both in above your usual request, with one line: "Match the voice and structure of these two, not the topic."
Then test it properly. Run your normal request twice, once with the old setup and once with the new, and ask one question of each result: could this have been written for somebody else? If yes, your examples are still too generic.
I'm collecting the before-and-afters. Reply and send me yours, especially if the second version surprised you.
✦ Closing
Thirteen Saturdays in. The arc so far: ship, constrain, build, compound, patch, subtract, ask, remember, price, staff, equip, doubt, and now show. Telling a machine what you want is the easy half. Showing it is the half that decides what comes back.
That is the thing these tools have quietly taught me about my own work. I used to think being clear was the same as being understood. It is not. Clear instructions plus one lazy example will lose to a good example every time, and that turns out to be just as true of the people I work with as the software.
The long-form is finally shot. It has slipped more times than I would like to admit, so I am not putting a date on it. You will know when it lands.
If this was useful, forward it to one person who keeps blaming the model. They can read every issue free here: https://sundaysignalhq.beehiiv.com/?utm_source=sunventures&utm_medium=site&utm_campaign=launch
Quietly, deliberately. One Saturday at a time.
— Kalpesh
P.S. Work with me. Everything I build runs on the same idea as this issue: show the machine what good looks like, then let it work.
The products, including the free Permission Pack: https://stan.store/buildwithkp
Hire me to build it with you: https://sunventures.studio/hire
Everything else: https://sunventures.studio/

