
Welcome back to the AI Edge weekly newsletter!
In this publication, we're dropping our full guide on how to reduce your Claude token spend by 80% or even higher.
The sad reality is that most people are significantly overpaying for Claude.
With these simple tricks, you’ll be able to get the most out of your Claude subscription and usage limits.
After our main anchor piece, be sure to stick around for the Midweek Edge for all the latest AI news updates and our Looking Ahead section, where we compile our top AI research.
Let’s get right into this week’s edition!
Table of Contents
Planning 101
Here's the thing most people get wrong:
Simple chats aren't what drain your Claude usage.
Building is. Tasks like coding, drawing, generating - that's where most people’s token spend goes (especially if context is heavy).
So the fix is almost embarrassingly simple. We recommend splitting your Claude tasks into two phases:
Planning
Do your thinking first, before Claude builds anything. Map out what you actually want, and do it on a cheaper model. You don't need a heavy model like Fable 5 to help brainstorm on basic tasks.
Good planning models:
Kimi K3
GLM-5.2
Haiku
Building
Only once the plan is confirmed do you switch up and let Claude build.
The goal here is to get one clean pass instead of five messy ones.

Tip: If you're in Claude Code, there's a built-in Plan Mode for exactly this. Hit Shift + Tab twice or type /plan, and Claude focuses on planning before it touches anything.
For more complex tasks, we recommend a three-phase strategy:
Planning (heavy model)
Building (cheaper model)
Verification (heavy model)

Context Engineering
Long chats are a silent killer.
Every time you send a message in a long thread, Claude re-reads the entire conversation history to respond.
Example: If you’re a hundred messages deep, that's a hundred messages of old context getting re-read on every single reply.
Not only are you burning more tokens, but you’re also getting slower responses and worse answers because all that old clutter dilutes the important stuff.

Two ways to fix this:
Use Projects
Set up a project for a recurring task, load your instructions/context into it once, and open a fresh chat every time you start something new.
This way, each chat stays short and sharp, but they all share the same context.
Hand offs
When a chat does get long, don't just keep pushing.
Prompt Claude:
"I'm moving to a new chat. Give me a prompt that restarts this session with all our context, so I don't lose anything."
Paste that into a fresh chat, and you carry the important context without dragging the whole history along.
One more trick: tell your project instructions you're watching usage. Something like "be concise, and tell me when I should start a new chat."
Giving Claude Better Memory
By default, Claude is a black box. Meaning, it doesn't hold on to much about you between sessions, so you end up re-explaining who you are and how you like things
done. Every re-explanation is tokens you didn't need to spend.
The fix is a two-step local desktop folder setup that anyone can run:
instructions.md - your rules. Who you are, what you do, how you like
things done. Make sure to include “update this memory file with new preferences over time.”
memory.md - the running brain. This is where Claude writes down what it
learns about you. Things like your preferences, your corrections, the patterns it
notices. This grows on its own because of the line above in instructions.md.

This system in practice: You tell Claude, "Stop using em dashes." It reads your instructions, sees the rule about updating memory, and writes that preference into memory.md.
You can also expand this folder setup by creating folders for specific context pieces like your business, health goals, etc.
Model Selection
This is a big one.
Running your most expensive Claude model for everything is the single biggest way
to overspend.
You don't need a heavyweight to summarise an email, and 70% of daily tasks should rely on models like Sonnet 5.
Our advice: use the “escalate” method.
Start cheap, and only climb when the task actually demands it.
Haiku → light tasks (run in Chrome extension for easy access)
Sonnet → medium tasks. Most of your day-to-day work lives here.
Opus → heavy tasks. The hard reasoning, intense coding, etc.
Think of it as a ladder. You start on the bottom rung and only step up
when you hit something that genuinely needs more.
A few settings that quietly cut your spend:
Turn thinking off for most tasks.
Set your style to Concise on the Claude homepage.
Use Low effort in Claude Code for routine work.
Final Tips
A few quick wins to round things out.
Buy credits, don't jump plans. If you're only occasionally exceeding your usage limit, topping up with extra credits is cheaper than upgrading your entire plan tier. Something we often do is just add $100 in credits every 2-3 months.
Build Skills for repetitive work. Anything you do again and again, turn into a Claude Skill so it runs lean every time instead of being re-explained from scratch.
Track your usage. Run /usage in Claude Code, or check the new Overview section, so you always know where you stand. Just seeing that you're close to the limit is often enough to make you prompt smarter.
Start a fresh chat sooner than you think. When in doubt, open a new one. Anthropic has confirmed that short chats always beat long ones in Claude 5 models.
Midweek Edge
Our manually curated list of the most important news updates across AI, robotics & tech.
xAI Releases Grok Imagine Image 2.0
xAI just released Grok Imagine Image 2.0, its new model for image generation and editing.
→ Ranked #2 on the Text-to-Image Arena
→ Better text and typography
→ More precise image editing
Try it now at grok.com/imagine

OpenAI Releases GPT-5.6-Cyber
OpenAI has released GPT-5.6-Cyber, a new version of Sol built for advanced cybersecurity work.
It can handle tasks such as vulnerability research, exploit validation, and other complex defensive workflows.
Access is available to approved cybersecurity researchers and teams through Daybreak Red (expected to go public in the future).

Claude is ‘Watermarking’ AI-Generated Content
Anthropic is adding invisible, machine-readable watermarks to text generated by the newly supported Claude models.
Generated images and files will also include signed metadata indicating Claude processed them.

Meta Releases Muse Code
Meta is gaining attention again after recently releasing Muse Code, its new terminal coding agent powered by Muse Spark 1.2.
It can plan a task, write code, test its work, and run multiple subagents across larger projects.
Available now in beta. You can get started through Meta Developer.

Anthropic Cuts Sonnet 5 Pricing
Anthropic is making Sonnet 5's introductory API pricing permanent.
→ $2 per million input tokens
→ $10 per million output tokens
The planned September price increase is no longer happening.

Buzz for AI Agents
Buzz, built by Twitter’s former founder, Jack Dorsey, is a free, open-source workspace where humans and AI agents can work together inside the same channels.
Think Slack, but your AI agents can join the conversation as teammates with their own identities, permissions, and tasks.
Try it now at buzz.xyz on Mac, Windows, or Linux.

Looking Ahead
Our manually curated list of the top AI trends, research, workflows & more.
Fun Fact: An AI Agent Hacked a Gym to Book a Class
A man asked his AI agent to book a popular gym class.
Instead, the agent found a vulnerability in the booking system, booked beyond the allowed window, and removed another person from the waitlist without being asked to.
As agents gain more autonomy, be sure to give them approval checkpoints before they can delete, cancel, buy, or submit anything.

Claude May Drop Weekly Usage Limit
Some Claude users are reportedly seeing accounts without the normal weekly usage limit, leaving only the shorter session limit.
Anthropic has not announced any change yet, so this could simply be a test.

The Next AI Wave is Robotics
NVIDIA recently released Alpamayo 2 Super, an open reasoning model designed for autonomous vehicles.
Instead of only answering questions, these models can perceive the physical world, reason about what is happening, and decide what to do next.
AI is moving beyond purely digital applications into cars, robots, factories, and other machines.
Robotics is clearly one of the biggest AI categories to watch next.

Astra May Be OpenAI’s First Critical Model
OpenAI says its upcoming Astra model is powerful enough that it cannot rule out a "Critical" cybersecurity capability.
That would mean a model capable of detecting and finding serious vulnerabilities or carrying out complex cyberattacks with very little human help.
OpenAI has already paused some Astra work while it strengthens its security controls.

How to Kill AI Writing
AI-generated writing is getting easier to produce, but generic AI writing patterns are also becoming easier to notice.
One simple fix is adding an anti-AI-slop Skill as the final review step in your writing workflow.
It works across Claude Code, Codex, Cursor, Gemini CLI, and other agents.
Run your final draft through it before you publish. We recommend using this repo.

Claude 5.5 Incoming?
Leaks suggest Anthropic may be working on Sonnet 5.5, with faster responses, lower latency, and better price-to-performance.
Anthropic has not confirmed the model or a release date, so this remains a watch item.

Closing out
If you made it this far, thank you for reading, and we hope you found this week’s edition valuable.
If you enjoy reading, please forward our newsletter to someone you think would benefit from it.💙
Our promise to you: Every Wednesday, at 7 am EST, we’ll cut through the AI noise and send you human-curated AI content to make sure you stay ahead.
Content Pipeline
YouTube
Live Now: I Turned ChatGPT Into My Real-Time Voice Assistant (by Miles)
X (Twitter)
Live Now: How to Automate Your Life With Claude (Full Guide)
Free AI Asset Library
For all the AI prompts, cheatsheets, and PDF guides mentioned on the AI Edge YouTube channel, grab them by browsing our free assets library here:
See you next Wednesday!
Interested in sponsorship or a partnership? Get in touch at [email protected]