Anthropic and Andrew Ng built an agent that uses 90% fewer tokens...

they dropped the entire book of Frankenstein into a prompt - 108,000 tokens asked one question
the input dropped from 108,000 tokens to 11
here's how:
step 1 → order matters: tools first, then system prompt, then your docs. anything that changes goes last
step 2 → drop the breakpoint after the last stable block. everything above it gets stored
step 3 → the match has to be byte-perfect. one extra space and you're back to full price
step 4 → 5 minutes of idle and it's gone. but every hit restarts the timer
step 5 → cache reads don't touch your rate limit. that's throughput you're not paying for
99% people never configure this - one afternoon and it pays for itself
save this - the full 1-hour course is right below ↓