How to slash token costs with context caching in agent harnesses

Coding agents do heavy engineering work. The agent harness manages runtime state, sandbox execution, and prompt preparation. Providing static surrounding context on every turn creates a multi-turn scaling bottleneck, wastes tokens, and increases cost.
There are a few ways to prevent this. One way is through Gemini Enterprise Agent Platform Context Caching, which helps eliminate this waste. By keeping the static codebase in server cache and sending only the dynamic updates on subsequent turns, the harness cuts transmitted input volume.
by Balaji Subramaniam, DevRel Engineer, Google Cloud
Here are the reasons why naive harness implementations incur unnecessary costs. Then we’ll cover how to avoid these mistakes by building a production caching pipeline with Google ADK 2.0 and the Gemini Enterprise Agent Platform.
Cost Drivers of Naive Harness Implementations
The Token Accumulation Problem
Large language models are stateless. They process an input token sequence and return generated text. The agent harness maintains conversational memory and serializes the full context window on each model invocation.
Consider the arithmetic of a typical multi-turn coding loop. If your static reference code contains 37,500 tokens and each turn appends a 300-token error traceback, a 5-turn run transmits 37,800 tokens on turn one, 37,800 on turn two, and so on. By the fifth turn, you have billed 189,000 prompt tokens. A ten-turn refactoring task against a large codebase consumes nearly 400,000 prompt tokens. The network spends time re-uploading identical files, and the server spends compute re-tokenizing them.
The Prefix-Breaking Trap in Harness Design
Context caching requires an exact, byte-for-byte token match starting from token zero of the prompt. If the harness places any mutable runtime metadata before or inside the static text, the prompt hash changes. The server cannot match the request to the pre-computed cache. It drops into a cache miss and the full prompt is charged at standard rates.
Dynamic variables belong at the end of the prompt, never at the beginning.
Behind the scenes: How server-side context caching works
Let’s take a look at how caching works. The sequence diagram shows the context caching lifecycle within Agent Platform.
After the context is cached, the harness passes only the dynamic suffix (such as a unit test traceback or compiler error) along with the cache reference.
Understanding the Dual Savings Model
Context caching delivers two distinct types of savings: bandwidth and cost savings. Caching cuts physical network data transmission and under Agent Platform, cached content reads are billed at a 75% discount, meaning you pay 0.25x the standard prompt rate for cached tokens.
How to implement a successful caching pipeline
A successful context caching architecture relies on deterministic payload structure and centralized lifecycle management. Below is the implementation breakdown of core components to build a caching pipeline.
Enforcing Prefix Invariance in the Harness
To prevent prefix corruption, the harness uses CachePayloadBuilder to strictly isolate immutable prompt headers from dynamic execution suffixes. Here are the key methods from CachePayloadBuilder:
Implementing the Cache Manager in Your Harness
The ContextCacheManager handles the lifecycle of cached resources on Agent Platform. It computes content hashes to reuse existing active caches, extends time-to-live (TTL) settings when needed, and dispatches cached inference requests.
Here are the key lifecycle operations from ContextCacheManager:
End-to-End Walkthrough: Multi-Target Modernization
To analyze caching in a realistic scenario, we constructed a five-module batch modernization workload. The agent harness upgrades five Python 2.7 modules, which together are around 37k tokens.
See context-caching/multi_target_suite.py for the complete target definitions and sandbox test harness.
Measured Results Across Three Topologies
We executed empirical analyses across three distinct multi-agent topologies in Python 3.11 sandbox environments on Google Cloud Run.
Scenario 1: Multi-Target Batch Modernization (5 Turns)
Modernizing five independent legacy Python files against a 37,659-token monorepo prefix.
Scenario 2: Adversarial SQLi Red/Blue Debate (4 Turns)
A four-turn debate between a Red Team exploiter agent and a Blue Team fixer agent evaluating database security against a 38,367-token OWASP specification.
Scenario 3: Multi-File Dependency Graph Modernization (4 Turns)
Cascading refactoring across four interdependent microservice layers against a 38,441-token Object-Relational Mapping (ORM) database SDK.
Decision Tree for Context Caching
Context caching provides massive structural advantages, but platform engineers must apply it selectively based on workflow characteristics.
Follow these operational rules:
Running the Reproducible Analyses
You can run these empirical analyses locally using the test suite and CLI runner:
Get started on your own
Context caching can turn multi-turn agent loops from an expensive token drain into a predictable and scalable architecture. The rules are simple, if you pass large static codebases across three or more turns, isolate your dynamic context to suffixes, pin a server-side cache and let the model reuse it.
Use the code and test suites in this repository to get started.










