API
Prompt Caching
Understand how MaxLabs prompt caching reduces repeated input processing and effective usage.
Prompt Caching
When you work on the same repository, service, or group of components for an extended period, much of the model input remains unchanged between requests.
Instead of processing all of that unchanged context at full input cost every time, MaxLabs can reuse previously cached portions of the prompt.
Cache writes may cost approximately the same as normal input processing, but subsequent cache reads are substantially cheaper.
The longer a session remains cache-friendly, the more valuable this becomes.
How Prompt Caching Works
Prompt caching works best when consecutive requests share the same large prefix.
For example, an agent may already have:
- System instructions
- Tool definitions
- Repository instructions
- Relevant source files
- Architecture notes
- Conversation history
If the next request keeps that content unchanged and only appends new information, much of the existing prompt can potentially be reused from cache.
The important principle is:
Keep stable context stable, and append what changed.
Cache Lifetime
MaxLabs prompt cache remains warm for approximately one hour.
If cached context is reused within that period, eligible prompt content can be served through cache reads.
If a session remains inactive long enough for the cache to go cold, the relevant context may need to be processed and written again when work resumes.
Cache Pricing Is Automatic
You do not need to calculate cache discounts manually.
Usage is billed according to the token path actually used by the request.
If tokens are served through cache reads, the cache-read rate is applied automatically.
If new context needs to be processed or written into cache, that portion is billed accordingly.
Why Coding Workloads Cache Well
Software development is particularly suitable for prompt caching because a large amount of context often stays unchanged between consecutive requests.
Consider an agent working on a billing service. It may initially load repository instructions, package configuration, billing models, payment services, database schema, tests, coding conventions, and architecture notes.
Subsequent requests may add only small amounts of new information.
That is an ideal cache pattern.
Why High Cache Reuse Matters
With effective caching, the expensive part of each new request becomes primarily:
- New user instructions
- Newly requested files
- New tool output
- Newly introduced repository context
- Newly generated conversation content
The stable prefix can continue to be reused through cache reads.
This means a long-running development session can consume significantly less billable input than its apparent context size suggests.
Cache Performance in MaxLabs Apps
The MaxLabs Desktop application and CLI are heavily optimized for prompt caching.
In internal testing, long-running development sessions focused on the same microservice, module, or closely related components have achieved effective cache-hit rates of up to 99.9%.
That does not mean every session will achieve 99.9%.
Cache effectiveness depends on session continuity, how much context remains reusable, and how often prompt structure changes.
Cache Affinity
The Desktop application and CLI are designed to preserve cache affinity wherever possible.
This includes avoiding unnecessary changes to:
- Prompt ordering
- System instructions
- Repository context
- Tool definitions
- Conversation structure
- Serialized metadata
Prompt caches generally depend on matching existing prompt prefixes.
Even semantically identical content may fail to reuse a cache entry if it is reordered, reformatted, serialized differently, or routed inconsistently.
Using MaxLabs Through Third-Party Agents
Third-party agents can still benefit from MaxLabs caching, but cache-hit rates may be lower than in the MaxLabs Desktop application or CLI.
Third-party agents may:
- Reorder messages
- Rebuild prompts between requests
- Serialize tool output differently
- Insert changing metadata
- Modify system instructions
- Reconstruct conversation history
- Rewrite repository context
- Route related requests differently
Any of these can reduce prompt-prefix reuse.
API Session Affinity
API customers should provide a stable session identifier when several requests belong to the same logical conversation.
Use:
X-Session-Id: conversation-unique-id
For example:
X-Session-Id: project-billing-debug-42
The identifier should remain stable throughout one conversation and should not be reused across unrelated conversations.
Building Cache-Friendly API Requests
For API customers, the most important principle is:
Keep the beginning of the prompt identical whenever the underlying information has not changed.
Prefer Byte-Exact Stability
For best results, unchanged prompt content should remain byte-identical.
Differences in whitespace, ordering, serialization, metadata, and formatting can matter.
A useful rule is:
Semantically identical is not enough. Prefer byte-identical.
Use Append-Only Conversation History
Where practical, build conversations by appending new content instead of reconstructing previous content differently.
Append-only structures naturally preserve long stable prefixes.
Keep Static Content First
A cache-friendly prompt structure is:
1. Stable system instructions
2. Stable tool definitions
3. Stable repository instructions
4. Stable or slowly changing project context
5. Existing conversation history
6. New tool output
7. Latest user message
Stable first, dynamic later.
Use Deterministic Serialization
Avoid unnecessary changing values such as:
- Timestamps
- Request IDs
- Random UUIDs
- Non-deterministically ordered JSON
- Randomly ordered file lists
- Continuously regenerated summaries
- Changing environment metadata
For structured data, use stable key ordering, consistent whitespace and encoding, and deterministic ordering where practical.
Preserve File and Tool Ordering
Keep file lists and tool definitions in a stable order unless the underlying data actually changes.
Tool definitions can represent a large portion of the prompt, so keeping them stable can materially improve cache reuse.
Be Careful With Summarization
Prefer stable summaries, periodic compaction, append-only history between compactions, and deterministic summarization where practical.
MaxLabs' own clients handle this balance automatically.
Dynamic Content Is Normal
Not every token needs to be cacheable.
Tool results such as test output, shell commands, repository searches, compiler errors, and logs naturally change.
The goal is to prevent dynamic content from unnecessarily invalidating the stable context that came before it.
A Cache-Friendly API Architecture
Stable system prompt
↓
Stable tool definitions
↓
Stable project/repository context
↓
Append-only conversation
↓
New user message
↓
Deterministic serialization
↓
Stable X-Session-Id
↓
MaxLabs API
In short:
Byte-exact, append-only, deterministic prompt construction with cache-affine routing.
Caching and Usage Quotas
High cache reuse can make your usage quota go substantially further.
The benefit becomes more noticeable as:
- Repository context grows
- Sessions become longer
- Tool usage increases
- The same components remain relevant
- Prompt prefixes remain stable
When Cache Reuse Will Be Lower
Expect lower reuse when:
- Every request concerns a different repository
- Conversations are very short
- Session IDs change constantly
- Prompts are rebuilt from scratch
- System instructions frequently change
- Tool definitions change between requests
- Context is repeatedly reordered
- Summaries are aggressively regenerated
- Related requests lack session affinity
- Work jumps constantly between unrelated components
API Best Practices
For API integrations:
- Use one stable
X-Session-Idper logical conversation. - Keep unchanged system instructions byte-identical.
- Prefer append-only conversation history.
- Keep stable context before dynamic context.
- Serialize structured content deterministically.
- Keep tool definitions and ordering stable.
- Avoid unnecessary timestamps, UUIDs, and changing metadata in stable prefixes.
- Preserve repository formatting and deterministic file ordering.
- Avoid regenerating unchanged summaries.
- Continue active sessions within the cache's warm period when practical.
The underlying rule is simple:
Do not change prompt bytes unless the underlying information actually changed.
A Practical Rule
For MaxLabs caching, remember three things:
Keep the session stable.
Keep the prompt stable.
Append what changed.
The MaxLabs Desktop application and CLI handle most of this automatically.
If you are building directly on the API, following the same principles can materially improve cache reuse, reduce effective usage, and make long-running development sessions far more economical.
