Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

Prompt Caching: How It Reduces LLM Cost and Latency

Prompt caching reuses identical prompt prefixes to reduce repeated processing. See how cache hits work, how to structure prompts, and how to verify reuse in the OpenAI API.

Table of Contents

Prompt caching lets a large language model API reuse work already performed for the unchanged beginning of a prompt. It is most useful when an application repeatedly sends long system instructions, examples, tool definitions, or reference material. A matching request can reach its first output token faster and may be billed at a lower cached-input rate.

Prompt caching does not reuse the model's final answer. The variable part of the request is still processed, and the model still generates a new response. What is reused is the computation associated with an eligible, identical prompt prefix.

 

How prompt caching works

LLM inference starts by processing the input prompt before generating output tokens. Long prompts make this initial processing stage more expensive. If several requests begin with exactly the same content, a provider can retain reusable intermediate data for that shared prefix instead of processing it from scratch every time.

A request can therefore produce one of two outcomes:

  • Cache hit: an eligible matching prefix is available and can be reused.
  • Cache miss: no matching entry is available, so the full prompt must be processed. An eligible prefix may then be cached for a later request.

This is different from an application-level response cache. A response cache returns a previously stored answer for the same query; prompt caching only reduces repeated model computation and still allows a different answer to be generated.

The exact-prefix rule

The most important requirement is an identical prefix. Put stable content first and changing content last.

Stable system instructions
Shared examples and reference material
User question: What should I cook for dinner?
Stable system instructions
Shared examples and reference material
User question: What should I cook for lunch?

The two requests share the same opening content, so that section may be reusable. By contrast, placing a timestamp, request ID, user-specific value, or changing question near the beginning alters the prefix and can prevent a cache hit.

For multi-turn chats, append new messages while keeping earlier messages unchanged. Reordering or rewriting previous conversation content changes the prefix. Tool definitions, tool order, images, files, and structured-output schemas may also contribute to cache matching, so keep them stable when reuse matters.

What content is worth caching?

Prompt caching provides the greatest benefit when the reusable prefix is both long and frequently repeated. Common candidates include:

  • System and developer instructions used for every request.
  • Few-shot examples that define the desired response format.
  • Large reference passages used by a RAG application.
  • Tool definitions and schemas used by an AI agent.
  • Stable conversation history in a long-running chat.

Short, one-off prompts are unlikely to benefit. If you are comparing providers or model families for an application, the guide to leading large language models provides broader selection criteria. Developers planning the rest of an LLM stack may also find these Python libraries for building LLM applications useful.

 

Prompt caching in the OpenAI API

OpenAI enables prompt caching for eligible requests on recent models. The exact rules depend on the model family. Current OpenAI documentation states that GPT-5.6 and later models require a cacheable prefix of at least 1,024 tokens, while earlier supported models may require between 1,024 and 2,048 tokens. Applications should check the requirements for the model they actually use instead of assuming one threshold or retention period applies to every model.

For GPT-5.6 and later models, developers can use cache breakpoints to mark the end of reusable content and a consistent prompt_cache_key to improve matching. Earlier supported models use automatic, best-effort reuse of matching prefixes. Because these details can change, consult the official OpenAI prompt caching guide before relying on a specific model, lifetime, or pricing rule.

How to verify a cache hit

Do not infer caching from response speed alone. Inspect the usage details returned by the API. For the Responses API, cache statistics appear under usage.input_tokens_details. The key fields are:

  • cached_tokens: input tokens read from a cache entry.
  • cache_write_tokens: input tokens written to cache on model families that report cache writes.
{
  "usage": {
    "input_tokens": 2600,
    "input_tokens_details": {
      "cached_tokens": 2000,
      "cache_write_tokens": 400
    }
  }
}

A practical test uses two requests with the same long prefix and different text only at the end. The first request generally creates or warms the eligible cache entry. The later request can reuse it if the prefix, model settings, routing key, and cache lifetime all meet the provider's requirements.

What is Prompt Caching? How to Reduce Costs and Speed Up LLM Picture 1

The second request should keep the reusable instructions byte-for-byte stable and change only the final user question.

What is Prompt Caching? How to Reduce Costs and Speed Up LLM Picture 2

The screenshots illustrate the kind of usage comparison to make, but the exact token counts, prices, and model behavior depend on the current API configuration.

 

Why cached tokens may remain at zero

  • The shared prefix is too short. The overall request may be long enough while the identical opening section is not.
  • Dynamic content appears too early. Timestamps, user IDs, request IDs, or changing context before the reusable section alter the prefix.
  • Tools or schemas changed. A small difference in a tool description, parameter schema, or ordering can break the match.
  • The cache entry expired or was not available. Retention and routing behavior vary by model and provider.
  • The wrong usage field is being checked. Read the cache details in the response's token-usage object rather than estimating from total input tokens.

Practical optimization checklist

  1. Measure how much of each prompt is genuinely repeated.
  2. Move stable instructions, examples, tools, and reference content to the beginning.
  3. Place user input and other changing values at the end.
  4. Keep the reusable prefix and its ordering identical across requests.
  5. Use the model's supported cache key, breakpoint, and retention options.
  6. Track cached_tokens, latency, and actual cost instead of assuming every repeated prompt produces a hit.

Prompt caching is a targeted optimization, not a substitute for shorter prompts or better context management. It is most valuable when an application repeatedly sends a substantial, stable prefix and can verify that later requests are actually reading from the cache.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.