BurbageLaw

Developments

OpenAI's prompt caching application claims the hash that routes a request to the server already holding its cache

2026-09-17Filings

United States Patent Application Publication No. 2026/0278026 A1, OpenAI OpCo, LLC, "Prompt Caching in Generative Response Engines," filed February 18, 2026, a continuation of Application No. 19/240,708, filed June 17, 2025, in a family that also identifies Application No. 19/077,920, filed March 12, 2025, published September 17, 2026

As published, claim 1 is a method in which a cloud computing service receives a request containing a prompt with a natural language task and an access key, generates a hash from a portion of the prompt, identifies the generative response engine that will handle the request on the basis of that hash, transmits the prompt to that engine, and receives a response reporting how many of the input tokens were already cached, from which a credit is determined. The specification describes computing the hash from a prefix of the prompt combined with a user identifier, so that repeated requests with the same opening are sent to the machine that already holds the computed attention state for those tokens, and describes a caching window with a minimum duration of 300 seconds. Other independent claims extend the scheme to multimodal inputs, to encoder tokens, and to a centralized cache within a data center. Both OpenAI OpCo, LLC and Anthropic, PBC announced prompt caching as a product feature before the earliest filing date in this family, Anthropic on August 14, 2024 and OpenAI on October 1, 2024, as reported by SiliconANGLE and in coverage of OpenAI's 2024 developer day. The trade site Patentlyze lists five OpenAI publications in September 2026, against a portfolio that stood at four documents when Justia Patents was checked on September 21, 2026.

What it changesWhat is claimed is not the idea of reusing a computed prefix, which was public from both companies months before the earliest filing date, but the routing mechanism around it: hashing the prefix to choose the server, and reporting cached token counts back for billing. That is a lesson in where to put the claim when the feature is already public, because the two 2024 announcements are prior art against the family under Section 102(a)(1) of Title 35 of the United States Code for everything they disclosed, and the hash-based routing is what they did not disclose. Any company running an inference service with prefix caching across more than one server should compare its request routing against these claims now rather than after grant, and companies buying inference from OpenAI should note that the billing credit for cached tokens is itself a claimed step. The wider signal is that OpenAI has started filing in volume on the serving infrastructure rather than on the models.

← All recent developments  ·  Full archive

This is general information about a published decision, agency guidance or piece of legislation. It is not legal advice and does not create an attorney-client relationship. Items concerning offices outside the United States are reported for information only and are not counsel on the law of those countries.