# LLM inference retries > Why retrying an LLM inference call normally bills you twice, what a safe retry actually requires, and how to make the retry your SDK already performs idempotent. Source: https://docs.stg.perdurance.dev/llm-inference-retries/ --- **Retrying an LLM inference call is not safe by default.** A retried request is a second call to your provider: a second generation, a second bill, and — because sampling is stochastic — a different answer. The first call usually keeps running on the provider's side, so the retry does not replace it. It adds to it. This page is about making the retry safe: what breaks, what a correct retry needs from the server, and what changes in your code (nothing, if you are using a vendor SDK). ## Why a retried LLM request bills twice Every HTTP retry rests on an assumption that inference violates. `POST` is not idempotent, so a client that does not hear an answer cannot tell these two apart: | What the client saw | What actually happened | Cost of retrying | |---|---|---| | Connection reset at 40s | The provider never received it | Correct — retry is free | | Connection reset at 40s | The provider is 40s into a generation | A second full generation | | Read timeout | The answer was produced and lost in transit | A second full generation | | `502` from a proxy | Ambiguous | Possibly a second generation | Three of those four rows charge you again, and the client cannot see which row it is in. So it retries, and pays. On a long reasoning generation that is not a rounding error — it is the whole cost of the request, doubled, for a failure that had nothing to do with the model. You do not have to write a retry loop to hit this. The OpenAI and Anthropic SDKs both retry automatically by default — two attempts, with backoff, on connection errors and on `408`, `409`, `429` and `5xx`. If you have never set `max_retries=0`, your application is already retrying LLM inference, and already paying for it when it does. ## What a safe retry requires A retry is safe when the server can recognise the second call as *the same call* and attach it to work already in progress rather than starting more. That needs three things, and a provider API gives you none of them: 1. **An identity for the request** that the client can reproduce without having been told one. 2. **A record of the execution** that outlives the connection that started it. 3. **A way to read the result back** once the connection is gone for good. Perdurance supplies all three, and the first is the one that makes it work with clients that have never heard of it. ## How Perdurance makes the retry idempotent **A request is identified by the bytes of its body**, for ten minutes from submission. Send the same bytes again inside that window and you attach to the execution already running; the provider is called once, no matter how many times your client asks. ``` attempt 1 ──POST body────▶ hash(body) → new record ──▶ provider call starts ◀──connection dies── (still running) attempt 2 ──POST same body──▶ hash(body) → same record ──▶ no second provider call ◀──the stored answer, or a live tail of it── ``` Nothing in that requires the client to hold state. It does not need an idempotency key it must generate, store and re-send — the body it is already re-sending *is* the key. That is precisely what the SDK's built-in retry does, which is why an unmodified SDK becomes durable by pointing its `base_url` at Perdurance. ```python tab="Python (OpenAI)" from openai import OpenAI client = OpenAI( base_url="https://api.perdurance.dev/acme/prod/v1", api_key=PERDURANCE_KEY, max_retries=5, # now safe to raise: each retry attaches, it does not re-generate ) # If the connection drops, the SDK re-sends the identical body and lands on the # execution already in flight. One provider call, one bill. answer = client.chat.completions.create( model="gpt-4.1", messages=[{"role": "user", "content": "Explain idempotency."}], ) ``` ```python tab="Python (Anthropic)" from anthropic import Anthropic client = Anthropic( base_url="https://api.perdurance.dev/acme/prod/v1", api_key=PERDURANCE_KEY, max_retries=5, ) answer = client.messages.create( model="claude-sonnet-4-5", max_tokens=1024, messages=[{"role": "user", "content": "Explain idempotency."}], ) ``` ```typescript tab="TypeScript" import OpenAI from 'openai' const client = new OpenAI({ baseURL: 'https://api.perdurance.dev/acme/prod/v1', apiKey: process.env.PERDURANCE_KEY, maxRetries: 5, }) const answer = await client.chat.completions.create({ model: 'gpt-4.1', messages: [{ role: 'user', content: 'Explain idempotency.' }], }) ``` ## What a retry does, per state The second call's behaviour depends only on where the first one got to. | State of the record | Non-streaming retry | Streaming retry | |---|---|---| | `running` | Holds until it settles, then returns the stored answer | Replays the chunks stored so far, then live-tails to completion | | `succeeded` | Returns the stored answer | Replays the stored chunks | | `failed` | Returns the stored error | Replays, ending with `event: failed` | Note the streaming row: a retry mid-generation does not restart the generation and does not lose the part that already arrived. It replays what was stored and then joins the live tail. ## Retries you should still make, and ones you should not **Retry the transport.** A dropped connection, a read timeout, a `502` from something between you and us — re-send the identical body. That is free. **Do not retry by changing the body.** A retry that re-serialises the request, adds a timestamp, regenerates a UUID or reorders JSON keys is a *different* request by definition, and will be executed and billed as one. Keep the exact bytes. **Do not treat `504 held_too_long` as a failure.** A synchronous call is held for at most the deployment's **synchronous hold** (five minutes by default). When a generation outlasts that, the connection is released with a `504` — but the execution continues and its answer is stored. The message carries the request id. Re-send the identical body, or read it back by id. Submitting a *different* body here is the one mistake that reliably costs real money. **Do not retry `422`.** A dialect mismatch is a configuration problem, and no record is written. See [Routing](/routing). ## Retries you do not have to make at all Some retries are not yours. When a worker dies holding a request — a pod evicted, a node drained, a deploy rolling — its lease expires, the request is re-dispatched, and the execution picks back up without any client involvement. The record counts it: | Field | Means | |---|---| | `attempts` | How many times a worker took the request up | | `recoveries` | How many of those followed a worker dying mid-flight | | `next_retry_at` | When the next attempt is due, if one is | A non-zero `recoveries` is the durability machinery working, not a fault to alert on. ## The ten-minute window The body hash maps to the same request for **600 seconds** from submission. Inside it, identical bytes are the same request. Outside it, identical bytes are a new one and your provider is called again. That length is a balance: longer than a slow generation, shorter than a session. Long enough that a client which dropped and reconnected finds its own execution; short enough that the same question asked tomorrow is not silently answered from a recording. For anything longer-lived than the window, use the id. `GET /requests/{id}` works until your retention policy removes the record — see [Idempotency and resume](/idempotency). - [Agent SDK retries](https://docs.stg.perdurance.dev/agent-sdk-retries/): Why an agent loop multiplies this problem, and what to do about it. - [Idempotency and resume](https://docs.stg.perdurance.dev/idempotency/): The window, what counts as identical, and the request statuses. - [Unstable AI inference](https://docs.stg.perdurance.dev/unstable-ai-inference/): The failures that are not the model's fault, and what each costs. - [Getting started](https://docs.stg.perdurance.dev/getting-started/): Point an existing SDK at Perdurance in six steps.