# Durable LLM inference > What Perdurance does to an LLM request: stores it, runs the call server-side, persists every chunk, and hands the result back by id whenever you return. Source: https://docs.perdurance.dev/ --- Perdurance sits between your code and your model provider. You send it the request body you would have sent the provider, on the route you would have sent it to, and it answers with the provider's own response — unchanged, no envelope, nothing your SDK has to be taught. What differs is what happens underneath. Perdurance writes the request down, runs the upstream call server-side, stores every chunk of the answer as it arrives, and records what the call cost in tokens. All of that survives your connection. It survives your process. ## What that buys - **A dropped connection loses nothing**: The upstream call keeps running when your client disappears. Send the identical body again and you attach to the execution already in flight instead of starting a second one. - **Retries do not double-bill**: Inside the idempotency window an identical body hashes to the same record. Your provider is called once, however many times your client asks. - **Every request is a record**: The body, the answer, each chunk, the usage counts and the execution history are stored and readable by id, for as long as you keep them. ## Two ways to call it Which one you want depends on whether you have something to hold a connection with. **Synchronously**, on your provider's own route. `POST /chat/completions` in OpenAI's dialect, `POST /messages` in Anthropic's. The connection is held and you get the provider's answer on it — a completed response, or a stream of the provider's own events when the body says `"stream": true`. An unmodified vendor SDK works against this, and its own retry is what makes it durable. **Fire-and-forget**, on `POST /requests`. You get `202` and a request id straight back, and the execution runs without you. This is the shape a batch job wants, or anything that cannot hold a socket open for a long generation. Either way the record is the same, and either way you can come back for it later with `GET /requests/{id}` — whole, or as a stream replayed from the beginning and resumable from wherever you stopped reading. ## The problems it is for Three failures cost you a generation you already paid for, and none of them are the model's fault. Each has a page. - [Unstable AI inference](https://docs.perdurance.dev/unstable-ai-inference/): The six ways an inference call fails without the model failing — dropped sockets, serverless limits, evicted workers — and what each one costs. - [LLM inference retries](https://docs.perdurance.dev/llm-inference-retries/): Why retrying an LLM call bills you twice by default, and what makes a retry safe. Your SDK is already retrying whether or not you wrote the loop. - [Agent SDK retries](https://docs.perdurance.dev/agent-sdk-retries/): Why an agent loop multiplies the problem across three layers of retry, and how to wire the frameworks up so it stops. [Getting started](/getting-started) takes you from an empty account to a stored request you can fetch back. If you are operating a deployment rather than calling one, [Deployment](/deployment) is the shorter road, and the [FAQ](/faq) answers what people ask before either.