Hey everyone, welcome to the fifty-fourth issue of The Main Thread. This issue also marks the beginning of an 8-part series on inference engineering.
Inference engineering is the work of using a trained model to handle requests in a running product. It means deciding how to use available computing power and memory when multiple requests arrive at once. These decisions determine how long a user waits for an answer, how many requests our system can handle, and how much an answer costs.
Let’s take an example of an e-commerce site that sells laptops. It adds an AI shopping assistant to help customers decide what to buy. A customer asks, “Which laptop should I buy for college under ₹60,000?” The assistant compares the products and explains its recommendation.
In this example, the model runs on one server with one GPU. Our website sends the customer’s question and relevant product details to that server. The answer streams in the shopping assistant as the model writes it.
Now, three customers ask for recommendations at once. One wants a laptop for college, another has a long conversation comparing several laptops, and the third one wants one for video editing. All three requests need the same GPU, but the long conversation and its product details take more work to process. Writing each reply requires repeated GPU work, and unfinished replies also keep calculations in GPU memory for later reuse. The serving software must decide how to share that time and memory. Processing the long conversation could delay the next words in the other two replies.
We will keep our example small enough to follow: one loaded text model, one GPU, and 3 requests. We will trace what each request needs, where it waits, and what must happen before the user receives a complete answer.
An earlier issue, “What Happens in the 200ms After You Hit Enter on Your LLM?”, followed the computation inside one request. Here, we follow the requests sharing that computation.
The Model is Loaded Before the Request Arrives
Before a customer asks anything, our GPU worker has already loaded the model. What did it load?
The model runs layers of mathematical operations. Its weights are a large array of numbers used in those operations: the model multiplies input numbers by these weights to calculate what comes next. During model training, these weights are adjusted over many examples so the model gets better at predicting the next token. When a customer asks a question, the server uses the learned weights to generate answers. Running a trained model this way is inference.
Our model is a transformer language model, which generates an answer one token at a time. When a customer asks a question, a tokeniser splits the question and product details into tokens and assigns each token a number (token_id) for the model to process. A token might be a word, part of a word, or punctuation.
In our example, all three customers use the same model weights. If we load them separately for each question, every answer would wait while the server read those weights into GPU memory before the model even touched its input. We load them when a worker starts and keep them there while it handles the request. The GPU runs many of the model’s numerical calculations in parallel.
When a customer asks a question, our website sends it to the server running the model. If 3 customers ask at once, the server receives all 3 questions and decides when the GPU works on each one.
Why An Answer Keeps Using the GPU
A customer asks which laptop to buy for college. The system sends that question, details about 2 laptops, and any earlier messages to the model server. This is what the preparation looks like with a Hugging Face tokeniser:
messages = [
{
"role": "system",
"content": "Recommend a laptop from this catalog:\nA: ...\nB: ...",
},
{
"role": "user",
"content": "Which laptop should I buy for college?",
},
]
formatted_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
token_ids = tokenizer(
formatted_text,
add_special_tokens=False
).input_idsformatted_text contains special markers that separate the instructions, the customer's message, and the place where the model should begin its reply. The exact markers depend on the model. The final line converts that entire sequence into token IDs, which are the input to the model. Earlier messages would appear as more entries in messages, so a short final question can still produce a long list of token IDs.
Prefill: read the whole input
Before the model can answer, it must process every input token: the instructions, laptop details, and conversation. All these tokens are known before the work starts, so the GPU can pass many token positions through each model layer together. The model still moves through its layers in order, but it does not have to wait for another input token to
become available.
By the end of prefill, the model has processed the complete input and predicted the first token of the answer. The worker has also saved reusable results for the input tokens, so the model does not repeat all that work as the answer grows. A longer input has more tokens to process before this first prediction, so the customer waits longer for the answer to begin.
Decode: write the answer one token at a time
The first answer token now becomes the model's next input. The worker passes it through the model's layers, reuses the calculations saved during prefill, and selects the second answer token. It then saves the new calculations and repeats the same process for the third token.
The worker cannot process all the answer tokens together because each one depends on the tokens selected before it. This repeated phase is decode. If the answer contains 10 tokens, prefill selects the first, and the worker needs 9 decode steps to produce the rest.

Prefill and decode phase
Once one request’s answer starts appearing in the browser, it still needs GPU time for the later tokens. If another request arrives with a long conversation, its input needs GPU time too. The first request’s saved calculations also occupy GPU memory while its answer continues.
Why An Unfinished Answer Occupies GPU Memory
What the model saves
To predict the next token, the model combines information from the input and the answer written so far. It calculates which earlier tokens matter for the current token and how much information to take from each one. This mechanism is called attention.
The model writes text from left to right. When processing a token, it can use that token and the tokens before it. It is blocked from using tokens that come after it, even during prefill when the full input is available. This rule is called causal attention, and it prevents a prediction from using future text. See the original transformer paper for this.
At every attention layer, the current token produces a query, a list of numbers used to find relevant earlier tokens. Each processed token has a key, used to match against that query, and a value, which contains the information it can contribute. The model compares the query with the keys, then combines more information from the values whose keys match more closely.
The keys and values for an earlier token remain useful as the answer grows because later tokens do not change the results already calculated for it. The worker keeps them in the key/value cache, or KV cache, instead of calculating them again during every decode step. The KV cache is numerical attention state stored for reuse.
Prefill creates KV-cache entries for every input token. Each decode step adds an entry for the new answer token. A request with a longer input starts with a larger cache, and its cache continues to grow until the answer ends.
Why this limits concurrent requests
Our GPU memory now contains three kinds of data:
Data in GPU memory | Who uses it | How long it stays |
|---|---|---|
Model weights | All requests share them | While the worker serves this model |
KV cache | Each active request has its own | Until that request ends or its state is removed |
Working memory | The model execution running now | While that work is being performed |
This is why fitting the model weights into memory is only the first check. The GPU also needs room for the KV caches of all requests, plus the memory needed to run the next model step. More active requests keep more caches alive, and longer requests need more space. The PagedAttention paper explains how these growing allocations affect serving capacity.
If one request is paused, it stops using GPU computation for the moment, but its KV cache remains in memory so generation can resume from the same point. Removing that cache frees memory, but the worker would have to rebuild or restore it before A could continue. The scheduler therefore manages both GPU time and GPU memory.
Memory capacity is the amount of usable GPU memory. Comparing it with the memory required by the model weights, active requests’ KV caches, and temporary execution buffers tells us whether the workload fits.
Memory bandwidth tells us how quickly the GPU can read that data while running the model. A request can fit and still generate slowly because every decode step must read model weights and attention state.
Accepting the Requests Creates An Obligation
Imagine our 3 requests arrive close together with their hypothetical counts:
Request | customer's question and supporting information | Input tokens | Maximum output tokens |
|---|---|---|---|
A | College laptop question and 2 short product descriptions | 200 | 128 |
B | Detailed comparison with product details and chat history | 12,000 | 256 |
C | Video-editing laptop question and 2 product descriptions | 300 | 128 |
Input counts include the complete formatted prompt and supporting information. We begin after the website has selected relevant products from its catalogue; the model receives their details with the question. Output limits are ceilings; the model can finish earlier.
Before GPU execution, the service validates the request and decides whether to accept it. Validation checks that the
requested model exists, the input is valid, and the prompt plus allowed generation fit the supported context limits. The context limit bounds the sequence the model can handle. The output limit bounds how much new text this request may generate. An output limit alone doesn't show how much GPU memory the request will need: B already brings 12,000 input tokens.
An answer can end in three ways:
The model produces a special token meaning it is done
The service finds configured stop text
The request reaches its output limit.
The response should say which one happened because an output limit can cut off an unfinished answer.
After validation, the service decides whether it has enough GPU time and memory to accept the request. This decision
is admission control. Request B can be valid and still be queued or rejected while requests A and C are active.
When the service is full, it can reject new requests or place them in a queue with a maximum length and wait time. If
requests arrive faster than the GPU completes them, an unbounded queue keeps growing. Customers time out and retry,
which adds more work to the overloaded service. Queue length alone does not show how much work is waiting. B has 12,000 input tokens, while A has 200, so processing
B's input requires far more GPU work.
The service knows the input length after tokenisation. It learns the final output length only when generation ends;
the output limit is the maximum. Reserving that maximum for every request can waste memory when answers end early.
Letting each request take more memory as its answer grows requires a policy for what happens when no space remains.
Queue time counts toward the request's deadline. Before starting queued work, the service should check that time
remains and the customer still wants the answer. Otherwise, the GPU may generate a response that nobody wants to receive.
Where a Request Runs, and Where
Before a request can run, the service needs a worker that has finished loading the model. The API may already be
reachable while the weights are still loading, so the readiness check must confirm that the worker can execute model work.
First choose the worker
Our example has one GPU worker, so every request goes there. A larger service may have several workers, each running
its own copy of the model. Choosing one of them for a new request is called routing.
After a request starts, its KV cache lives on the chosen worker. Its later decode steps usually run on that worker too. Moving this request elsewhere would require transferring the cache or rebuilding it from its tokens. Some systems deliberately use one set of workers for prefill and another for decode. They must transfer the KV cache
between those workers.
Then choose the next GPU work
Several requests can wait at the same worker. The scheduler decides what the GPU runs next: Request A's next decode step,
request B's prefill, or work for request C. It makes this decision repeatedly because each unfinished answer needs more GPU steps.

The scheduler runs work from several requests together in one GPU operation. We call that group a batch. Each request still keeps its own tokens, KV cache, and stopping condition. The group can change between model steps. When one request finishes, it leaves; a waiting request can then join. This is called continuous batching.
A larger batch can use more of the GPU's parallel hardware because it runs calculations for more requests at once. It
also keeps more request state in memory and makes more requests compete for each step. The useful batch size depends
on the workload and the latency customers can tolerate.
How generated text reaches the customer
From token IDs to visible text
The GPU produces token IDs, which are numbers. The tokeniser converts those IDs back into text, and the server sends that text to the browser while the model continues generating the answer.
The text travels in chunks. A chunk is one piece of data sent over the network, and it may contain text from one
token or several. A server between the model and the customer's browser, often called a proxy, may collect several
chunks before forwarding them. This is response buffering.
Suppose request A's model steps run at a steady pace, but the proxy waits for several chunks before sending them. Request A sees the answer arrive in bursts. A faster GPU would not remove those pauses because they happen after the model generates the text.
Time to first token (TTFT)
This is the time from the browser sending the request to the first generated text reaching the browser. It measures the wait the customer actually feels. To investigate a high TTFT, we also record when the inference engine produced the first token. If the token was ready quickly but reached the browser late, the delay happened after inference, during text conversion, proxy buffering, or network delivery.
Time to first byte (TTFB)
It is different from TTFT. TTFB ends when the browser receives the first byte of the HTTP response, which may belong to a header rather than the generated answer. A low TTFB tells us that the response started quickly. It does not tell us that the customer saw the answer quickly.
The request must end and release its state
A request can end because the answer finished, the customer closed the page, or the model worker failed.
If a user closes the page, that cancellation must reach the scheduler. A GPU step already in progress may finish, but the
scheduler should stop choosing that cancelled request for later steps and release its KV cache. Otherwise, the service keeps spending GPU time and memory on an answer nobody will read.
If the worker fails after a request has received part of the answer, that partial text remains in the browser. Retrying starts a new generation, which may produce different text. The application should mark the first answer as interrupted and decide whether to keep or replace it.
Once a request finishes or is cancelled, it must stop reserving GPU time and memory for active generation.
When Many Requests Need the Same GPU
Let’s return to our initial example where three requests are made to our system. The first request’s reply has started streaming. The second request is still waiting because its 12,000-token prompt needs prefill. The scheduler has also started the third request, whose 300-token prompt is much shorter, so its reply is streaming too.
The timeline below shows how the service reaches its next scheduling decision:

The vertical direction shows the order of events, not how long each one takes. At the final line, the GPU becomes
available again, and the scheduler must choose between B's prefill and the next decode steps for A and C.
If the scheduler runs the decode steps for A and C, their replies keep moving, but B's TTFT keeps growing. If it gives the GPU to B's entire prefill, B moves closer to its first token, but A and C may receive no new text during that work. Their inter-token latency grows instead.
Some serving engines can divide B's prefill into smaller pieces and run them alongside decode steps for A and C. This
can reduce the pause for A and C without leaving B untouched. The scheduler still has to decide how much GPU time each kind of work receives. We will cover this in the next part of this series.
GPU memory limits these choices as well. A and C already have KV caches, and those caches grow as their replies gain
tokens. Starting B creates a much larger KV cache for its 12,000-token prompt, followed by more state as B generates
its answer.
The model may support B's prompt when B runs alone, yet the KV caches for A, B, and C may not fit at the same time. If
the engine accepts more active work than GPU memory can hold, it may pause one request, discard some of its saved state, and rebuild that state later. This is called preemption, and rebuilding adds more GPU work and delay.
One average latency number would hide the trade-off. B is waiting for its first token, while A and C are waiting for
their next tokens. The scheduler changes one wait whenever it chooses the next GPU work.
Measure the waits the customer can see
A streaming reply can feel slow in two different ways.
It may take too long to begin
It may begin quickly and then keep pausing.
We need a separate measurement for each problem.
1. Time to first token
Time to first token (TTFT) runs from the browser sending the request to the first generated text reaching the browser. It includes network travel, request preparation, queueing, prefill, and delivery of the first output. Response headers and empty stream events don’t count because the customer hasn’t received any answer text.
In our example, request B's TTFT grows while its long prompt waits for prefill.
2. Inter-token latency
Inter-token latency (ITL) is the time between one generated token and the next. It reveals pauses after a reply has
started. If the scheduler gives the GPU to request B's prefill for too long, the ITL for A and C increases.
The browser usually receives chunks rather than individual tokens, so it directly observes inter-chunk latency: the
time between one text chunk and the next. A chunk may contain several tokens, which means inter-chunk latency and ITL are not always equal. The serving engine can record token timing, while the browser records the experience the customer actually sees.
3. Goodput
Throughput counts how many requests the service completes per second. That number includes requests whose replies
started late or repeatedly stalled. Goodput counts only successful requests that also met the latency targets we set
for TTFT and ITL.
Suppose 120 requests finish during a 60-second test. Throughput is 2 requests per second. If only 90 of them meet a 1-second TTFT target and a 100-millisecond ITL target, goodput is 1.5 requests per second. The targets are examples,
not recommendations. Rejected, failed, and timed-out requests do not count as goodput.
Goodput exposes the trade-off in our timeline. Serving request A and C may protect their ITL while B misses its TTFT target. Running request B's prefill may improve its TTFT while causing A or C to miss an ITL target. The better schedule is the one that keeps more requests within the targets chosen for this application.
These three measurements tell us what the customer experienced. Internal measurements help explain why: queue time, prefill time, waits between decode steps, active request count, and KV-cache use.
When comparing serving settings, keep the model, inputs, output limits, generation settings, and request arrival pattern fixed. Shortening request B's prompt or limiting A's answer reduces the work, but it also changes what the app can ask the model to do.
For requests A, B, and C, record the input length, queue time, prefill time, first delivered output, later output gaps, and final status. Then return to the last line of the diagram and ask which wait the scheduler should reduce next.
The service decides who waits
A reachable model becomes a dependable service only when its behaviour under contention is deliberate. We must decide
how much work can enter, which request runs next, how active requests share GPU memory, how generated text reaches the
browser, and when abandoned work releases its resources.
The model calculates what token could come next. The serving system decides when that calculation runs and when the customer sees its result. This is why a model can answer one request perfectly and still produce a poor experience under load. The delay may come from the queue, prefill, decode scheduling, KV-cache pressure, or the network stream. Each one needs a different fix.
Our timeline ends at the first difficult choice. Request B needs a long prefill, while request A and C need their next decode steps. In the next issue, we will place the possible schedules side by side and see exactly how processing B's prompt can interrupt the replies already streaming to A and C.
Namaste!
Subscribe to The Main Thread to receive the remaining 7 issues in the inference engineering series. In the next issue, we will take the unresolved scheduling decision above, draw both timelines, and work out which customer waits under each one.



