Inference & performance · Concept

Paged attention

Manage attention KV-cache memory in blocks to improve allocation and sharing efficiency.

Related terms

TTFT · ITL · TPOT

Primary references

Reference definition