lAItest / AI field guide · edition 2026-09-26
Know the language. Build with clarity.
Plain-English definitions for the concepts, patterns, protocols, and tools behind modern AI work.
414 defined terms · 16 topic areas · 50 essentials
Inference & performance terms
28 matching terms.
- KV cache Concept · Inference & performance
Stored attention keys and values reused to avoid recomputing prior-token attention state.
- Model routing Concept · Inference & performance
Choose a model based on task needs, constraints, policy, or observed difficulty.
- Prompt caching Concept · Inference & performance
A provider's mechanism for reusing eligible prompt processing; exact behavior and billing are platform-specific.
- TTFT Concept · Inference & performance
Time to first token: delay from a defined request start until the first generated token arrives.
- Batch inference Concept · Inference & performance
Process a collection of requests as an offline or scheduled workload.
- Continuous batching Concept · Inference & performance
Dynamically batch requests as they enter and leave inference execution.
- CPU/GPU offloading Concept · Inference & performance
Move model state or computation between CPU and GPU resources.
- Decode Concept · Inference & performance
Generate subsequent output tokens using the model and available state.
- End-to-end latency Concept · Inference & performance
Total request or task duration, including queues, model calls, tools, retrieval, and retries.
- Expert parallelism Concept · Inference & performance
Distribute mixture-of-experts components across devices.
- Fast path / slow path Pattern · Inference & performance
Separate low-latency handling from more expensive reasoning or processing.
- FlashAttention Concept · Inference & performance
An IO-aware exact attention algorithm designed to reduce memory traffic and improve execution efficiency.
- ITL Concept · Inference & performance
Inter-token latency: elapsed time between consecutive emitted tokens.
- Model cascade Pattern · Inference & performance
Try a less expensive path first and escalate when specified criteria require it.
- Paged attention Concept · Inference & performance
Manage attention KV-cache memory in blocks to improve allocation and sharing efficiency.
- Prefill Concept · Inference & performance
Process input tokens to establish the state used for subsequent generation.
- Prefill-decode disaggregation Concept · Inference & performance
Manage input processing and token generation on separately allocated serving resources.
- Prefix caching Concept · Inference & performance
Reuse computation for matching input prefixes.
- Quantization Concept · Inference & performance
Represent model values with reduced precision to lower memory or compute requirements.
- Rate limit Concept · Inference & performance
A restriction on request, token, concurrency, or resource consumption over a defined interval.
- Reasoning budget Concept · Inference & performance
A configured limit or allocation for inference-time reasoning effort.
- Semantic cache Concept · Inference & performance
Reuse a previous result for a sufficiently similar request, subject to validity and permission checks.
- Speculative decoding Concept · Inference & performance
Use a cheaper draft mechanism to propose tokens that a target model verifies.
- Speculative execution Concept · Inference & performance
Start potentially useful application work before knowing whether it will be needed.
- Tail latency Concept · Inference & performance
The slow end of a latency distribution, commonly summarized by p95 or p99.
- Tensor parallelism Concept · Inference & performance
Split operations within model layers across multiple devices.
- TPOT Concept · Inference & performance
Time per output token: an average token-generation latency under a specified measurement method.
- TPS Concept · Inference & performance
Tokens per second; specify whether it measures one request or aggregate system throughput.
Browse by topic
- Agents & architecture 23 — Who decides, what executes, and where control lives.
- Goals, specs & plans 23 — Describe the outcome before delegating the implementation.
- Loops, critics & adversaries 33 — Understand how work is challenged, repaired, and stopped.
- Multi-agent coordination 21 — Delegation, ownership, context boundaries, and aggregation.
- Context & memory 29 — What the model sees now, and what persists for later.
- Tools & output contracts 20 — Model proposals become validated, authorized operations.
- Protocols & interoperability 19 — Name the boundary: tools, agents, editors, or interfaces.
- Coding-agent internals 25 — Instructions, skills, hooks, tools, and durable artifacts.
- Retrieval & knowledge 31 — Find evidence, rank it, and preserve source boundaries.
- Evals & observability 29 — Measure outcomes and inspect the execution path.
- Inference & performance 28 — Latency, throughput, compute, and memory are different constraints.
- Models, reasoning & training 36 — Separate weight changes from context and inference-time work.
- Security & reliability 27 — Make privileges explicit and side effects recoverable.
- Generative UI & voice 18 — The interaction layer has its own contracts and timing.
- Generative & physical AI 8 — Broader model families beyond text-based assistants.
- Tools & ecosystem 44 — Recognize the role before choosing the dependency.