lAItest / AI field guide · edition 2026-09-26
Know the language. Build with clarity.
Plain-English definitions for the concepts, patterns, protocols, and tools behind modern AI work.
414 defined terms · 16 topic areas · 50 essentials
Evals & observability terms
29 matching terms.
- Cost per successful task Concept · Evals & observability
Total execution cost divided by successful outcomes, including the cost of unsuccessful attempts.
- Eval Concept · Evals & observability
A systematic test of a model or application's behavior against defined criteria.
- LLM-as-a-judge Concept · Evals & observability
Use a model to assess outputs or execution trajectories against a rubric.
- Task success rate Concept · Evals & observability
The fraction of evaluated tasks meeting the specified completion criteria.
- Trace Concept · Evals & observability
A structured record of model calls, tool actions, timing, errors, and related execution events.
- Ablation Concept · Evals & observability
Remove or change a component to estimate its contribution to system performance.
- Benchmark contamination Concept · Evals & observability
Evaluation material leaks into training or optimization, weakening the validity of the evaluation.
- Correctness Concept · Evals & observability
Whether an answer or action satisfies the actual problem's requirements or truth conditions.
- Cost optimization Concept · Evals & observability
Reducing execution spend while preserving the required outcome quality, latency, reliability, and coverage.
- Deterministic grader Concept · Evals & observability
A code-based check with explicitly defined decision logic.
- Eval suite Concept · Evals & observability
A set of cases covering relevant capabilities, risks, and failure modes.
- Eval-driven development Concept · Evals & observability
Use systematic evaluation results to guide application and model-integration changes.
- Evaluation frameworks Concept · Evals & observability
Software or process that organizes evaluation cases, runners, graders, metrics, and reporting for model or application behavior.
- Faithfulness Concept · Evals & observability
Whether an answer remains supported by the evidence supplied to it.
- Golden dataset Concept · Evals & observability
A curated reference set with expected outcomes, properties, or grading criteria.
- Hallucination Concept · Evals & observability
Generated content that is unsupported or incorrect despite being presented as an answer.
- Judge calibration Concept · Evals & observability
Compare a model judge with trusted assessments and investigate disagreements or biases.
- Observability Concept · Evals & observability
The ability to infer system behavior from instrumentation such as logs, metrics, traces, and artifacts.
- Offline eval Concept · Evals & observability
Evaluate controlled examples outside live user traffic.
- Online eval Concept · Evals & observability
Assess behavior observed during production use.
- OpenTelemetry Concept · Evals & observability
A vendor-neutral framework and specifications for producing and exporting observability data.
- pass@k Concept · Evals & observability
The probability or estimated rate that at least one of k sampled attempts succeeds.
- pass^k Concept · Evals & observability
A consistency measure asking whether all k trials succeed under the evaluation's definition.
- Precision / recall Concept · Evals & observability
Measures of how many selected items are relevant and how many relevant items were found.
- Regression eval Concept · Evals & observability
Check whether a change harms previously acceptable behavior.
- Rubric Concept · Evals & observability
Explicit criteria defining how an output is evaluated.
- Span Concept · Evals & observability
A timed operation within a distributed or application trace.
- Trajectory Concept · Evals & observability
The sequence of actions and observations in an execution run.
- Trajectory evaluation Concept · Evals & observability
Assess the execution path, not only the final answer.
Browse by topic
- Agents & architecture 23 — Who decides, what executes, and where control lives.
- Goals, specs & plans 23 — Describe the outcome before delegating the implementation.
- Loops, critics & adversaries 33 — Understand how work is challenged, repaired, and stopped.
- Multi-agent coordination 21 — Delegation, ownership, context boundaries, and aggregation.
- Context & memory 29 — What the model sees now, and what persists for later.
- Tools & output contracts 20 — Model proposals become validated, authorized operations.
- Protocols & interoperability 19 — Name the boundary: tools, agents, editors, or interfaces.
- Coding-agent internals 25 — Instructions, skills, hooks, tools, and durable artifacts.
- Retrieval & knowledge 31 — Find evidence, rank it, and preserve source boundaries.
- Evals & observability 29 — Measure outcomes and inspect the execution path.
- Inference & performance 28 — Latency, throughput, compute, and memory are different constraints.
- Models, reasoning & training 36 — Separate weight changes from context and inference-time work.
- Security & reliability 27 — Make privileges explicit and side effects recoverable.
- Generative UI & voice 18 — The interaction layer has its own contracts and timing.
- Generative & physical AI 8 — Broader model families beyond text-based assistants.
- Tools & ecosystem 44 — Recognize the role before choosing the dependency.