/

Metrics & KPIs

/

Time-to-First-Token (TTFT)

Time-to-First-Token (TTFT)

/ time-to-first-token-ttft /

Delay from prompt submission to the model’s first output token — the LLM’s share of response latency.

Delay from prompt submission to the model’s first output token — the LLM’s share of response latency.

Why it matters

TTFT isolates the model’s share of latency. Tracking it separately makes it obvious when a model change is what slowed everything down.

Related — Metrics & KPIs