Skip to content

latenzy

Per-model LLM latency monitoring for enterprises.

latenzy is a synthetic prober and Prometheus exporter, with prebuilt Grafana dashboards, that measures what lab-level status pages can't: the latency your account gets from each modelclaude-sonnet-4-6 vs gpt-4o vs gemini-2.0-flash, not "Anthropic is up".

Latency is tenant-specific: it depends on your rate-limit tier, your region, and the path you take to the model (direct API, Bedrock, Vertex). latenzy runs inside your network on your keys and exports per-model metrics your existing Prometheus + Grafana stack can alert on.

DOI PyPI

What it measures

Every probe cycle, for each (source, provider, model, endpoint, prompt_class):

Metric Meaning
latenzy_ttft_seconds time to first streamed token (histogram)
latenzy_request_duration_seconds total request duration (histogram)
latenzy_output_tokens_per_second streaming throughput (histogram)
latenzy_probes_total{outcome=...} count by ok / rate_limited / timeout / error
latenzy_last_success_timestamp_seconds staleness signal for alerting

The source label is synthetic for the prober's canaries and live for real application traffic (Live traffic) — one dashboard shows both.

Where to go next

License: AGPL-3.0-only. Dual licensing available for enterprises.

— amitpatole