Measure cache use
See fresh input, cache writes, cache reads, and output tokens.
LLM PROMPT-CACHE ANALYTICS
CacheEconomics analyzes LiteLLM usage metadata, shows where repeated input misses the cache, and recommends the first fix to test.
Keep stable instructions together so more requests can reuse them.
BASED ON REQUEST STRUCTUREWHAT CACHEECONOMICS SHOWS
System prompts, tool definitions, documents, and conversation history often repeat. See how much is reused and what still pays the standard input rate.
See fresh input, cache writes, cache reads, and output tokens.
Spot repeated input that expires or changes before it can be reused.
Apply your pricing to measured traffic and see how each estimate was calculated.
Track request volume, latency, errors, analysis jobs, and ingestion lag.
SAVINGS CALCULATOR
Use three numbers from your LLM bill and traffic pattern.
DASHBOARD TOUR
These short tours use example data, so you can see the workflow without exposing customer traffic.
HOW IT WORKS
The collector records token counts, timing, outcomes, and private identifiers for repeated prompt sections.
Every event is checked before storage. Prompt text, response text, and unknown fields are rejected.
The analyzer compares fresh input, cache writes, cache reads, request timing, and repeated sections.
The dashboard explains the observed problem, the change to test, and any risk to response quality.
PRIVACY BOUNDARY
The collector sends token counts, timing, model names, outcomes, and one-way identifiers. It rejects prompt and response text.
Read the security modelOPEN ENGINEERING
The tests, system design, security decisions, and remaining production work are public.
FAQ
You can connect a LiteLLM proxy callback, upload existing LiteLLM JSONL logs, or upload a normalized CacheEconomics trace.
No. The collector sends usage metadata and one-way identifiers. Prompt text and response text are rejected before they can enter the delivery queue or hosted API.
It multiplies monthly input-token spend by the share of reusable input and the cache-read price reduction. The estimate excludes output tokens, cache-write charges, negotiated rates, and implementation cost.
The dashboard needs enough complete request data and applicable pricing or invoice information. When either is missing, it explains what is needed instead of showing an unsupported dollar figure.
Yes. Organizations keep each team’s data separate, and roles control who can view results, manage sources, run analyses, and administer members.
The public demo uses generated test traffic. It lets visitors explore the product without exposing customer prompts, usage, invoices, or performance data.
READY TO ANALYZE YOUR TRAFFIC?
Connect LiteLLM or upload a usage file. See where cached input is being missed and what to change.
DASHBOARD SIGN-IN
Review cache use, cost estimates, recommended fixes, request latency, and data-delivery health for your organization.
SECURE WORKSPACE
Sign in with the identity attached to your organization. Your role controls the sources, analyses, and operational actions you can use.
CACHE PERFORMANCE