PROMPT-CACHE COST INTELLIGENCE

Spend less on repeated context.

Find where prompt-cache reuse breaks, what it costs, and which engineering change is worth testing first.

  • Prompt bodies rejected
  • Provider facts source-dated
  • Core library stays offline
CacheEconomics Synthetic walkthrough
RESEARCH AGENTCache overview
Receiving data
Accepted events286Checked synthetic trace
Input from cache16.7%Observed token ratio
Ranked findings4Evidence-labelled actions
Monthly impactWithheldNo invoice supplied
Request signals7 days
Top actionTTL-1

Test a longer cache lifetime for the stable prefix.

STRUCTURAL EVIDENCE

Real dashboard contract · checked synthetic data · monetary values withheld

Apache-2.0 source Prompt-free ingest schema Reproducible evaluation packet Scanned, digest-pinned deployments

THE QUESTIONS THAT MATTER

Know where cache value leaks.

Follow the cost, the cause, and the next action in one view.

01

See the write/read balance

Separate writes, reads, fresh input, and output.

02

Find broken reuse windows

Match stable prefixes and request cadence to the right TTL.

03

Keep money defensible

See the evidence behind every released cost figure.

04

Watch the pipeline

Track volume, latency, errors, jobs, and ingestion lag.

QUICK ESTIMATE

See the opportunity in seconds.

Move three sliders. Start with the example or use your own workload.

ILLUSTRATIVE STARTING POINT Adjust the assumptions

Example assumptions only. Check the cache-read rate in your provider contract.

POTENTIAL MONTHLY SAVINGS $11,250 45% of standard-rate input spend
Standard input$25,000
With cache reads$13,750
12-month opportunity$135,000 Input-cost reduction45%

Rough planning estimate before cache-write premiums, output tokens, volume terms, and implementation cost. The workspace uses measured traffic and stronger evidence gates.

PRODUCT TOUR

See the product move.

Short tours built from the checked synthetic trace. Dollar values stay withheld because the fixture has no invoice.

SYNTHETIC PRODUCT WALKTHROUGH00:12
Read the synthetic case study

HOW THE ANALYSIS WORKS

Four steps. Every claim traceable.

  1. 01

    Observe locally

    Collect counters, timing, outcomes, and keyed segment fingerprints.

  2. 02

    Validate before storage

    Reject prompt bodies and every field outside the allow-list.

  3. 03

    Analyze offline

    Evaluate token classes, structure, reuse timing, and dated provider facts.

  4. 04

    Publish the evidence

    Pair each action with its evidence, quality risk, and release state.

PRIVACY BOUNDARY

Your prompts stay where your models run.

Only operational metadata and keyed fingerprints cross the boundary. Prompt and completion bodies are rejected twice.

Read the security model
YOUR ENVIRONMENT Prompt bodies Completion bodies Local HMAC key
HOSTED WORKSPACE Token counters Timing & outcomes Keyed fingerprints

BUILT FOR SCRUTINY

Inspect the work behind the claim.

Architecture, assumptions, results, and deployment proof live beside the code.

FAQ

Before you connect traffic.

What data leaves my environment?

Timestamps, token counters, normalized outcome and timing fields, model and surface labels, and locally keyed segment fingerprints. The collector validates that shape before it writes to its local delivery queue.

Does the collector change live requests?

The documented LiteLLM setup uses observe-only mode. Marker placement is a separate opt-in behavior that should be evaluated against the workload before use.

Which ingestion paths work today?

The implemented live path is a LiteLLM proxy callback. Existing LiteLLM JSONL logs and normalized CacheEconomics traces can also be uploaded explicitly.

Why would the dashboard withhold a dollar figure?

Monetary results require sufficient request coverage, structural evidence, reconciled usage, and applicable pricing or invoice evidence. The API returns the reason when a value stays withheld.

How are teams separated?

Every workspace request carries an organization context. Membership and roles are enforced by the API, with PostgreSQL row-level security providing a second tenant boundary.

What is the current deployment meant for?

The public environment is a low-traffic portfolio staging system using synthetic data. The runbook lists the backup, telemetry, capacity, and incident exercises required before customer production use.

CACHE ECONOMICS YOU CAN DEFEND

Bring the evidence into one workspace.

Connect a prompt-free source and turn repeated context into ranked, reviewable actions.

cacheeconomics

YOUR CACHE CONTROL ROOM

Pick up where the evidence leads.

Review spend evidence, cache ratios, ranked recommendations, request latency, and ingestion health for your organization.

The local boundary stays intact The installed Python package remains offline and socket-free.
Product overview

SECURE WORKSPACE

Open your workspace

Sign in with the identity attached to your organization. Your role controls the sources, analyses, and operational actions you can use.

Loading sign-in configuration…
Authorization enforced by the API
No browser token persistence
Prompt-free hosted schema