> ## Documentation Index
> Fetch the complete documentation index at: https://www.truefoundry.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Cache Metrics

> Query Gateway cache lookup, hit, and savings metrics via API.

The **Gateway Cache Metrics Query API** provides a flexible way to query cache-eligible requests: cache lookups, hits, savings, and the model fields that go with them. Internally this is the same underlying table as `modelMetrics`, restricted to rows where `CacheLookupStatus` is set. You can retrieve either **distribution** (aggregated) or **timeseries** results with powerful filtering and grouping.

<Info>
  This page covers `datasource: "cacheMetrics"`. For other datasources, see [Model](/docs/ai-gateway/fetch-model-metrics), [MCP](/docs/ai-gateway/fetch-mcp-metrics), [Guardrail](/docs/ai-gateway/fetch-guardrail-metrics), [Routing](/docs/ai-gateway/fetch-routing-metrics), and [Agent](/docs/ai-gateway/fetch-agent-metrics) metrics.
</Info>

All requests go to a single endpoint:

```
POST https://{your_control_plane_url}/api/svc/v1/llm-gateway/metrics/query
```

Send JSON with `Authorization: Bearer <your_api_key>` and `Content-Type: application/json`.

## Access control

Access to metrics is governed by the **data access rules configured by your tenant**. The server applies these rules automatically based on the caller's identity—you don't pass any RBAC or scoping fields in the request. What a caller can query (their own data, their team's data, or tenant-wide data) depends entirely on the rules an admin has set up.

See [Configure Data Access](/docs/ai-gateway/data-access) for how these rules are defined and evaluated.

## Authentication

<Accordion title="Get your API key">
  Authenticate with your TrueFoundry API key. You can use either a Personal Access Token **(PAT)** or Virtual Account Token **(VAT)**.

  1. **Personal Access Token (PAT)**: Go to Access → Personal Access Tokens in your TrueFoundry dashboard
  2. **Virtual Account Token (VAT)**: Go to Access → Virtual Account Tokens (requires admin permissions)

  For detailed authentication setup, see our [Authentication guide](/docs/ai-gateway/authentication).
</Accordion>

## Quick start

<Warning>
  By default, cache metrics include **both models and virtual models**. To restrict to one, use `{"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}` for model-only metrics, or `value: false` for virtual-model-only metrics.
</Warning>

<Note>
  The server automatically adds `WHERE "CacheLookupStatus" IS NOT NULL` to every cache query; you do not (and should not) add it yourself. Because cache shares its underlying table with `modelMetrics`, every model field is reachable in addition to the cache-specific ones below.
</Note>

<Note>
  The virtual-model column has two aliases. In `groupBy` and `aggregations[].column` use `virtualModel`. In `filters[].fieldName` and in response keys, the name is `virtualModelName`. They refer to the same underlying database column.
</Note>

<Tabs>
  <Tab title="Distribution query">
    Cost savings and tokens read from cache, grouped by cache type and namespace:

    ```python theme={"dark"}
    import requests

    response = requests.post(
        "https://{your_control_plane_url}/api/svc/v1/llm-gateway/metrics/query",
        headers={
            "Authorization": "Bearer <your_api_key>",
            "Content-Type": "application/json"
        },
        json={
            "startTs": "2026-04-21T00:00:00.000Z",
            "endTs": "2026-04-22T00:00:00.000Z",
            "datasource": "cacheMetrics",
            "type": "distribution",
            "aggregations": [
                {"type": "sum", "column": "potentialCostSavings"},
                {"type": "sum", "column": "cacheReadInputTokens"},
                {"type": "p50", "column": "cacheLookupLatencyMs"}
            ],
            "groupBy": ["cacheType", "cacheNamespace"],
            "filters": [
                {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
            ]
        }
    )

    print(response.json())
    ```
  </Tab>

  <Tab title="Timeseries query">
    Hourly cache savings and p99 lookup latency, grouped by cache type:

    ```python theme={"dark"}
    import requests

    response = requests.post(
        "https://{your_control_plane_url}/api/svc/v1/llm-gateway/metrics/query",
        headers={
            "Authorization": "Bearer <your_api_key>",
            "Content-Type": "application/json"
        },
        json={
            "startTs": "2026-04-21T00:00:00.000Z",
            "endTs": "2026-04-22T00:00:00.000Z",
            "datasource": "cacheMetrics",
            "type": "timeseries",
            "interval": "1 hour",
            "aggregations": [
                {"type": "sum", "column": "potentialCostSavings"},
                {"type": "p99", "column": "cacheLookupLatencyMs"}
            ],
            "groupBy": ["cacheType"],
            "filters": [
                {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
            ]
        }
    )

    print(response.json())
    ```
  </Tab>
</Tabs>

## API reference

Post JSON to the endpoint above with `Authorization: Bearer <your_api_key>` and `Content-Type: application/json`.

### Request parameters

<ParamField path="startTs" type="string" required>
  ISO 8601 timestamp marking the **inclusive** lower bound of the query window (e.g. `"2026-04-21T00:00:00.000Z"`).
</ParamField>

<ParamField path="endTs" type="string" required>
  ISO 8601 timestamp marking the **exclusive** upper bound of the query window (e.g. `"2026-04-22T00:00:00.000Z"`).
</ParamField>

<ParamField path="datasource" type="string" required>
  The data source to query. Use `"cacheMetrics"` for Gateway cache metrics.
</ParamField>

<ParamField path="type" type="string" required>
  The type of query to execute:

  * `"distribution"`: returns aggregated rows (one row per `groupBy` combination).
  * `"timeseries"`: returns time-bucketed rows (one row per bucket per `groupBy` combination). Requires `interval`.
</ParamField>

<ParamField path="aggregations" type="array">
  Array of `{ type, column }` objects describing the aggregations to compute. When omitted, only the implicit `total = COUNT(*)` is returned.

  ```json theme={"dark"}
  "aggregations": [
      {"type": "sum", "column": "potentialCostSavings"},
      {"type": "sum", "column": "cacheReadInputTokens"},
      {"type": "p50", "column": "cacheLookupLatencyMs"}
  ]
  ```

  <Accordion title="Supported aggregation types">
    | Type                                                          | Description                                                   |
    | ------------------------------------------------------------- | ------------------------------------------------------------- |
    | `sum`                                                         | Sum of values                                                 |
    | `count`                                                       | Non-null count of the column                                  |
    | `countDistinct`                                               | Distinct count                                                |
    | `min`                                                         | Minimum value                                                 |
    | `max`                                                         | Maximum value                                                 |
    | `avg`                                                         | Average                                                       |
    | `p5`, `p10`, `p25`, `p50`, `p75`, `p90`, `p95`, `p99`, `p999` | Percentiles (approximate)                                     |
    | `rateSum`                                                     | `sum` normalised by the interval in seconds (timeseries only) |
    | `rateAvg`                                                     | `avg` normalised by the interval in seconds (timeseries only) |
    | `rateMin`                                                     | `min` normalised by the interval in seconds (timeseries only) |
    | `rateMax`                                                     | `max` normalised by the interval in seconds (timeseries only) |
    | `ratePerMinute`                                               | Value divided by the interval in minutes (timeseries only)    |
  </Accordion>

  <Accordion title="Supported aggregation columns">
    | Column                        | Notes                                             |
    | ----------------------------- | ------------------------------------------------- |
    | `cacheLookupLatencyMs`        | The cache-lookup latency itself                   |
    | `potentialCostSavings`        | Cost saved by cache hits (USD)                    |
    | `cacheCreationInputTokens`    | Tokens written into the cache                     |
    | `cacheReadInputTokens`        | Tokens read from the cache                        |
    | `costInUSD`                   | Cost incurred (USD)                               |
    | `inputTokens`                 | Number of input tokens                            |
    | `outputTokens`                | Number of output tokens                           |
    | `latencyMs`                   | Total request latency (ms)                        |
    | `timeToFirstTokenMs`          | Time to the first generated token (ms)            |
    | `interTokenLatencyMs`         | Latency between consecutive generated tokens (ms) |
    | `timePerOutputTokenLatencyMs` | Latency per output token (ms)                     |

    All scalar and percentile aggregation types apply to every column above.
  </Accordion>
</ParamField>

<ParamField path="groupBy" type="array">
  Array of field names to group results by. Custom metadata keys are supported with a `metadata.` prefix (e.g. `"metadata.environment"`).

  ```json theme={"dark"}
  "groupBy": ["cacheType", "cacheNamespace", "metadata.environment"]
  ```

  <Accordion title="Available group-by fields">
    | Field                  | Notes                                                                         |
    | ---------------------- | ----------------------------------------------------------------------------- |
    | `cacheType`            | e.g. `semantic`, `simple`                                                     |
    | `cacheNamespace`       | Logical bucket within a cache type                                            |
    | `modelName`            | The underlying model name                                                     |
    | `virtualModel`         | The virtual-model name (when the request was routed through one)              |
    | `requestType`          | Type of request, e.g. `ChatCompletion`, `Embedding`                           |
    | `providerModelName`    | Underlying provider model name                                                |
    | `providerAccountType`  | Account type of the provider (e.g. `model`, `mcp-server`, `guardrail-config`) |
    | `errorCode`            | HTTP error code returned, when applicable                                     |
    | `userEmail`            | Group by user (response key: `createdBySubjectSlug`)                          |
    | `virtualaccount`       | Group by virtual account (response key: `createdBySubjectSlug`)               |
    | `team`                 | Unnests the `Teams` array                                                     |
    | `createdBySubjectType` | Distinguishes `user` vs `virtualaccount`                                      |
    | `metadata.<key>`       | Group by a custom metadata key                                                |

    When `groupBy` contains `userEmail` (without `virtualaccount`), the server auto-injects `WHERE CreatedBySubjectType = 'user'`. `virtualaccount` alone auto-injects `'virtualaccount'`. When both appear, scope it yourself with `createdBySubjectType` if needed.
  </Accordion>
</ParamField>

<ParamField path="filters" type="array">
  Array of filter objects, AND-combined. See [Filtering](#filtering) below for the full operator reference and the per-field allow-list.
</ParamField>

<ParamField path="interval" type="string">
  **Required for timeseries queries.** Bucket size as `<positive integer> <unit>`, where `<unit>` is one of `second`, `minute`, `hour`, `day`, `week`, `month`, `year` (with or without a trailing `s`). Examples: `"30 second"`, `"5 minute"`, `"1 hour"`, `"1 day"`. Compound expressions like `"1 hour 30 minute"` are rejected.
</ParamField>

<ParamField path="intervalInSeconds" type="number" deprecated>
  **Deprecated alias for `interval`.** Accepts a positive integer number of seconds (e.g. `3600` for hourly). Prefer `interval` in new code. If both are provided, `interval` wins.
</ParamField>

## Filtering

Filters narrow down the rows that go into each aggregation and group. They are AND-combined; there is no OR-group support. The server enforces a per-field operator allow-list, so the exact subset of operators you can use depends on the field.

<Tabs>
  <Tab title="Field filters">
    For standard datasource fields, use `fieldName`:

    ```json theme={"dark"}
    {
        "fieldName": "cacheType",
        "operator": "IN",
        "value": ["semantic"]
    }
    ```
  </Tab>

  <Tab title="Metadata filters">
    For custom request-metadata keys, use `metadataKey`. Works on every datasource:

    ```json theme={"dark"}
    {
        "metadataKey": "environment",
        "operator": "IN",
        "value": ["production"]
    }
    ```
  </Tab>
</Tabs>

<Accordion title="Filterable fields and allowed operators">
  Cache metrics share the underlying table with model metrics, so all model filter fields are also reachable here. The cache-specific fields are `cacheType`, `cacheNamespace`, `cacheLookupStatus`, and the cache token columns. `cacheType` is narrow (`IN`/`NOT_IN` only); the others accept the full string or numeric operator set as documented below.

  | Field                            | Type    | Allowed operators                                                           |
  | -------------------------------- | ------- | --------------------------------------------------------------------------- |
  | `cacheType`                      | string  | `IN`, `NOT_IN`                                                              |
  | `cacheNamespace`                 | string  | `IN`, `NOT_IN`, `STRING_CONTAINS`, `STRING_STARTS_WITH`, `STRING_ENDS_WITH` |
  | `cacheLookupStatus`              | string  | full string operator set                                                    |
  | `modelName`                      | string  | full string operator set                                                    |
  | `requestType`                    | string  | full string operator set                                                    |
  | `virtualModelName`               | string  | full string operator set, including `IS_NULL`                               |
  | `httpStatusCode`                 | number  | full numeric operator set                                                   |
  | `errorCode`                      | string  | full string operator set                                                    |
  | `isFailure`                      | boolean | `EQUAL`, `IS_NULL`                                                          |
  | `providerAccountType`            | string  | full string operator set                                                    |
  | `providerModelName`              | string  | full string operator set                                                    |
  | `costInUSD`                      | number  | numeric range operators                                                     |
  | `inputTokens`                    | number  | numeric range operators                                                     |
  | `outputTokens`                   | number  | numeric range operators                                                     |
  | `cacheCreationInputTokens`       | number  | numeric range operators                                                     |
  | `cacheReadInputTokens`           | number  | numeric range operators                                                     |
  | `userEmail`                      | string  | full string operator set                                                    |
  | `virtualAccount`                 | string  | full string operator set                                                    |
  | `team`                           | array   | `ARRAY_HAS_ANY`, `ARRAY_HAS_NONE`                                           |
  | `latencyMs`                      | number  | numeric range operators                                                     |
  | `conversationID`                 | string  | full string operator set                                                    |
  | `traceId`                        | string  | `EQUAL`                                                                     |
  | `metadataKey` / `metadata.<key>` | string  | full string operator set                                                    |
</Accordion>

<Accordion title="Filter operators reference">
  **String field operators**

  | Operator                 | Description                                                                        | Example value            |
  | ------------------------ | ---------------------------------------------------------------------------------- | ------------------------ |
  | `EQUAL`                  | Exact match                                                                        | `"hit"`                  |
  | `NOT_EQUAL`              | Not equal to value                                                                 | `"miss"`                 |
  | `IN`                     | Match any value in the list                                                        | `["semantic", "simple"]` |
  | `NOT_IN`                 | Exclude values in the list                                                         | `["deprecated"]`         |
  | `STRING_CONTAINS`        | Contains substring                                                                 | `"prod"`                 |
  | `STRING_NOT_CONTAINS`    | Does not contain substring                                                         | `"staging"`              |
  | `STRING_STARTS_WITH`     | Starts with prefix                                                                 | `"prod-"`                |
  | `STRING_NOT_STARTS_WITH` | Does not start with prefix                                                         | `"internal-"`            |
  | `STRING_ENDS_WITH`       | Ends with suffix                                                                   | `"-v1"`                  |
  | `STRING_NOT_ENDS_WITH`   | Does not end with suffix                                                           | `"-deprecated"`          |
  | `IS_NULL`                | `true` matches rows where the field is unset; `false` matches rows where it is set | `true`                   |

  **Numeric field operators**

  | Operator             | Description                                                                        | Example value     |
  | -------------------- | ---------------------------------------------------------------------------------- | ----------------- |
  | `EQUAL`              | Exact match                                                                        | `1000`            |
  | `NOT_EQUAL`          | Not equal to value                                                                 | `0`               |
  | `IN`                 | Match any value in the list                                                        | `[200, 404, 500]` |
  | `NOT_IN`             | Exclude values in the list                                                         | `[0]`             |
  | `GREATER_THAN`       | Strictly greater than                                                              | `1000`            |
  | `LESS_THAN`          | Strictly less than                                                                 | `5000`            |
  | `GREATER_THAN_EQUAL` | Greater than or equal to                                                           | `100`             |
  | `LESS_THAN_EQUAL`    | Less than or equal to                                                              | `1000`            |
  | `BETWEEN`            | Between two values (inclusive)                                                     | `[500, 5000]`     |
  | `IS_NULL`            | `true` matches rows where the field is unset; `false` matches rows where it is set | `true`            |

  **Boolean field operators**

  | Operator  | Description                                                                        | Example value |
  | --------- | ---------------------------------------------------------------------------------- | ------------- |
  | `EQUAL`   | Exact match                                                                        | `true`        |
  | `IS_NULL` | `true` matches rows where the field is unset; `false` matches rows where it is set | `false`       |

  **Array field operators (used by `team`)**

  | Operator         | Description                                    | Example value                 |
  | ---------------- | ---------------------------------------------- | ----------------------------- |
  | `ARRAY_HAS_ANY`  | Match if the array contains any of the values  | `["team-alpha", "team-beta"]` |
  | `ARRAY_HAS_NONE` | Match if the array contains none of the values | `["excluded-team"]`           |
</Accordion>

<Accordion title="Custom metadata, team unnesting, and combining filters">
  **Custom metadata filtering and grouping.** Every datasource supports filtering and grouping by custom request-metadata keys:

  * **Filter:** `{ "metadataKey": "environment", "operator": "EQUAL", "value": "prod" }`
  * **Group:** include `"metadata.environment"` in the `groupBy` array (string literal, prefix is `metadata.`).

  Metadata fields are treated as strings; use the string field operators above.

  **Implicit team unnesting.** When `team` is in `groupBy` (or used as the column of an aggregation), the server transparently UNNESTs the `Teams` array CTE before applying RBAC. Callers don't need to do anything extra. Rows whose `Teams` array is NULL or empty drop out naturally.

  **Combining multiple filters.** Filters are AND-combined:

  ```json theme={"dark"}
  {
      "startTs": "2026-04-21T00:00:00.000Z",
      "endTs": "2026-04-22T00:00:00.000Z",
      "datasource": "cacheMetrics",
      "type": "distribution",
      "filters": [
          {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true},
          {"fieldName": "cacheType", "operator": "IN", "value": ["semantic"]},
          {"fieldName": "cacheNamespace", "operator": "STRING_STARTS_WITH", "value": "prod-"},
          {"fieldName": "team", "operator": "ARRAY_HAS_ANY", "value": ["team-alpha"]}
      ],
      "groupBy": ["cacheNamespace"]
  }
  ```
</Accordion>

## Query examples

Every example posts a JSON body to the endpoint above. To keep the snippets short, only the `json` body is shown; the request wrapper is identical to the [Quick start](#quick-start).

<Note>
  The examples below pin the model side with `{"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}`. Flip to `false` (and swap `groupBy: ["modelName"]` for `groupBy: ["virtualModel"]`) for virtual-model-only cache stats.
</Note>

### Distribution examples

<AccordionGroup>
  <Accordion title="Top namespaces by tokens served from cache">
    Sum of `cacheReadInputTokens` per namespace. Surfaces which buckets do the most work:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "sum", "column": "cacheReadInputTokens"}
        ],
        "groupBy": ["cacheNamespace"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Lookup latency percentiles by cache type">
    p50, p90, and p99 lookup latency grouped by cache type:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "p50", "column": "cacheLookupLatencyMs"},
            {"type": "p90", "column": "cacheLookupLatencyMs"},
            {"type": "p99", "column": "cacheLookupLatencyMs"}
        ],
        "groupBy": ["cacheType"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Savings by model">
    Sum of cost savings per underlying model:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"}
        ],
        "groupBy": ["modelName"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Semantic cache only">
    Restrict to a specific cache type and break savings down by model:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"},
            {"type": "sum", "column": "cacheReadInputTokens"}
        ],
        "groupBy": ["modelName"],
        "filters": [
            {"fieldName": "cacheType", "operator": "IN", "value": ["semantic"]},
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Filter namespaces by prefix">
    Use `STRING_STARTS_WITH` on `cacheNamespace`, handy when prod and staging share a cache type:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"}
        ],
        "groupBy": ["cacheNamespace"],
        "filters": [
            {"fieldName": "cacheNamespace", "operator": "STRING_STARTS_WITH", "value": "prod-"},
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Hit vs miss breakdown">
    Group by `cacheLookupStatus` to see hits vs misses per cache type:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [],
        "groupBy": ["cacheType", "cacheLookupStatus"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Tokens written vs read">
    Compare cache-creation tokens to cache-read tokens per namespace:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "distribution",
        "aggregations": [
            {"type": "sum", "column": "cacheCreationInputTokens"},
            {"type": "sum", "column": "cacheReadInputTokens"}
        ],
        "groupBy": ["cacheNamespace"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>
</AccordionGroup>

### Timeseries examples

Every timeseries query must include `interval` (or the deprecated `intervalInSeconds`). Buckets are expressed as `<positive integer> <unit>` strings like `"5 minute"`, `"1 hour"`, or `"1 day"`.

<AccordionGroup>
  <Accordion title="Hourly savings by namespace">
    Track cost savings per namespace over time:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 hour",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"}
        ],
        "groupBy": ["cacheNamespace"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Hourly p99 lookup latency by cache type">
    Watch for regressions in cache lookups:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 hour",
        "aggregations": [
            {"type": "p99", "column": "cacheLookupLatencyMs"}
        ],
        "groupBy": ["cacheType"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Hourly tokens read from cache">
    Track cache read volume per cache type:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 hour",
        "aggregations": [
            {"type": "sum", "column": "cacheReadInputTokens"}
        ],
        "groupBy": ["cacheType"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Hourly hits vs misses">
    Volume by `cacheLookupStatus` over time:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 hour",
        "aggregations": [],
        "groupBy": ["cacheLookupStatus"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Daily savings over a week">
    Daily cost savings across a 7-day window:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-14T00:00:00.000Z",
        "endTs": "2026-04-21T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 day",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"}
        ],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="5-minute lookup latency in an incident window">
    Fine-grained breakdown to investigate a regression:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T14:00:00.000Z",
        "endTs": "2026-04-21T16:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "5 minute",
        "aggregations": [
            {"type": "p99", "column": "cacheLookupLatencyMs"}
        ],
        "groupBy": ["cacheType"],
        "filters": [
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>

  <Accordion title="Hourly savings for semantic cache only">
    Combine cache-type filter with namespace breakdown:

    ```python theme={"dark"}
    json={
        "startTs": "2026-04-21T00:00:00.000Z",
        "endTs": "2026-04-22T00:00:00.000Z",
        "datasource": "cacheMetrics",
        "type": "timeseries",
        "interval": "1 hour",
        "aggregations": [
            {"type": "sum", "column": "potentialCostSavings"}
        ],
        "groupBy": ["cacheNamespace"],
        "filters": [
            {"fieldName": "cacheType", "operator": "IN", "value": ["semantic"]},
            {"fieldName": "virtualModelName", "operator": "IS_NULL", "value": true}
        ]
    }
    ```
  </Accordion>
</AccordionGroup>

## Response format

Every successful response has the same outer shape:

```json theme={"dark"}
{
  "data": {
    "dataPoints": [
      {
        "startTimestamp": "2026-04-29T12:00:00.000Z",
        "endTimestamp": "2026-04-29T13:00:00.000Z",
        "total": 1234,
        "<aggregationKey>": 0,
        "<groupByKey>": "value-or-null"
      }
    ]
  }
}
```

* **`total`**: implicit `COUNT(*)` for the row. Always present.
* **`<aggregationKey>`**: one key per requested aggregation. The key is `<type><Column>` in camelCase (e.g. `sumPotentialCostSavings`, `p99CacheLookupLatencyMs`, `sumCacheReadInputTokens`).
* **`<groupByKey>`**: one key per `groupBy` entry. The key is the lowerCamelCase form of the underlying column. Two special mappings:
  * `userEmail` and `virtualaccount` both map to `createdBySubjectSlug` in the response.
  * `team` maps to `team` (the value is a single unnested scalar, not an array).
    All other `groupBy` keys preserve their lowerCamelCase name.
* **`startTimestamp`** / **`endTimestamp`**: present only for timeseries responses. Bucket start and end as ISO 8601 timestamp strings; `endTimestamp` equals the next bucket's `startTimestamp`. Distribution responses omit both.

<AccordionGroup>
  <Accordion title="Distribution response example">
    ```json theme={"dark"}
    {
      "data": {
        "dataPoints": [
          {
            "cacheType": "semantic",
            "cacheNamespace": "prod-chat",
            "total": 4200,
            "sumPotentialCostSavings": 12.34,
            "sumCacheReadInputTokens": 480000,
            "p50CacheLookupLatencyMs": 8.2
          },
          {
            "cacheType": "semantic",
            "cacheNamespace": "prod-search",
            "total": 1850,
            "sumPotentialCostSavings": 4.81,
            "sumCacheReadInputTokens": 210000,
            "p50CacheLookupLatencyMs": 11.4
          }
        ]
      }
    }
    ```
  </Accordion>

  <Accordion title="Timeseries response example">
    ```json theme={"dark"}
    {
      "data": {
        "dataPoints": [
          {
            "startTimestamp": "2026-04-21T00:00:00.000Z",
            "endTimestamp": "2026-04-21T01:00:00.000Z",
            "cacheType": "semantic",
            "total": 180,
            "sumPotentialCostSavings": 0.52,
            "p99CacheLookupLatencyMs": 22.0
          },
          {
            "startTimestamp": "2026-04-21T01:00:00.000Z",
            "endTimestamp": "2026-04-21T02:00:00.000Z",
            "cacheType": "semantic",
            "total": 210,
            "sumPotentialCostSavings": 0.61,
            "p99CacheLookupLatencyMs": 24.5
          }
        ]
      }
    }
    ```
  </Accordion>
</AccordionGroup>

<Info>
  If `groupBy` is empty or omitted, the response collapses to a single row (or one row per timeseries bucket) summarising every cache-eligible request inside the window. Note that the server always pins `WHERE "CacheLookupStatus" IS NOT NULL`, so non-cache rows are filtered out automatically.
</Info>

<Note>
  Virtual-model rows surface under the `virtualModelName` key in the response (not `virtualModel`), because the response key is the lowerCamelCase of the underlying database column.
</Note>

### Error responses

A malformed query returns `400 Bad Request`:

```json theme={"dark"}
{
  "statusCode": 400,
  "message": "Invalid query",
  "details": ["..."]
}
```

<Accordion title="Common causes and other status codes">
  Common causes of `400`:

  * Operator not allowed on this field, for example, `STRING_CONTAINS` on `cacheType` (it supports only `IN`/`NOT_IN`).
  * Missing required `value` (or wrong shape, e.g. scalar where array is expected for `IN` / `BETWEEN`).
  * Unknown field name for the datasource.
  * Invalid `interval` format (compound expressions, unrecognised unit, non-positive integer).
  * Missing required `interval` for a timeseries query.

  Other status codes:

  * `401 Unauthorized`: missing or invalid bearer token.
  * `403 Forbidden`: caller does not have permission for the requested scope.
  * `500 Internal Server Error`: unexpected server error while executing the query.
</Accordion>
