Usage and limits
Infrared calls two AI providers: Claude (Anthropic) for every agent step and for summary video scripts, and Gemini (Google) for summary video narration, stage images and presenter clips. Both have limits, and some of Google's are small and daily. Infrared shows where you stand on every page.
The AI usage chip
At the top right of every page, the AI usage chip shows the org's AI spend today, Claude and Gemini together, with a status dot:
| Dot | Means |
|---|---|
| Green | Within every limit |
| Amber | A Gemini model has used 80% of its daily limit, or Anthropic's per-minute rate limit is nearly spent |
| Red | A daily limit is reached, or the Anthropic key is out of credit or rejected |
Click it for today's Claude spend and tokens, Gemini spend and summary videos, a meter for each Gemini daily limit, the reasons for an amber or red dot, how long until Google's limits reset, and a link to the Usage page. It refreshes every minute.
The Usage page
Usage in the navigation shows, for the last 7 or 30 days:
- Today: spend, Claude tokens (and how many were read from cache), and Gemini spend with the number of summary videos.
- Gemini daily limits: each model's requests since Google's last reset against its daily limit. Google's quota is per Google project, so the counts include every org; the note says how many were this org's.
- Spend by day: Claude and Gemini stacked per day; hover a day for its numbers, or open the table.
- Claude by model and Claude by AgentRole: steps, tokens and cost. Summary video scripts are their own row.
- Anthropic rate limits: the requests and tokens per minute Anthropic last reported for the key, and what was left.
- Settings links: the model provider key, summary videos and their monthly cap, models per AgentRole and model tiers.
Platform admins also see Every org (each org's spend today and over the window) and can set the Gemini daily limits.
Buttons that would spend a Gemini video request, like making a summary video with a fresh presenter, can show how many video requests are left today and what it costs.
Measured or estimated
| Figure | Where it comes from |
|---|---|
| Claude spend and tokens | Measured: each agent step's usage as the Agent SDK reports it, and each summary video script's |
| Gemini spend | Measured per summary video (clips, image, narration), priced from Google's list prices |
| Gemini requests per model | Counted by Infrared's summary video Jobs, since midnight Pacific |
| Gemini daily limits | Configured, not reported: Google doesn't expose remaining quota. Defaults are the limits seen on an AI Studio Tier 1 project (Veo 3.1 Lite about 10 requests a day, Gemini 2.5 Flash TTS 100). Set the real ones from AI Studio → Usage → Rate limits |
| A limit reached | Google refused a request with a per-day 429 since the last reset (or since the Gemini key changed), or the count reached the configured limit. Google's 429 names the quota it hit: a per-minute limit clears after the short wait Google asks for (Infrared waits and retries, and counts it apart), so it doesn't mark the day as used up |
| Anthropic rate limits | Anthropic's anthropic-ratelimit-* response headers, recorded by Check now (Settings → Model provider) and by each summary video's script call. Agent steps don't record them |
Requests Infrared doesn't make (anything else using the same Google project or Anthropic key) don't appear in the counts.
The API
curl https://<host>/api/v1/orgs/<org>/usage?days=7 -H "Authorization: Bearer $INFRARED_TOKEN"
curl https://<host>/api/v1/usage?days=7 -H "Authorization: Bearer $INFRARED_TOKEN" # platform admins: every org
curl -X PUT https://<host>/api/v1/ai-limits -H "Authorization: Bearer $INFRARED_TOKEN" \
-H 'Content-Type: application/json' -d '{"gemini": {"veo-3.1-lite-generate-preview": 10}}' # platform admins
The daily limits live in the ConfigMap infrared-ai-limits in Infrared's namespace. The label of the Gemini key in use comes from the annotation infrared.darkshift.io/key-label on its Secret; the key itself is never shown.