Skip to main content
Version: 0.1 (next)

Usage and limits

Infrared calls two AI providers: Claude (Anthropic) for every agent step and for summary video scripts, and Gemini (Google) for summary video narration, stage images and presenter clips. Both have limits, and some of Google's are small and daily. Infrared shows where you stand on every page.

The AI usage chip​

At the top right of every page, the AI usage chip shows the org's AI spend today, Claude and Gemini together, with a status dot:

DotMeans
GreenWithin every limit
AmberA Gemini model has used 80% of its daily limit, or Anthropic's per-minute rate limit is nearly spent
RedA daily limit is reached, or the Anthropic key is out of credit or rejected

Click it for today's Claude spend and tokens, Gemini spend and summary videos, a meter for each Gemini daily limit, the reasons for an amber or red dot, how long until Google's limits reset, and a link to the Usage page. It refreshes every minute.

The Usage page​

Usage in the navigation shows, for the last 7 or 30 days:

  • Today: spend, Claude tokens (and how many were read from cache), and Gemini spend with the number of summary videos.
  • Gemini daily limits: each model's requests since Google's last reset against its daily limit. Google's quota is per Google project, so the counts include every org; the note says how many were this org's.
  • Spend by day: Claude and Gemini stacked per day; hover a day for its numbers, or open the table.
  • Claude by model and Claude by AgentRole: steps, tokens and cost. Summary video scripts are their own row.
  • Anthropic rate limits: the requests and tokens per minute Anthropic last reported for the key, and what was left.
  • Settings links: the model provider key, summary videos and their monthly cap, models per AgentRole and model tiers.

Platform admins also see Every org (each org's spend today and over the window) and can set the Gemini daily limits.

Buttons that would spend a Gemini video request, like making a summary video with a fresh presenter, can show how many video requests are left today and what it costs.

Measured or estimated​

FigureWhere it comes from
Claude spend and tokensMeasured: each agent step's usage as the Agent SDK reports it, and each summary video script's
Gemini spendMeasured per summary video (clips, image, narration), priced from Google's list prices
Gemini requests per modelCounted by Infrared's summary video Jobs, since midnight Pacific
Gemini daily limitsConfigured, not reported: Google doesn't expose remaining quota. Defaults are the limits seen on an AI Studio Tier 1 project (Veo 3.1 Lite about 10 requests a day, Gemini 2.5 Flash TTS 100). Set the real ones from AI Studio → Usage → Rate limits
A limit reachedGoogle refused a request with a per-day 429 since the last reset (or since the Gemini key changed), or the count reached the configured limit. Google's 429 names the quota it hit: a per-minute limit clears after the short wait Google asks for (Infrared waits and retries, and counts it apart), so it doesn't mark the day as used up
Anthropic rate limitsAnthropic's anthropic-ratelimit-* response headers, recorded by Check now (Settings → Model provider) and by each summary video's script call. Agent steps don't record them

Requests Infrared doesn't make (anything else using the same Google project or Anthropic key) don't appear in the counts.

The API​

curl https://<host>/api/v1/orgs/<org>/usage?days=7 -H "Authorization: Bearer $INFRARED_TOKEN"
curl https://<host>/api/v1/usage?days=7 -H "Authorization: Bearer $INFRARED_TOKEN" # platform admins: every org
curl -X PUT https://<host>/api/v1/ai-limits -H "Authorization: Bearer $INFRARED_TOKEN" \
-H 'Content-Type: application/json' -d '{"gemini": {"veo-3.1-lite-generate-preview": 10}}' # platform admins

The daily limits live in the ConfigMap infrared-ai-limits in Infrared's namespace. The label of the Gemini key in use comes from the annotation infrared.darkshift.io/key-label on its Secret; the key itself is never shown.