Define image costs in fallback context estimates
What problem would this solve?
Fallback context estimates are shared by multiple providers and agent interfaces, while supported image representations and vision costs vary.
What would a good outcome look like?
Agree how image-containing conversations should contribute to approximate context management when authoritative usage is unavailable.
Possible approaches
Should fallback accounting use a documented conservative per-image allowance, or provider/model-aware estimates? Specify audience filtering, providers that omit images, top-level and tool-result images, and behavior when dimensions or capabilities are unavailable. Verification should cover both agent loops, legitimate multimodal histories, visible versus UI-only content, and authoritative-usage precedence.
Additional context
This is an approximation and compatibility decision, not a promise to impose a hard provider billing or request-byte limit.
- I have verified this does not duplicate an existing feature request
Do not begin implementation until the issue reaches Ready on the Goose Issues board.
Source: aaif-goose/goose