Long cache TTL for Claude models

Feature request for product/service

AI Models

Describe the request

I’ve heard that Anthropic offers one-hour cache retention for an increased inference cost. I assume that, given my usage pattern, for me this would be cheaper than the current behavior - I’d like to see a toggle for this in the model selection menu in Chat.

For example, here less than 190k tokens

Hey @Artemonim!

Interestingly, we ran an experiment where longer Claude cache retention was evaluated, and due to the cost of cache writes, this didn’t reduce costs on average!

Individual users might benefit based on their usage patterns, and it remains something for us to consider!

Your team recently conducted an experiment on rewriting SQL in Rust — In such cases, the leading agent can wait a long time for a response from subagents, and in such cases, OpenAI models benefit in cost simply by storing the cache 6 times longer.