I’ve heard that Anthropic offers one-hour cache retention for an increased inference cost. I assume that, given my usage pattern, for me this would be cheaper than the current behavior - I’d like to see a toggle for this in the model selection menu in Chat.
Interestingly, we ran an experiment where longer Claude cache retention was evaluated, and due to the cost of cache writes, this didn’t reduce costs on average!
Individual users might benefit based on their usage patterns, and it remains something for us to consider!
Your team recently conducted an experiment on rewriting SQL in Rust — In such cases, the leading agent can wait a long time for a response from subagents, and in such cases, OpenAI models benefit in cost simply by storing the cache 6 times longer.