Inferrail: Self-hosted LLM gateway to track your Cursor/AI agent costs (Payload-free)

I kept running into a surprisingly simple question while building with AI:

What did that piece of work actually cost?

Not the monthly bill. Not total token usage. The actual job.

So I built Inferrail.

It turns model calls into something closer to:

20 model calls → 1 job → $0.43

It tracks token usage and cost across the work, while keeping prompts and responses out of the receipts.

It’s self-hosted, open source, and still early. I’m not looking to sell anyone here anything. I’d actually rather have people break it and tell me what’s missing.

GitHub: GitHub - domondi1/inferrail: Self-hosted, OpenAI-compatible LLM gateway that turns every request into a payload-free, attributable cost receipt — no stored prompts or responses. Apache-2.0, zero dependency on any Inferrail-operated service. Also offers a hosted x402 capability, Work Economics, for AI job cost receipts (Base Sepolia testnet). · GitHub

Curious whether other people building heavily with Cursor/agents have the same problem: do you actually know what one completed piece of AI work costs you?