gin

Gin by Winding Labs

Pay 60% less
for the same AI.

A dead simple, drop-in wrapper for OpenRouter and every AI API.

See a sample email

Free to watch. Your keys, your bills. Works with Claude Code, Codex, Cursor, OpenCode and Gemini CLI.

src/ai.ts+1 −1
const ai = createOpenAI({  baseURL: 'https://openrouter.ai/api/v1',- fetch: fetch,+ fetch: gin({ useCase: 'review' }),})

In our own apps

Measured on IonWarp production traffic, blind-graded on real code reviews.

66%

cheaper performance reviews

GLM 5.3 Flash, equal quality over 21 paired runs
82%

cheaper follow-up turns

Session affinity keeps the provider cache warm
15×

cheaper security reviews

$0.166 → $0.011 per run, no findings lost

How it works

Four steps. Two of them are yours.

You run one command and press one button. Your coding agent does the wiring, and the email does the math.

  1. Install the skill

    Adds a Gin skill to the coding agents you already use and signs you in with GitHub.

    npx @winding-labs/gin
  2. Your agent wraps the calls

    It finds each place your app calls AI and labels it with a use case. One line each.

    fetch: gin({ useCase: 'review' })
  3. Read the morning email

    Spend across every router, spend per use case, and what routing would save, shown on one of your own calls.

    Tumbly: CarPlay Generation · $8.40
  4. Press “Save with Gin”

    Calls go through Gin’s Cloudflare edge to your own OpenRouter key, on a cheaper model only where your replays held.

    falls back to your model on any error

The daily email

Your AI bill, explained before coffee.

Every morning, to you and your teammates. Free forever, even if you never route a single call.

FromGin <daily@trygin.ai> Toteam@windinglabs.com
Winding Labs: $38.12 yesterday, $417/mo you could save

Spend yesterday · all routers

$38.12
▲ 15% vs 7-day average
OpenRouter $31.40OpenAI $4.92Anthropic $1.80
By use caseSpend7-dayCould save
IonWarp: Code Reviewdeepseek-v4.1-flash$12.84▲ 9%keep
Tumbly: CarPlay Generationclaude-sonnet-5.5$8.40▲ 31%$3.02
IonWarp: Performance Reviewdeepseek-v4.1-flash$6.92▼ 4%$4.57
IonWarp: Security Reviewgpt-5.6-terra$4.31▲ 2%$4.03
Tumbly: Live Q&Aclaude-haiku-5.5$2.95▼ 6%keep
IonWarp: Checkerdeepseek-v4.1-flash$2.70▲ 12%$2.29

Routing would have saved $13.91 yesterday.

That is about $417 a month across 4 use cases. Code Review and Live Q&A stay on your models: cheaper ones missed findings in replays.

One of your calls, replayed · IonWarp: Performance Review

Review this diff for performance regressions. src/feed/rank.ts (+48 −12)
deepseek-v4.1-flash$0.081

N+1 query: getAuthor() inside the map at rank.ts:41 runs once per post, about 200 queries per page. Batch with getAuthors(ids).

yours
glm-5.3-flash$0.027 · −66%

rank.ts:41 calls getAuthor() per post inside map: an N+1, roughly 200 round trips per page. Prefetch once with getAuthors(ids).

same finding · equal in 21 of 21 replays
Save $417/mo with Gin

Routes through your OpenRouter key. Turn it off per use case, any time.

  • Every router, one numberOpenRouter, OpenAI, Anthropic and the rest, added up, with the trend against last week.
  • Named by use case“Tumbly: CarPlay Generation”, not “sk-…3f9a”. Your agent labels calls when it wraps them.
  • Proof, not promisesSavings come with a real call of yours, answered by the cheaper model, side by side.
  • Honest about “keep”If a cheaper model missed things in replays, the email says so and leaves that use case alone.
  • Whole teamTeammates in the workspace get the email too. One button, anyone can press it.

The dashboard

Every call. Replay any of them.

Activity and logs in the style you know from OpenRouter, across all your routers. Press Replay to run any real call on another model and compare.

Winding Labs / all repos
Spend today$14.62
Requests18,204
Routed41%
Saved today$5.7128% off
Recent AI calls
TimeUse caseModelTokensCostLatencyActions
09:41:12IonWarp: Performance Reviewglm-5.3-flashrouted41.2k / 1.9k$0.0273.1s
09:41:08IonWarp: Code Reviewdeepseek-v4.1-flash62.4k / 3.1k$0.24118.4s
09:40:55Tumbly: CarPlay Generationclaude-sonnet-5.52.1k / 1.4k$0.0276.2s
09:40:51Tumbly: Live Q&Aclaude-haiku-5.51.3k / 180$0.0020.9s
09:40:33IonWarp: Checkerglm-5.3-flashrouted8.8k / 0.6k$0.0041.8s

Workspaces and repos

One workspace per team, repos underneath. Spend rolls up; use cases stay separate.

One-click replay

Send any logged call to other models and see the answers and the cost next to yours.

Routing you can read

Every routed call says which model served it and why. Turn any use case off with one switch.

Pricing

Free to see. A small cut when we save you money.

You keep your own provider keys and your own provider bills. Gin never resells tokens.

Observe

$0

Forever. Every call, every router, every teammate.

  • Spend by use case and router
  • Daily email for the whole workspace
  • Activity and logs
  • Unlimited repos

Replay

$10 free

In replay credits to start. Top up when you want more.

  • One-click replay from any log
  • Paired replays per use case
  • Side-by-side answers and cost
  • Graded for “same quality” or not

Route

5%

Of routed spend. Your first $50 of routed spend is free.

  • Cheaper model only where replays held
  • Falls back to your model on any error
  • Your OpenRouter key, your bill
  • Off per use case, any time

The 5% is the only thing Gin charges for, and only on calls that go through routing. Observing and the daily email never cost anything.

Privacy and security

We keep as little as we can, for as short as we can.

Prompts and responses are what make replays possible. Here is exactly how we treat them. Read the privacy policy.

Encrypted per workspace

Request and response bodies are encrypted with AES-GCM using a key unique to your workspace.

Deleted after 7 days

Bodies expire automatically after 7 days. Spend and token counts stay so your charts keep working.

Metadata-only mode

One setting and bodies are dropped on arrival. You still get spend, trends and the daily email.

Keys never stored

Your provider key is forwarded to the provider on each routed call and never written down. The SDK never sends it on observed calls.

FAQ

Questions, answered plainly.

Does it add latency?

Observing adds zero. Your call goes straight to your provider, and Gin’s event is sent after the response, outside your request path.

Routing adds one edge hop on Cloudflare’s network, typically about 10–30 ms. The cheaper model is often faster than that.

What happens if Gin is down?

Your calls go direct. If the edge can’t be reached, times out connecting, or reports its own error, the SDK calls your provider exactly as it did before Gin. Any error on a routed model falls back to the model you asked for.

Which providers work?

OpenRouter, OpenAI, Anthropic, Google Gemini, Together, Groq, Fireworks, DeepSeek, xAI and Mistral. Observing works for any HTTP API you call through the wrapped fetch. Routing uses your own OpenRouter key so Gin can pick across models.

Which languages?

The SDK is TypeScript and wraps fetch, so it works with the OpenAI SDK, the Vercel AI SDK, or plain fetch, on Node, Bun, Deno and Cloudflare Workers.

Anything else (Python, Go, Ruby) can point an OpenAI-compatible client at https://trygin.ai/v1/openrouter with an x-gin-key header.

Will Gin switch my models without asking?

No. Nothing is routed until someone presses “Save with Gin”, and then only for use cases where replays of your own calls held quality. Everything else stays on the model you chose.

What does my coding agent actually change?

It adds the @winding-labs/gin package and changes the line where your app passes fetch to its AI client, adding a use-case label. You review the diff like any other.

See tomorrow morning what your AI costs.

Sign in with GitHub