FFluxen
Product

Everything you need to run AI traffic efficiently

One self-hosted deployment: gateway, application intelligence, cost intelligence, four detectors, simulation, control, and measurement.

Gateway

A single OpenAI-compatible ingress in front of OpenAI, Google Gemini, and Ollama. Application-scoped API keys, streaming and non-streaming, usage extraction, integer micro-USD cost calculation, exact caching, rate limiting, budgets, and percentage-based routing.

Application Efficiency Profile

Applications, not raw requests, are the unit of optimization in Fluxen. Every application gets an Efficiency Score — five weighted components (model, token, cache, traffic stability, cost efficiency) — plus its own opportunities, policies, and request history.

Four detectors

Model cost

Finds traffic on an expensive model that a cheaper, catalog-declared alternative could plausibly handle, based on real token-size and feature-profile evidence.

Repeated request

Finds exact-duplicate request traffic that a TTL cache would have served for free, with a savings curve across four candidate TTLs.

Token efficiency

Finds sustained drift in input or output token usage per request, with a change-point estimate for when the drift began.

Traffic anomaly

Finds statistically unusual request volume, cost, token usage, or error rate against an application's own seasonal baseline — a stability signal, not a savings claim.

Simulate, control, measure

Every recommendation can be replayed against real historical traffic before anything changes production. Applying a change is one human-confirmed action across five controls — routing, caching, budgets, rate limits, model restrictions. Every applied change gets an honest before/after verdict, with revert always one click away.