The control plane
for platform teams

Put agents and models behind one plane. See every call, route by policy, and catch quality drift. When quality drops, retrain and roll forward — without guessing.

Self-host in minutes. No credit card.

Built for teams running agents and owned models

  • Platform engineering
  • ML ops
  • Agent builders
  • Inference operators
  • AI product teams

Why Inferix

Give your whole org control of agents and inference — then see, route, and improve every request.

  1. 01

    Observe: See every call

    Latency, cost, tokens, and errors for agents and models — one dashboard, not five tools.

    • p95 / p99 latency and error rate by model, agent, and tenant
    • Cost and token volume attributed on every call
    • Anomaly signals when spend or latency leave the baseline
    • One panel for owned models and providers together

    LensAI

    LensAI — live traffic

    tenant · acme-platform

    Healthy
    312 msp95 latency
    $48.20Cost / 1h
    0.4%Error rate
    1,840Calls / min
    • Top modelowned/general-llm
    • Hot agentsupport-refund
    • Alertslatency-spikecost-budget-80%
  2. 02

    Trace: Follow the whole agent

    Tools and model steps in one graph. Debug a bad run without guessing which hop failed.

    • Parent/child spans for tools, models, memory, and HITL waits
    • Business success vs HTTP 200 — know if the job actually worked
    • Cost and exclusive time on every hop
    • Replay a single bad run from the waterfall

    TraceForge

    TraceForge — run waterfalls

    trace · tr_8f2a…c91

    Failed hop
    14Spans
    2.4 sDuration
    6Tools
    $0.018Cost
    • PathTool callModel hopFinal response
    • Failedpayments.refund
    • Taxonomyllmexternalmemory
  3. 03

    Route: Send each call by policy

    Cheap owned SLM for simple asks. Strong path or provider when you need it. Policy decides — not a spreadsheet.

    • Default, fallback, and intent-based routes in one policy
    • Cheap SLM vs strong owned vs provider hard path
    • Budgets, cache hits, and capability deny matrices
    • Sticky routing mid-task so agents don’t flip models mid-run

    RouteIQ

    RouteIQ — policy map

    policy · cx-default-v3

    Active
    68%Cheap path
    24%Strong path
    8%Provider
    31%Cache hit
    • Defaultowned/slm-support
    • Escalateowned/general-llmopenai/gpt-4o
    • Denyheal-tools on slm-apiheal
  4. 04

    Detect: Catch quality before tickets

    Alert when answers go bad versus a teacher or golden set. Know before users complain.

    • Golden-set and teacher-gap scores by intent and locale
    • Prompt, tool, provider, and retrieval drift types
    • Deploy markers so you see which change moved quality
    • Alert windows tuned for operators — not noise floods

    DriftWatch

    DriftWatch — quality

    slice · refund_request · en-IN

    Drift alert
    0.81 ↓Quality
    +12%Teacher gap
    86%Golden pass
    3.2 minTTD
    • Typeprompt-drift
    • Baselinegolden-v12claude-teacher
    • Actionpage-oncallopen FineForge
  5. 05

    Improve: Retrain and ship safely

    Turn drift into the next fine-tune. Promote or roll back without guessing.

    • Jobs from DriftWatch failures and TraceForge bad runs
    • Prompt, adapter, and model artifacts in one promote bundle
    • Canary → promote → rollback with eval gates
    • Owned model registry stays in sync with RouteIQ

    FineForge

    FineForge — ship loop

    job · ff_retrain_19

    Canary
    v19Candidate
    94%Eval pass
    5%Canary
    1-clickRollback
    • Bundleprompt-v19adapter-v3route-patch
    • Gatesgoldenteacher-gaplatency
    • NextPromote versionSafe rollback

What you get

Five products. One loop.

See traffic. Trace agents. Route by policy. Catch drift. Retrain. Each product is a clear job — not another jargon stack.

Product details →

How it works

Connect once. See everything. Improve what drifts.

Three steps from first request to a better model.

  1. 1

    Connect

    Send agent and app traffic through Inferix. One plane sits in front of your models and providers.

  2. 2

    See & route

    LensAI and TraceForge show each run. RouteIQ picks owned models or providers by policy.

  3. 3

    Detect & retrain

    DriftWatch alerts when quality slips. FineForge turns that into the next fine-tune job.

04

What Inferix measures.
Design targets for the open suite.

These are control-plane KPIs we design for — not published race results. Production numbers ship with the open benchmark suite.

Design targets · not shipped benchmarks

  • <50 ms

    Trace ingest lag · p99 target

  • 100%

    Cost attribution coverage target

  • <5 min

    Drift alert time-to-detect target

  • <2 ms

    Route decision overhead target

Time to find a bad agent run — minutes (lower is better)

  • Inferix traces~2 min
  • Logs only~25 min
  • Spreadsheet opshours+

Illustrative comparison of workflows — not a lab benchmark against other vendors.

Route decision overhead — ms added (lower is better)

  • Inferix RouteIQ · target<2 ms
  • Ad-hoc app if/else10–40 ms
  • Manual triage queuehuman delay

Target vs typical agent-stack approaches. Live measurements publish with the open suite.

Illustrative layout — production benchmarks publish with the open suite. Numbers above are design targets and workflow contrasts, not claimed wins over named gateways.

Where Inferix sits

Next to your gateway and your traces — not instead of them

  • vs LiteLLM / Portkey — gateways give you one key and many providers. Inferix is the plane around that layer: see calls, route with policy, catch drift, and retrain.
  • vs Langfuse / Helicone — observability shows what apps did. Inferix adds operator routing, drift alerts, and a fine-tune loop for models you run.

Run it locally in a few minutes

One Docker command. Then add routing and drift when you need them.