Skip to content

Observe · Trace · Route · Drift · Retrain

The control plane
for platform teams

Put agents and owned models behind one plane. See every call, route by policy, catch quality drift, and retrain. A closed loop, not a dashboard that stops at charts.

Self-host in minutes. No credit card.

Built for teams running agents and owned models

  • Platform engineering
  • ML ops
  • Agent builders
  • Inference operators
  • AI product teams

The gap

Observability without a loop leaves you staring at charts

Gateways move traffic. Trace tools show what happened. Neither closes the loop when quality slips on agents you operate and models you own.

  • Dashboards stop at “what happened”

    You see latency and cost, then open a ticket. No policy route, no drift signal, no path to the next fine-tune.

  • Gateways stop at “where it went”

    One key, many providers. Useful. Still missing quality drift and a retrain path for owned models.

  • Inferix closes observe → route → retrain

    See the run, send the next call by policy, catch drift early, and ship an improved model — one plane for agents and inference.

The loop

Observe → Trace → Route → Drift → Retrain. Five products. One spine.

Each step is a clear job, not another jargon stack. Scroll the spine, open the product.

  1. 01

    Observe: See every call

    Latency, cost, tokens, and errors for agents and models. One dashboard, not five tools.

    • p95 / p99 latency and error rate by model, agent, and tenant
    • Cost and token volume attributed on every call
    • Anomaly signals when spend or latency leave the baseline
    • One panel for owned models and providers together

    LensAI

    LensAI — live traffic

    tenant · acme-platform

    Example UIHealthy
    312 msp95 latency
    $48.20Cost / 1h
    0.4%Error rate
    1,840Calls / min
    • Top modelowned/general-llm
    • Hot agentsupport-refund
    • Alertslatency-spikecost-budget-80%
  2. 02

    Trace: Follow the whole agent

    Tools and model steps in one graph. Debug a bad run without guessing which hop failed.

    • Parent/child spans for tools, models, memory, and HITL waits
    • Business success vs HTTP 200: know if the job actually worked
    • Cost and exclusive time on every hop
    • Replay a single bad run from the waterfall

    TraceForge

    TraceForge — run waterfalls

    trace · tr_8f2a…c91

    Example UIFailed hop
    14Spans
    2.4 sDuration
    6Tools
    $0.018Cost
    • PathTool callModel hopFinal response
    • Failedpayments.refund
    • Taxonomyllmexternalmemory
  3. 03

    Route: Send each call by policy

    Cheap owned SLM for simple asks. Strong path or provider when you need it. Policy decides, not a spreadsheet.

    • Default, fallback, and intent-based routes in one policy
    • Cheap SLM vs strong owned vs provider hard path
    • Budgets, cache hits, and capability deny matrices
    • Sticky routing mid-task so agents don’t flip models mid-run

    RouteIQ

    RouteIQ — policy map

    policy · cx-default-v3

    Example UIActive
    68%Cheap path
    24%Strong path
    8%Provider
    31%Cache hit
    • Defaultowned/slm-support
    • Escalateowned/general-llmopenai/gpt-4o
    • Denyheal-tools on slm-apiheal
  4. 04

    Drift: Catch quality before tickets

    Alert when answers go bad versus a teacher or golden set. Know before users complain.

    • Golden-set and teacher-gap scores by intent and locale
    • Prompt, tool, provider, and retrieval drift types
    • Deploy markers so you see which change moved quality
    • Alert windows tuned for operators, not noise floods

    DriftWatch

    DriftWatch — quality

    slice · refund_request · en-IN

    Example UIDrift alert
    0.81 ↓Quality
    +12%Teacher gap
    86%Golden pass
    3.2 minTTD
    • Typeprompt-drift
    • Baselinegolden-v12claude-teacher
    • Actionpage-oncallopen FineForge
  5. 05

    Retrain: Ship the next model safely

    Turn drift into the next fine-tune. Promote or roll back without guessing.

    • Jobs from DriftWatch failures and TraceForge bad runs
    • Prompt, adapter, and model artifacts in one promote bundle
    • Canary → promote → rollback with eval gates
    • Owned model registry stays in sync with RouteIQ

    FineForge

    FineForge — ship loop

    job · ff_retrain_19

    Example UICanary
    v19Candidate
    94%Eval pass
    5%Canary
    1-clickRollback
    • Bundleprompt-v19adapter-v3route-patch
    • Gatesgoldenteacher-gaplatency
    • NextPromote versionSafe rollback

What you get

Five products. One loop.

LensAI · TraceForge · RouteIQ · DriftWatch · FineForge. Observe, route, catch drift, retrain.

Product details →

How it works

Install → instrument → see → route → catch drift → retrain

Six steps from first install to a better model in production.

  1. 1

    Install

    One Docker (or Helm) path. Plane up on your laptop or in your VPC.

  2. 2

    Instrument

    Point agent and app traffic at Inferix. Spans and metrics start flowing.

  3. 3

    See

    LensAI shows cost and latency. TraceForge follows the full agent run.

  4. 4

    Route

    RouteIQ sends each call to the right owned model or provider by policy.

  5. 5

    Catch drift

    DriftWatch alerts when quality slips versus teacher or golden set.

  6. 6

    Retrain

    FineForge turns that signal into the next fine-tune you can promote.

Fits next to the stack you already run

  • OpenAI API
  • Anthropic
  • vLLM
  • LiteLLM
  • LangChain
  • OpenTelemetry
  • Kafka
  • ClickHouse
  • Kubernetes

Integration names, not customer logos. Connect via API, SDK, or OpenTelemetry. Details in docs.

04

Enterprise scale.
These are our targets.

Same shape of ambition as mature LLM platforms: volume, latency, deploy speed, and one plane for agents and models. Labeled targets until published production numbers replace them.

Platform targets · not live cloud stats

  • Target

    Billions+

    Observations / month · platform scale

  • Target

    <100 ms

    Ingest P99 · accept + WAL + enqueue

  • Target

    Minutes

    Self-host to first dashboard · your laptop or VPC

  • Target

    1 plane

    Agents + owned models · observe · route · drift · retrain

What “scale” means for Inferix

  • Observations Design for billions of inference and agent observations per month across the platform (Langfuse-class monthly volume), not a vanity events/min headline
  • Ingest P99 <100 ms at accept + WAL + enqueue so the plane stays off the critical path of your apps
  • RouteIQ Policy decision overhead <2 ms when the router is measured
  • DriftWatch Catch quality drift vs teacher / golden set before tickets pile up (time-to-detect target <5 min)
  • Attribution Cost and latency by tenant, model, and agent. Fail closed when tenant is missing

When a target is proven in the open repos, the label upgrades from Target to a measured number. We do not invent customer counts or Fortune logos.

Where Inferix sits

Next to your gateway and your traces — not instead of them

  • vs LiteLLM / Portkey Gateways give you one key and many providers. Inferix is the plane around that layer: see calls, route with policy, catch drift, and retrain.
  • vs Langfuse / Helicone Observability shows what apps did. Inferix adds operator routing, drift alerts, and a fine-tune loop for models you run.
  • vs Prefactor Prefactor evaluates agents and can hold / block at runtime. Inferix is adjacent: closed-loop observe → route → drift → retrain for agents and owned models. We do not ship runtime hold/block as the product.

FAQ

Straight answers

What is Inferix?
The control plane for agents and owned models. One place to observe traffic, route by policy, catch quality drift, and retrain. Not another dashboard that stops at charts.
How is this different from Langfuse or Helicone?
Those tools excel at traces and app observability. Inferix keeps that visibility and adds operator routing (RouteIQ), drift alerts (DriftWatch), and a retrain loop (FineForge) for models you run.
How is this different from LiteLLM or Portkey?
Gateways give you one key and many providers. Inferix sits around that layer: see every call, route with policy, catch drift, and ship the next model. It does not replace the gateway.
Do you hold or block agent actions at runtime?
No. That is a different category (runtime evaluation + enforcement). Inferix’s loop is observe → route → drift → retrain. We do not ship hold / block / HITL enforcement as the product.
What does “control plane” mean here?
A single operator surface for agents and inference: LensAI (observe), TraceForge (trace), RouteIQ (route), DriftWatch (drift), FineForge (retrain). Connect once; improve what drifts.
Can I self-host?
Yes. Start local in minutes from quick start. No credit card. Bring your own VPC when you are ready.

Run it locally in a few minutes

One Docker command. Then add routing and drift when you need them.