Skip to content

Showcase agents / Incident Investigator

SRE / on-call

Incident Investigator

What problem? Incident bots produce long tool chains with no trace depth visibility — wrong RCA hypotheses look identical to correct ones in logs.

SRE reference: alert → metrics/logs/deploys → RCA hypothesis → draft PR. Long chains with Claude-primary reasoning.

tenant: tenant-incident-opsport :9104scenario: latency_regression
Inferix console view for Incident Investigator

What Inferix observes

Every run dual-writes to LensAI and TraceForge. Switch to this tenant in the console to see live KPIs.

  • TraceForge — deep tool waterfall and chain depth
  • LensAI — token cost per investigation
  • DriftWatch — RCA quality vs golden set
  • FineForge — promote better retrieval prompts

Tools & models

Real SQLite backends, not mocked HTTP stubs.

  • pager.get_alert
  • metrics.query
  • logs.search
  • git.recent_deploys
  • k8s.get_events
  • github.open_draft_pr

Model routing

  • Claude (long context, reasoning)
  • limited cheap path

Try it

Start the stack, then run the golden scenario. Traffic appears in the console within seconds.

1. Start stack + agents

cd inferix && ./run-inferix.sh --full --agents

2. Golden scenario

curl -s -X POST http://localhost:9104/scenarios/latency_regression | jq

Expect HTTP 200 with JSON status and tool steps. Traces appear under this tenant in ~10s.

3. All agents at once

cd inferix && ./scripts/generate-agent-traffic.sh

Expected: JSON with status and tool steps. Then open the tenant console and filter traces by tenant-incident-ops. In-browser runner stays deferred — curl is the verified path.

Console: /console/overview?tenant_id=tenant-incident-ops · API traces: GET /v1/traces?tenant_id=tenant-incident-ops

See all seven agents in the console

Multi-tenant super-dashboard. Filter by tenant, drill into traces and costs.