Skip to content

Observability, tracing and monitoring #24

Description

@Glockner00

Part of #11

Question

What do we need to see about a turn in production, and how do we see it?

app/app_utils/telemetry.py and src/deadlock_coach/telemetry.py write local event traces to artifacts/telemetry/. That is enough for one developer debugging locally and not enough for a public product.

Decide:

  • what one turn's trace must contain: tool calls, retrieved evidence, model, tokens, cost, latency, and the provenance chain behind each claim
  • where traces go, retained how long, and what that implies once real users' questions are in them
  • what is monitored continuously versus inspected on demand
  • what constitutes an incident worth alerting on: groundedness failures, tool errors, cost spikes, latency
  • how a bad answer a user complains about is traced back to its cause
  • how this connects to the eval harness, so production failures feed golden datasets

Depends on the topology: what a trace looks like differs sharply between one agent and several.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions