Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ mkdir -p .bob && ln -s ../.agents/skills .bob/skills
- Prefer primitives over classes
- **Friendly Dependency Errors**: Wraps optional backend imports in `try/except ImportError` with a helpful message (e.g., "Please pip install mellea[hf]"). See `mellea/stdlib/session.py` for examples.
- **CLI command docstrings**: Typer command functions in `cli/` follow an enriched convention with `Prerequisites:` and `See Also:` sections — these feed the auto-generated CLI reference page. See [`docs/CONTRIBUTING_DOCS.md`](docs/CONTRIBUTING_DOCS.md) for the full pattern. Regenerate after changes: `uv run poe clidocs`. Test the generator: `uv run pytest tooling/docs-autogen/test_cli_reference.py -v`. Full pipeline docs: [`tooling/docs-autogen/README.md`](tooling/docs-autogen/README.md).
- **Backend telemetry fields**: All backends must populate `mot.generation.usage` (dict with `prompt_tokens`, `completion_tokens`, `total_tokens`), `mot.generation.model` (str), and `mot.generation.provider` (str) in their `post_processing()` method. These fields live on `mot.generation`, a `GenerationMetadata` dataclass. `mot.generation.streaming` (bool) is set in `astream()`; `mot.generation.ttfb_ms` (float | None) is stamped at the provider's first-chunk receipt inside `send_to_queue()` — backends set neither manually. Metrics are automatically recorded by `TokenMetricsPlugin`, `LatencyMetricsPlugin`, and `ErrorMetricsPlugin` — don't add manual `record_token_usage_metrics()`, `record_request_duration()`, or `record_error()` calls.
- **Backend telemetry fields**: All backends must populate `mot.generation.usage` (dict with `prompt_tokens`, `completion_tokens`, `total_tokens`), `mot.generation.model` (str), and `mot.generation.provider` (str) in their `post_processing()` method. These fields live on `mot.generation`, a `GenerationMetadata` dataclass. `mot.generation.streaming` (bool) is set in `astream()`; `mot.generation.ttfb_ms` (float | None) is stamped at the provider's first-chunk receipt inside `send_to_queue()` — backends set neither manually. `mot.generation.prompt_template` (str | None) and `mot.generation.prompt_template_variables` (dict | None) are set at render time in `_generate_from_context` (not `post_processing`), beside the early `model`/`provider` assignment, via `self.formatter.prompt_template_for(action)` — only on the chat path (`generate_from_raw` has no single unified template). Metrics are automatically recorded by `TokenMetricsPlugin`, `LatencyMetricsPlugin`, and `ErrorMetricsPlugin` — don't add manual `record_token_usage_metrics()`, `record_request_duration()`, or `record_error()` calls.
- **Adding or editing telemetry (spans/metrics)**: Telemetry is emitted by **hook-fired plugins**, not direct calls. Core fires lifecycle hooks; a `*TracingPlugin` in `mellea/telemetry/tracing_plugins.py` emits spans and a `*MetricsPlugin` in `mellea/telemetry/metrics_plugins.py` emits metrics, both subscribing to those hooks. Core does **not** call `start_*_span`/`finish_*_span` from `mellea/telemetry/tracing.py` directly (the only exception is sync code that can't fire paired hooks). Matching helper names in `tracing.py` is not enough — read the plugins, and use the existing one whose span shape matches yours as the template. Emitting a span needs a start/pre hook to open it and a matching end/post hook to close it; a hook that only fires at completion, with no paired opener, can feed a metric but cannot anchor a span.

## 6. Commits & Hooks
Expand Down
2 changes: 2 additions & 0 deletions docs/docs/advanced/custom-components.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,8 @@ class FeedbackForm(Component[dict[str, str]]):
return json.loads(raw.strip())
```

> **Tip:** Keep `format_for_llm` side-effect-free — it may be called more than once per generation.

Pass the component to `m.act()` to get a result:

```python
Expand Down
8 changes: 8 additions & 0 deletions docs/docs/observability/tracing.md
Original file line number Diff line number Diff line change
Expand Up @@ -231,6 +231,14 @@ Backend spans cover individual LLM API calls. They follow the
| `gen_ai.response.id` | Response identifier from the backend |
| `gen_ai.response.time_to_first_chunk` | Time to first chunk in seconds; streaming requests only |

Backend spans also carry the prompt-template attributes from the
[OpenInference semantic conventions](https://arize-ai.github.io/openinference/spec/semantic_conventions.html):

| Attribute | Description |
| --------- | ----------- |
| `llm.prompt_template.template` | Source of the prompt template rendered for the action; `chat` generations only |
| `llm.prompt_template.variables` | The template's post-substitution variables, truncated to 500 characters; recorded only when `MELLEA_TRACES_CONTENT=true` |

Mellea also adds context-specific attributes to backend spans:

| Attribute | Description |
Expand Down
8 changes: 8 additions & 0 deletions mellea/backends/huggingface.py
Original file line number Diff line number Diff line change
Expand Up @@ -1720,6 +1720,10 @@ async def _generate_from_context_with_kv_cache(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down Expand Up @@ -1912,6 +1916,10 @@ async def _generate_from_context_standard(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down
4 changes: 4 additions & 0 deletions mellea/backends/litellm.py
Original file line number Diff line number Diff line change
Expand Up @@ -519,6 +519,10 @@ async def _generate_from_chat_context_standard(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down
4 changes: 4 additions & 0 deletions mellea/backends/ollama.py
Original file line number Diff line number Diff line change
Expand Up @@ -1199,6 +1199,10 @@ async def generate_from_chat_context(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down
4 changes: 4 additions & 0 deletions mellea/backends/openai.py
Original file line number Diff line number Diff line change
Expand Up @@ -1500,6 +1500,10 @@ async def _generate_from_chat_context_standard(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down
4 changes: 4 additions & 0 deletions mellea/backends/watsonx.py
Original file line number Diff line number Diff line change
Expand Up @@ -572,6 +572,10 @@ async def generate_from_chat_context(
# Set model/provider early so they are available in the error path
output.generation.model = self._model_id
output.generation.provider = self._provider
(
output.generation.prompt_template,
output.generation.prompt_template_variables,
) = self.formatter.prompt_template_for(action)

try:
# To support lazy computation, will need to remove this create_task and store just the unexecuted coroutine.
Expand Down
19 changes: 19 additions & 0 deletions mellea/core/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -885,6 +885,8 @@ def parts(self) -> list[Span]:
def format_for_llm(self) -> TemplateRepresentation | str:
"""Formats the `Component` into a `TemplateRepresentation` or plain string for LLM consumption.

Implementations should be side-effect-free, as this may be called more than once per generation.

Returns:
TemplateRepresentation | str: A structured `TemplateRepresentation` (for components
with tools, fields, or templates) or a plain string for simple components.
Expand Down Expand Up @@ -943,6 +945,8 @@ class GenerationMetadata:
usage: Token usage dict with 'prompt_tokens', 'completion_tokens', 'total_tokens'.
model: Requested model identifier.
provider: Provider name (e.g. 'openai', 'ollama', 'huggingface', 'watsonx').
prompt_template: Source text of the template rendered for the action, if it has one.
prompt_template_variables: Post-substitution variables the template was rendered with.
ttfb_ms: Time to first token in milliseconds; None for non-streaming.
streaming: Whether this generation used streaming mode.
response_model: Model identifier reported on the response; may differ from the requested model.
Expand Down Expand Up @@ -971,6 +975,21 @@ class GenerationMetadata:
provider: str | None = None
"""Provider name (e.g. 'openai', 'ollama', 'huggingface', 'watsonx')."""

prompt_template: str | None = None
"""Source text of the template rendered for the action.

Set at generation time by template-rendering backends. `None` when the action
has no template (e.g. a `CBlock`, or a component whose `format_for_llm`
returns a plain string).
"""

prompt_template_variables: dict[str, Any] | None = None
"""Post-substitution variables the template was rendered with.

The stringified template arguments used at generation time; may contain user
data. `None` when the action has no template.
"""

ttfb_ms: float | None = None
"""Time to first token in milliseconds.

Expand Down
17 changes: 17 additions & 0 deletions mellea/formatters/chat_formatter.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,13 +11,30 @@
completion endpoint.
"""

from typing import Any

from ..core import Component, Formatter, ModelOutputThunk, Span, TemplateRepresentation
from ..stdlib.components.chat import Message, message_from_template_representation


class ChatFormatter(Formatter):
"""Formatter used by Legacy backends to format Contexts as Messages."""

def prompt_template_for(
self, action: Span
) -> tuple[str | None, dict[str, Any] | None]:
"""Return the template source and rendered variables for an action.

Chat formatters render no templates and return `(None, None)`.

Args:
action (Span): The component or block a backend generates from.

Returns:
`(None, None)`.
"""
return None, None

def to_chat_messages(self, cs: list[Span]) -> list[Message]:
"""Convert a linearized chat history into a list of chat messages.

Expand Down
58 changes: 50 additions & 8 deletions mellea/formatters/template_formatter.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@
import re
import sys
from collections.abc import Iterable, Mapping
from dataclasses import fields
from dataclasses import fields, replace
from typing import Any

import jinja2
Expand Down Expand Up @@ -179,6 +179,42 @@ def print(self, c: Span) -> str:
"""
return self._stringify(c)

def prompt_template_for(
self, action: Span
) -> tuple[str | None, dict[str, Any] | None]:
"""Return the template source and rendered variables for an action.

Recovers the Jinja source and the post-substitution variables the formatter
would render `action` with, without producing the rendered string.
Best-effort and never raises: returns `(None, None)` when `action` has no
template (a `CBlock`, a `ModelOutputThunk`, or a component whose
`format_for_llm` returns a plain string) or when lookup or rendering fails.

Args:
action: The component or block a backend generates from.

Returns:
A `(template_source, variables)` tuple, where `variables` maps each
template argument to its stringified value, or `(None, None)` when no
template applies.
"""
if not isinstance(action, Component):
return None, None
try:
representation = action.format_for_llm()
if not isinstance(representation, TemplateRepresentation):
return None, None
if representation.obj is None:
representation = replace(representation, obj=action)
template = self._load_template(representation)
source = self._template_source(template, representation)
variables = {
key: self._stringify(val) for key, val in representation.args.items()
}
except Exception:
return None, None
return source, variables

def _load_template(self, repr: TemplateRepresentation) -> jinja2.Template:
"""This method makes an attempt at auto-loading a Template for the Component.

Expand Down Expand Up @@ -283,19 +319,25 @@ def _get_expected_variables(
) -> set[str]:
"""Return the set of externally-expected variable names for a template."""
try:
if representation.template:
source = representation.template
else:
loader = template.environment.loader
assert loader is not None
assert template.name is not None
source, _, _ = loader.get_source(template.environment, template.name)
source = self._template_source(template, representation)
ast = template.environment.parse(source)
return jinja2.meta.find_undeclared_variables(ast)
except Exception:
# Return an empty set if something goes wrong here.
return set()

def _template_source(
self, template: jinja2.Template, representation: TemplateRepresentation
) -> str:
"""Return the Jinja source text for a loaded template."""
if representation.template:
return representation.template
loader = template.environment.loader
assert loader is not None
assert template.name is not None
source, _, _ = loader.get_source(template.environment, template.name)
return source

def _get_template(self, root_path: str, template_name: str) -> str:
"""Attempts to walk the provided directory structure to find the best matching template.

Expand Down
30 changes: 24 additions & 6 deletions mellea/telemetry/_tracing_helpers.py
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,24 @@ def set_mellea_attrs(span: Any, mot: Any) -> None:
span.set_attribute("mellea.request.context_size", len(ctx) if ctx else 0)


def set_prompt_template_attrs(span: Any, gen: GenerationMetadata) -> None:
"""Emit OpenInference `llm.prompt_template.*` attributes from `gen`.

The template text is always emitted; the variables are recorded only when
content capture is enabled (they may contain user data) and are length-bounded.

Args:
span: The span object.
gen: The generation metadata carrying the prompt template and variables.
"""
set_attribute_safe(span, "llm.prompt_template.template", gen.prompt_template)
set_attribute_safe(
span,
"llm.prompt_template.variables",
get_capture_content_value(_serialize_mapping(gen.prompt_template_variables)),
)


def set_conversation_id(span: Any) -> None:
"""Emit `gen_ai.conversation.id` from the current telemetry context, if set."""
from mellea.telemetry.context import get_session_id
Expand All @@ -171,17 +189,17 @@ def set_conversation_id(span: Any) -> None:
span.set_attribute("gen_ai.conversation.id", session_id)


def _serialize_arguments(arguments: Mapping[str, Any] | None) -> str | None:
"""Return a stable, key-sorted JSON string of tool arguments, or `None`.
def _serialize_mapping(mapping: Mapping[str, Any] | None) -> str | None:
"""Return a stable, key-sorted JSON string of a mapping, or `None`.

Falls back to `str()` for values JSON cannot serialize.
"""
if not arguments:
if not mapping:
return None
try:
return json.dumps(arguments, sort_keys=True, default=str)
return json.dumps(mapping, sort_keys=True, default=str)
except Exception:
return str(arguments)
return str(mapping)


def _tool_schema_attrs(tool_call: Any) -> tuple[str | None, str | None]:
Expand Down Expand Up @@ -221,7 +239,7 @@ def get_tool_call_attrs(tool_call: Any) -> dict[str, Any]:
Span attributes keyed by attribute name.
"""
tool_type, tool_description = _tool_schema_attrs(tool_call)
serialized = _serialize_arguments(getattr(tool_call, "args", None))
serialized = _serialize_mapping(getattr(tool_call, "args", None))
arguments_hash = (
hashlib.sha256(serialized.encode("utf-8")).hexdigest()[:16]
if serialized is not None
Expand Down
3 changes: 3 additions & 0 deletions mellea/telemetry/tracing.py
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,7 @@
set_attribute_safe,
set_conversation_id,
set_mellea_attrs,
set_prompt_template_attrs,
set_request_attrs,
set_response_attrs,
set_usage_attrs,
Expand Down Expand Up @@ -380,6 +381,7 @@ def finish_backend_span_success(
if gen is not None:
set_request_attrs(span, gen, operation)
set_response_attrs(span, gen)
set_prompt_template_attrs(span, gen)
set_usage_attrs(span, usage)
if mot is not None:
set_mellea_attrs(span, mot)
Expand Down Expand Up @@ -411,6 +413,7 @@ def finish_backend_span_error(
try:
if gen is not None:
set_request_attrs(span, gen, operation)
set_prompt_template_attrs(span, gen)
span.record_exception(exception)
span.set_status(trace.Status(trace.StatusCode.ERROR, str(exception)))
span.set_attribute("error.type", type(exception).__name__)
Expand Down
53 changes: 53 additions & 0 deletions test/formatters/test_template_formatter.py
Original file line number Diff line number Diff line change
Expand Up @@ -223,6 +223,59 @@ def format_for_llm(self) -> str:
)


def test_prompt_template_for_inline_template(tf: TemplateFormatter):
"""`prompt_template_for` returns the inline template source and its variables."""

class _TemplInstruction(Instruction):
def format_for_llm(self) -> TemplateRepresentation:
instr_args = super().format_for_llm().args
return TemplateRepresentation(
obj=self, args=instr_args, template="""{{description}}"""
)

c = _TemplInstruction("description text", ["req1"])
template, variables = tf.prompt_template_for(c)
assert template == "{{description}}"
assert variables is not None
assert variables["description"] == "description text"


def test_prompt_template_for_file_template(tf: TemplateFormatter, instr: Instruction):
"""`prompt_template_for` recovers a file-backed template source and its variables."""
template, variables = tf.prompt_template_for(instr)
assert isinstance(template, str) and template != ""
assert variables is not None
# The instruction's description is one of the rendered variables.
assert any("Write an essay about LLMs." in str(v) for v in variables.values()), (
"rendered variables should include the instruction description"
)


def test_prompt_template_for_cblock_returns_none(tf: TemplateFormatter):
"""A `CBlock` action has no template, so `prompt_template_for` returns `(None, None)`."""
assert tf.prompt_template_for(CBlock(value="raw")) == (None, None)


def test_prompt_template_for_string_repr_returns_none(tf: TemplateFormatter):
"""A component whose `format_for_llm` returns a plain string has no template."""

class _StringRepr(MObject):
def format_for_llm(self) -> str:
return "plain string"

assert tf.prompt_template_for(_StringRepr()) == (None, None)


def test_prompt_template_for_swallows_errors(tf: TemplateFormatter):
"""An internal failure yields `(None, None)`; the method never raises."""

class _RaisingRepr(MObject):
def format_for_llm(self) -> TemplateRepresentation:
raise RuntimeError("boom")

assert tf.prompt_template_for(_RaisingRepr()) == (None, None)


def test_user_path(instr: Instruction):
"""Ensures that paths with no templates don't prevent default template lookups.

Expand Down
Loading
Loading