Skip to content

[openai] Record cache-write and audio token usage - #681

Open
1fanwang wants to merge 10 commits into
open-telemetry:mainfrom
1fanwang:1fannnw/openai-token-details-603
Open

1fanwang wants to merge 10 commits into
open-telemetry:mainfrom
1fanwang:1fannnw/openai-token-details-603

Conversation

@1fanwang

@1fanwang 1fanwang commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Description

Chat Completions telemetry omits cache-write and audio token counts, so users cannot observe those reported usage breakdowns. This records them for normal and streamed calls without changing aggregate input/output totals; streamed usage is applied during cleanup.

Text/image mappings are excluded as requested in review. The SDK schema defines those fields, but I have not verified that Chat Completions returns them.

Fixes #603.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • This change requires a documentation update

How has this been tested?

Testing Done

The real OpenAI 3.13.0 SDK parsed JSON/SSE through a mock HTTP transport on macOS with Python 3.14.7. These synthetic responses check extraction, not provider behavior. No live API call was made.

Save the source below as sdk_probe.py. From the repository root:

uv run tox -e py314-test-instrumentation-genai-openai-latest --notest
before=$(mktemp -d)
git archive d859ebce191f93fe4675bdf354f55b03bdbe468c instrumentation/opentelemetry-instrumentation-genai-openai/src | tar -x -C "$before"
PYTHONPATH="$before/instrumentation/opentelemetry-instrumentation-genai-openai/src" .tox/py314-test-instrumentation-genai-openai-latest/bin/python sdk_probe.py
.tox/py314-test-instrumentation-genai-openai-latest/bin/python sdk_probe.py

Both probe runs exited 0. Before:

{"path": "sync-response", "usage": {"gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "sync-stream", "usage": {"gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "async-response", "usage": {"gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "async-stream", "usage": {"gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}

After:

{"path": "sync-response", "usage": {"gen_ai.usage.audio.input_tokens": 10, "gen_ai.usage.audio.output_tokens": 2, "gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.cache_write.input_tokens": 10, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "sync-stream", "usage": {"gen_ai.usage.audio.input_tokens": 10, "gen_ai.usage.audio.output_tokens": 2, "gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.cache_write.input_tokens": 10, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "async-response", "usage": {"gen_ai.usage.audio.input_tokens": 10, "gen_ai.usage.audio.output_tokens": 2, "gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.cache_write.input_tokens": 10, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
{"path": "async-stream", "usage": {"gen_ai.usage.audio.input_tokens": 10, "gen_ai.usage.audio.output_tokens": 2, "gen_ai.usage.cache_read.input_tokens": 5, "gen_ai.usage.cache_write.input_tokens": 10, "gen_ai.usage.input_tokens": 100, "gen_ai.usage.output_tokens": 20, "gen_ai.usage.reasoning.output_tokens": 3}}
Reproducer source: sdk_probe.py
import asyncio
import json

import httpx2
from openai import AsyncOpenAI, OpenAI

from opentelemetry.instrumentation.genai.openai import OpenAIInstrumentor
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.sdk.trace.sampling import ALWAYS_ON
from opentelemetry.test_util_genai.instrumentor import instrument


def respond(request: httpx2.Request) -> httpx2.Response:
    usage: dict[str, object] = {
        "prompt_tokens": 100, "completion_tokens": 20, "total_tokens": 120,
        "prompt_tokens_details": {
            "cached_tokens": 5, "cache_write_tokens": 10, "audio_tokens": 10,
            "text_tokens": 70, "image_tokens": 20,
        },
        "completion_tokens_details": {
            "audio_tokens": 2, "reasoning_tokens": 3, "text_tokens": 18,
        },
    }
    common: dict[str, object] = {
        "id": "chatcmpl-probe", "created": 1, "model": "gpt-4",
    }
    if json.loads(request.content).get("stream"):
        chunks = [
            {**common, "object": "chat.completion.chunk", "choices": [
                {"index": 0, "delta": {"content": "hello"}, "finish_reason": "stop"}
            ]},
            {**common, "object": "chat.completion.chunk", "choices": [], "usage": usage},
        ]
        return httpx2.Response(
            status_code=200, headers={"content-type": "text/event-stream"},
            content="".join(f"data: {json.dumps(chunk)}\n\n" for chunk in chunks)
            + "data: [DONE]\n\n",
        )
    return httpx2.Response(
        status_code=200,
        json={**common, "object": "chat.completion", "usage": usage, "choices": [
            {"index": 0, "message": {"role": "assistant", "content": "hello"},
             "finish_reason": "stop"}
        ]},
    )


exporter = InMemorySpanExporter()
provider = TracerProvider(sampler=ALWAYS_ON)
provider.add_span_processor(SimpleSpanProcessor(exporter))


def report(path: str) -> None:
    (span,) = exporter.get_finished_spans()
    attributes = {key: value for key, value in span.attributes.items()
                  if key.startswith("gen_ai.usage.")}
    print(json.dumps({"path": path, "usage": attributes}, sort_keys=True))
    exporter.clear()


async def run_async() -> None:
    async with AsyncOpenAI(
        api_key="test", base_url="http://openai.test/v1", max_retries=0,
        http_client=httpx2.AsyncClient(transport=httpx2.MockTransport(respond)),
    ) as client:
        for streaming in (False, True):
            response = await client.chat.completions.create(
                model="gpt-4", messages=[{"role": "user", "content": "hello"}],
                stream=streaming,
                **({"stream_options": {"include_usage": True}} if streaming else {}),
            )
            if streaming:
                async for _ in response:
                    pass
            report(path=f"async-{'stream' if streaming else 'response'}")


with instrument(
    instrumentor=OpenAIInstrumentor(), tracer_provider=provider,
    logger_provider=None, meter_provider=None,
):
    with OpenAI(
        api_key="test", base_url="http://openai.test/v1", max_retries=0,
        http_client=httpx2.Client(transport=httpx2.MockTransport(respond)),
    ) as client:
        for streaming in (False, True):
            response = client.chat.completions.create(
                model="gpt-4", messages=[{"role": "user", "content": "hello"}],
                stream=streaming,
                **({"stream_options": {"include_usage": True}} if streaming else {}),
            )
            if streaming:
                list(response)
            report(path=f"sync-{'stream' if streaming else 'response'}")
    asyncio.run(run_async())
provider.shutdown()
  • Test A: The SDK probe above exports cache-write and audio counts on sync/async responses and streams, with unchanged totals and no text/image attributes.

Local conformance scenarios skipped because Weaver is unavailable.

Checklist

See CONTRIBUTING.md
for the style guide, changelog guidance, and more.

  • Followed the style guidelines of this project
  • Changelog updated if the change requires an entry
  • Unit tests added
  • Documentation updated (changelog; no public API or configuration change)

Signed-off-by: 1fanwang <1fannnw@gmail.com>
@opentelemetry-pr-dashboard

opentelemetry-pr-dashboard Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

Pull request dashboard status

Waiting on the author · refreshed 2026-10-06 06:59 UTC

Respond to 2 review items (e.g. link a commit, explain why not, ask a follow-up):

  • Inline threads: 1
  • Top-level threads: 2
Status above doesn't look right?
  • Just replied or pushed? Anything around or after the refresh time above may not be picked up yet — give it a few minutes.
  • Should this be with reviewers? Comment /dashboard route:reviewers to route it to them.
  • Anything wrong — including the routing? Report it with what you expected; it helps us improve the dashboard.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

Add integer type assertions for exported usage attributes in the regression tests.

Pull request overview

Adds OpenAI Chat Completions cache-write and modality token usage recording for synchronous, asynchronous, and streaming calls.

Changes:

  • Extracts detailed usage fields into telemetry.
  • Adds local HTTP/SSE regression coverage.
  • Adds a changelog entry.
File summaries
File Description
instrumentation/opentelemetry-instrumentation-genai-openai/tests/test_chat_token_usage.py Tests usage-detail variants across call modes.
instrumentation/opentelemetry-instrumentation-genai-openai/src/opentelemetry/instrumentation/genai/openai/utils.py Extracts detailed usage fields.
instrumentation/opentelemetry-instrumentation-genai-openai/src/opentelemetry/instrumentation/genai/openai/patch.py Applies extraction to non-streaming responses.
instrumentation/opentelemetry-instrumentation-genai-openai/src/opentelemetry/instrumentation/genai/openai/chat_wrappers.py Applies extraction to streaming responses.
instrumentation/opentelemetry-instrumentation-genai-openai/.changelog/603.added Documents the feature.
Review details

Suppressed comments (1)

instrumentation/opentelemetry-instrumentation-genai-openai/tests/test_chat_token_usage.py:222

  • This only checks equality, so a bool or float usage value would still pass (True == 1 and 1.0 == 1). The new semconv usage attributes must be integers; assert the types of the exported values here so the regression test catches invalid attribute values.
    assert actual == expected
  • Files reviewed: 5/5 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@lmolkova lmolkova left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please make sure to follow existing test practices in this repo

)
invocation.text_input_tokens = get_property_value(
prompt_details, "text_tokens"
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OpenAI Chat Completions does not expose text_tokens or image_tokens in prompt_tokens_details or completion_tokens_details. Only cache_write_tokens and audio_tokens exist on PromptTokensDetails, and audio_tokens on CompletionTokensDetails. Please remove the non-existent modality mappings.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
@opentelemetry-pr-dashboard

Copy link
Copy Markdown

Hi @1fanwang — just a friendly reminder that this pull request is waiting on you. The dashboard status comment has the open items and is kept current.

  • Replying is enough to hand it off — answer, explain why no change is needed, or ask a follow-up. The dashboard routes it onward once nothing on the list is waiting on you.
  • To hand it back for any other reason, including the dashboard getting this wrong, comment /dashboard route:reviewers.

invocation.thinking_tokens = get_property_value(
obj=completion_details, property_name="reasoning_tokens"
)
invocation.set_input_tokens(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OpenAI Chat Completions does not expose text_tokens or image_tokens in prompt_tokens_details or completion_tokens_details. Only audio_tokens exists. Please remove text and image from the modality extractions and only pass audio.

Signed-off-by: 1fanwang <1fannnw@gmail.com>
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

This PR has been automatically marked as stale because it has not had any activity for 14 days. It will be closed if no further activity occurs within 14 days of this comment.
If you're still working on this, please add a comment or push new commits.

@github-actions github-actions Bot added the Stale Issue or PR has been inactive label Oct 6, 2026
Assisted-by: GitHub Copilot CLI (Auto)
Signed-off-by: 1fanwang <1fannnw@gmail.com>
@1fanwang 1fanwang changed the title [openai] Record cache-write and modality token usage [openai] Record cache-write and audio token usage Oct 6, 2026
@github-actions github-actions Bot removed the Stale Issue or PR has been inactive label Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

[openai] Capture detailed token usage (cache write and modality breakdown)

3 participants