Skip to content

fix(langchain): activate context at execution boundaries instead of in callbacks - #818

Open
sangkyoonnam wants to merge 28 commits into
open-telemetry:mainfrom
sangkyoonnam:fix/677-langchain-execution-context-split
Open

sangkyoonnam wants to merge 28 commits into
open-telemetry:mainfrom
sangkyoonnam:fix/677-langchain-execution-context-split

Conversation

@sangkyoonnam

Copy link
Copy Markdown
Contributor

Description

On top of #817 and the LangChain part of @lmolkova's prototype for #677, kept as the first commit with her authorship. Context activation moves from callbacks to the LangChain and LangGraph execution boundaries (chat models, tools, retrievers, runnables, composites, branches, single-input batches, streams, graph nodes), so attach and detach happen in the same Python context and nested SDK and HTTP spans correlate across asyncio tasks and threads. The commits after hers cover what the prototype missed: plain Runnable subclasses and composites, lazy stream config, missing-boundary handling, and main's chat span leak when ainvoke is cancelled or interrupted, fixed here because it sits on this boundary.

Known gaps: requires langchain >= 0.3.22 (core 0.3.46 added set_config_context) and the workspace util-genai 1.3b0.dev; graph-node propagation needs LangGraph >= 0.3.18. A plain invoke override under multi-input batch/abatch stays unparented, as on main; strict xfails. Private boundaries that are missing are skipped, with a warning for installed libraries. batch_as_completed has no dedicated test. A trivial RunnableLambda.invoke costs 18 to 19 us more than main on Python 3.12. The app probes ran on latest dependencies before the interrupt cleanup; examples/manual and zero-code are unrun.

Refs #677.

Type of change

  • Bug fix (non-breaking change which fixes an issue)

How has this been tested?

  • New tests in test_custom_execution_context.py, test_execution_context.py, test_stream_context.py, test_persistence_context.py cover span parenting, HTTP traceparent through instrumented httpx, cross-task stream reads, cancellation, interruption, lazy config and instrumentation restoration. Each fix has a test that fails on the commit before it.
  • tox -e py312-test-instrumentation-genai-langchain-latest: 620 passed, 2 xfailed; -oldest (langchain 0.3.22, langgraph 0.3.18): 565 passed, 21 skipped, 2 xfailed; util 520; openai 394; typecheck; precommit; langchain conformance 12 passed.
  • A local create_agent app with streaming ChatOpenAI, tools and a mock server, on main, fix(util-genai,openai): scope stream invocation context #817 and this branch: astream consumed in another task, batch, a thread pool, mid-stream cancel, a two-turn checkpointer and the util-genai: GenAIInvocation._finish detaches its context token unguarded; async LangChain runs log a traceback per invocation #677 shape. On main the async cases leave the provider chat span and its HTTP child as orphan roots; here every span has its parent and the caller's context is restored. All three states log zero detach errors.

Checklist

  • Followed the style guidelines of this project
  • Changelog updated if the change requires an entry
  • Unit tests added
  • Documentation updated

lmolkova and others added 21 commits September 30, 2026 22:04
…enAI part)

Keep callback-managed spans detached and activate their context around
framework execution boundaries, including custom tools and retrievers,
streaming, and graph persistence. Add inference and HTTP correlation tests
covering concurrency, cancellation, and context restoration.

Assisted-by: GPT-6
…art)

Keep callback-managed spans detached and activate their context around
framework execution boundaries, including custom tools and retrievers,
streaming, and graph persistence. Add inference and HTTP correlation tests
covering concurrency, cancellation, and context restoration.

Assisted-by: GPT-6
Copilot AI balanced review requested due to automatic review settings September 30, 2026 14:33
@sangkyoonnam
sangkyoonnam requested a review from a team as a code owner September 30, 2026 14:33
@opentelemetry-pr-dashboard

opentelemetry-pr-dashboard Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Pull request dashboard status

Waiting on the author · refreshed 2026-10-06 22:24 UTC

Respond to 1 review item (e.g. link a commit, explain why not, ask a follow-up):

  • Top-level threads: 1
Status above doesn't look right?
  • Just replied or pushed? Anything around or after the refresh time above may not be picked up yet — give it a few minutes.
  • Should this be with reviewers? Comment /dashboard route:reviewers to route it to them.
  • Anything wrong — including the routing? Report it with what you expected; it helps us improve the dashboard.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

LangChain telemetry finalization can still replace an underlying BaseException because its guard catches only Exception.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Moves LangChain/LangGraph span-context activation from callbacks to execution boundaries, while extending stream wrappers to scope context per read and cleanup.

Changes:

  • Adds execution-boundary instrumentation for LangChain/LangGraph calls, streams, batches, and graph nodes.
  • Adds reusable stream execution-context hooks and applies them to OpenAI streams.
  • Expands context propagation, cancellation, failure, and restoration tests.
File Description
uv.lock Updates LangChain dependency floor.
util/​opentelemetry-util-genai/​tests/​test_stream.py Tests scoped stream operations.
util/​opentelemetry-util-genai/​src/​opentelemetry/​util/​genai/​stream.py Adds execution-context stream hooks.
util/​opentelemetry-util-genai/​README.rst Documents the new stream hook.
util/​opentelemetry-util-genai/​AGENTS.md Updates stream guidance.
util/​opentelemetry-util-genai/​.changelog/​817.added Records the new hook.
README.md Updates the LangChain version floor.
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​test_responses_stream_context.py Tests Responses stream context.
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​test_response_wrappers.py Updates wrapper test fixtures.
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​test_chat_stream_context.py Tests chat stream context.
instrumentation/​opentelemetry-instrumentation-genai-openai/​tests/​requirements.oldest.txt Uses the workspace stream utility.
instrumentation/​opentelemetry-instrumentation-genai-openai/​src/​opentelemetry/​instrumentation/​genai/​openai/​response_wrappers.py Scopes Responses stream execution.
instrumentation/​opentelemetry-instrumentation-genai-openai/​src/​opentelemetry/​instrumentation/​genai/​openai/​chat_wrappers.py Scopes chat stream execution.
instrumentation/​opentelemetry-instrumentation-genai-openai/​pyproject.toml Raises the util dependency floor.
instrumentation/​opentelemetry-instrumentation-genai-openai/​.changelog/​817.fixed Records OpenAI context fixes.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_stream_context.py Tests lazy and cross-task streams.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_persistence_context.py Tests checkpoint/store context.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_execution_context.py Tests execution-boundary propagation.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_custom_execution_context.py Tests custom runnable boundaries.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_callback_handler.py Updates non-attaching callback expectations.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​test_agent_classification_corpus.py Updates agent invocation expectations.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​requirements.oldest.txt Adds oldest boundary-capable dependencies.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​tests/​requirements.latest.txt Adds integration-test dependencies.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​src/​opentelemetry/​instrumentation/​genai/​langchain/​package.py Raises the LangChain floor.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​src/​opentelemetry/​instrumentation/​genai/​langchain/​callback_handler.py Stops callbacks from attaching context.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​src/​opentelemetry/​instrumentation/​genai/​langchain/​_run_context.py Manages per-run execution scopes.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​src/​opentelemetry/​instrumentation/​genai/​langchain/​_execution_context.py Patches LangChain execution boundaries.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​src/​opentelemetry/​instrumentation/​genai/​langchain/​__init__.py Wires boundary instrumentation lifecycle.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​README.rst Documents support and limitations.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​pyproject.toml Raises dependency floors.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​.changelog/​818.fixed Records context propagation fixes.
instrumentation/​opentelemetry-instrumentation-genai-langchain/​.changelog/​818.changed Records dependency requirements.
AGENTS.md Updates repository stream guidance.
.github/​instructions/​instrumentation.instructions.md Updates instrumentation review guidance.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +90 to +91
except Exception:
_logger.exception("Failed to finalize LangChain run")
@lmolkova

lmolkova commented Oct 6, 2026

Copy link
Copy Markdown
Member

4 async composite tests fail on oldest (Python 3.10): test_plain_runnable_in_composite_correlates_http[async-*] and test_plain_runnable_fallback_correlates_http[async]. langchain-core < 1.5 on Python < 3.11 doesn't run async steps in the copied context. How about skipping them on that combination?

requires_async_step_context = pytest.mark.skipif(
    sys.version_info < (3, 11)
    and Version(version("langchain-core")) < Version("1.5.0"),
    reason="langchain-core < 1.5 on Python < 3.11 doesn't run async steps in the copied context",
)

# (a chat model invoked from a lambda body): attaching that one here would
# leave its ended span current for the rest of this frame.
config = dict(bound.arguments.get("config") or {})
scope.run_id = config.get("run_id") or uuid4()

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LangChain mints run ids with uuid7 (time-ordered; LangSmith says future versions will require v7), so injecting uuid4 here and in _wrap_run changes the ids every tracer sees for runnables, tools, retrievers and streams. Could we reuse LangChain's generator when it's available?

try:
    from langchain_core.utils.uuid import uuid7 as _new_run_id
except ImportError:  # older langchain-core
    from uuid import uuid4 as _new_run_id
Suggested change
scope.run_id = config.get("run_id") or uuid4()
scope.run_id = config.get("run_id") or _new_run_id()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8b3015b. _new_run_id is langchain_core.utils.uuid.uuid7 when the import works and uuid4 otherwise, and both _config_run and _wrap_run go through it. On langchain-core 1.6.6 the ids come out version 7; on the oldest pin (0.3.49) the module doesn't exist, so it falls back to uuid4, which is what those cores mint themselves.

Comment on lines +87 to +88
if isinstance(config, Sequence):
config = cast(Any, config[0])

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
if isinstance(config, Sequence):
config = cast(Any, config[0])
if isinstance(config, Sequence):
config = cast(Any, config[0]) if config else None

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 8b3015b. With config=[] and one input the old line raised IndexError from the wrapper before LangChain's own ValueError about the config count; now the wrapper returns no parent and LangChain's error surfaces.

@sangkyoonnam

Copy link
Copy Markdown
Contributor Author

Done in db80ddf, with your marker as written. Only the async params carry it (pytest.param("async", marks=requires_async_step_context)), so the sync cases still run there. py310 oldest (langchain-core 0.3.49): those 4 are now skipped (561 passed, 25 skipped in all, 21 of them pre-existing); without the marker the same 4 fail. py312 oldest: nothing new skipped, 565 passed. py312 latest (1.6.6): 620 passed.

@lmolkova lmolkova left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for working on it!

Let's get #817 in first and then rebase this one on it, so this PR only carries the LangChain changes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants