Skip to content

feat(observability): metrics end-to-end (gateway + UI panel) - #5380

Draft
Ma77Ball wants to merge 6 commits into
apache:mainfrom
Ma77Ball:obs/pr6/metrics
Draft

Ma77Ball wants to merge 6 commits into
apache:mainfrom
Ma77Ball:obs/pr6/metrics

Conversation

@Ma77Ball

@Ma77Ball Ma77Ball commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this PR?

Workflow metrics through the gateway plus an ECharts-based UI panel.
Backend:

  • Metrics query builder and parseMetrics.
  • MetricsResource serving a closed allowlist of named queries (throughput, outcome rates, and duration percentiles), enforced on both the gateway and the client.
    Frontend:
  • Metrics panel rendering each named query as a typed series with Apache ECharts. Chart data is bound as typed arrays; backend output is never interpreted as a formatter or template.

Any related issues, documentation, or discussions?

Closes: #5372
Part of #4070. Stacked on #5379.

How was this PR tested?

  • Backend specs for the metrics query builder and parser; sbt scalafmtCheckAll passes.
  • Frontend metrics-panel and service specs; prettier-eslint and eslint pass.
  • Compile and the full test suites run in this PR's CI.

Was this PR authored or co-authored using generative AI tooling?

Co-authored with Claude Opus 4.8 in compliance with ASF

@github-actions github-actions Bot added engine dependencies Pull requests that update a dependency file frontend Changes related to the frontend GUI docs Changes related to documentations infra common labels Jun 5, 2026
@codecov-commenter

codecov-commenter commented Jun 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 64.88332% with 617 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.72%. Comparing base (b143a4f) to head (3d0da09).
⚠️ Report is 8 commits behind head on main.
✅ All tests successful. No failed tests found.

Files with missing lines Patch % Lines
...observability/gateway/ObservabilityResources.scala 0.00% 241 Missing ⚠️
...observability/logs-panel/logs-panel.component.html 56.71% 58 Missing ⚠️
...apache/texera/web/observability/gateway/dtos.scala 73.40% 44 Missing and 6 partials ⚠️
...texera/web/observability/gateway/AuditLogger.scala 0.00% 41 Missing ⚠️
...ala/org/apache/texera/observability/OtelInit.scala 70.86% 25 Missing and 12 partials ⚠️
...xera/web/observability/gateway/ScopeResolver.scala 22.58% 22 Missing and 2 partials ⚠️
...ra/web/observability/gateway/ResponseParsers.scala 77.27% 9 Missing and 11 partials ⚠️
...web/observability/gateway/WorkflowRunCounter.scala 5.00% 19 Missing ⚠️
...r/observability/logs-panel/logs-panel.component.ts 85.18% 3 Missing and 13 partials ⚠️
...rg/apache/texera/web/service/WorkflowService.scala 74.13% 11 Missing and 4 partials ⚠️
... and 19 more
Additional details and impacted files
@@             Coverage Diff              @@
##               main    #5380      +/-   ##
============================================
- Coverage     92.60%   91.72%   -0.88%     
- Complexity     5043     5103      +60     
============================================
  Files          1252     1287      +35     
  Lines         53587    55330    +1743     
  Branches       6673     7025     +352     
============================================
+ Hits          49623    50754    +1131     
- Misses         2315     2831     +516     
- Partials       1649     1745      +96     
Flag Coverage Δ
access-control-service 77.44% <100.00%> (+0.06%) ⬆️
agent-service 99.16% <ø> (ø)
amber 86.30% <58.73%> (-1.80%) ⬇️
computing-unit-managing-service 60.52% <100.00%> (+0.03%) ⬆️
config-service 87.50% <100.00%> (+0.12%) ⬆️
file-service 81.45% <0.00%> (-0.09%) ⬇️
frontend 96.24% <80.12%> (-0.35%) ⬇️
notebook-migration-service 83.77% <100.00%> (+0.03%) ⬆️
pyamber 98.58% <ø> (ø)
workflow-compiling-service 74.25% <100.00%> (+0.15%) ⬆️

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions github-actions Bot added the platform Non-amber Scala service paths label Jun 5, 2026
@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Benchmark changes need a look

🟢 0 better · 🔴 7 worse · ⚪ 8 noise (<±5%) · 0 without baseline

Compared against main 8bfe034 benchmarked on this same runner, so the delta is largely free of cross-runner hardware noise. The "7d avg" column still reflects the gh-pages dashboard. Treat <±5% as noise unless repeated.

Dashboard · Run

config throughput MB/s latency max Δ latest / 7d
🔴 bs=10 sw=10 sl=64 515 0.315 18,041/32,243/32,243 us 🔴 +31.6% / 🔴 +115.9%
🔴 bs=100 sw=10 sl=64 1,210 0.738 82,034/94,733/94,733 us 🔴 -5.1% / 🟢 -19.5%
⚪ bs=1000 sw=10 sl=64 1,411 0.861 711,418/769,232/769,232 us ⚪ within ±5% / 🟢 +26.6%
Baseline details

Latest main 8bfe034 from same runner

config metric PR latest main 7d avg Δ latest Δ 7d
bs=10 sw=10 sl=64 throughput 515 tuples/sec 611 tuples/sec 850.28 tuples/sec -15.7% -39.4%
bs=10 sw=10 sl=64 MB/s 0.315 MB/s 0.373 MB/s 0.519 MB/s -15.5% -39.3%
bs=10 sw=10 sl=64 p50 18,041 us 15,730 us 12,159 us +14.7% +48.4%
bs=10 sw=10 sl=64 p95 32,243 us 24,495 us 14,933 us +31.6% +115.9%
bs=10 sw=10 sl=64 p99 32,243 us 24,495 us 18,219 us +31.6% +77.0%
bs=100 sw=10 sl=64 throughput 1,210 tuples/sec 1,275 tuples/sec 1,085 tuples/sec -5.1% +11.5%
bs=100 sw=10 sl=64 MB/s 0.738 MB/s 0.778 MB/s 0.662 MB/s -5.1% +11.4%
bs=100 sw=10 sl=64 p50 82,034 us 78,188 us 97,351 us +4.9% -15.7%
bs=100 sw=10 sl=64 p95 94,733 us 91,615 us 103,954 us +3.4% -8.9%
bs=100 sw=10 sl=64 p99 94,733 us 91,615 us 117,621 us +3.4% -19.5%
bs=1000 sw=10 sl=64 throughput 1,411 tuples/sec 1,403 tuples/sec 1,115 tuples/sec +0.6% +26.6%
bs=1000 sw=10 sl=64 MB/s 0.861 MB/s 0.856 MB/s 0.68 MB/s +0.6% +26.5%
bs=1000 sw=10 sl=64 p50 711,418 us 709,982 us 955,077 us +0.2% -25.5%
bs=1000 sw=10 sl=64 p95 769,232 us 753,992 us 1,000,375 us +2.0% -23.1%
bs=1000 sw=10 sl=64 p99 769,232 us 753,992 us 1,025,400 us +2.0% -25.0%
Raw CSV
config_idx,batch_size,schema_width,string_len,num_batches,total_ms,total_tuples,total_bytes,tuples_per_sec,mb_per_sec,lat_p50_us,lat_p95_us,lat_p99_us
0,10,10,64,20,388.07,200,128000,515,0.315,18041.05,32242.97,32242.97
1,100,10,64,20,1653.00,2000,1280000,1210,0.738,82033.97,94733.10,94733.10
2,1000,10,64,20,14175.76,20000,12800000,1411,0.861,711417.53,769232.40,769232.40

@github-actions

github-actions Bot commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Automated Reviewer Suggestions

Based on the git blame history of the changed files, we recommend the following reviewers:

  • Committers with relevant context: @parshimers
    You can request their reviews formally with /request-review @parshimers.

  • Contributors with relevant context: @bobbai00, @dzueck, @aicam
    You can notify them by mentioning @bobbai00, @dzueck, @aicam in a comment.

Introduce the OpenTelemetry foundation for Texera: an SDK bootstrap, a
Logback-to-OTel log bridge, and a log sanitizer, wired into every service
entry point. This is PR 1 of the observability stack; it ships logging only,
with the trace and metric exporters wired but not yet emitted (those arrive in
follow-up PRs).

New module (common/observability):
- OtelInit: one-call SDK bootstrap per service. Reads OTEL_* settings from
  observability.conf, validates the OTLP endpoint (scheme + host allowlist)
  before any exporter is built, wires the span/log/metric providers explicitly
  (no sdk-extension-autoconfigure), clamps the metric export interval, and
  attaches the log appender to the Logback ROOT logger. The whole init body is
  guarded so a missing or malformed config returns None with one WARN instead
  of throwing into a service's run(), even when telemetry is disabled.
- TexeraOtelLogAppender: bridges Logback events to OTel log records. Maps
  severity, forwards MDC, sets exception.type/message/stacktrace semantic
  attributes, and drops records from io.opentelemetry loggers so export-failure
  diagnostics are not fed back to the collector that just failed.
- LogSanitizer: strips C0 control characters, redacts secrets, and caps body
  length (MaxBodyChars) to keep individual records bounded.

Config (common/config):
- observability.conf with the OTEL_* defaults, ObservabilityConfig to read it,
  and ENV_OTEL_* entries in EnvironmentalVariable.

Wiring:
- build.sbt defines the Observability module and adds dependsOn(Observability)
  to the eight Dropwizard services.
- Each service entry point calls OtelInit.init with its own service name, and
  each service config gains a logging block.

Deployment:
- OTEL_* env entries in bin/single-node/.env, bin/k8s/values.yaml, and
  bin/k8s/values-development.yaml (kept a name-for-name mirror).
- LICENSE-binary manifests updated with the pinned OTel 1.50.0 jars.

Tests: OtelInitSpec, TexeraOtelLogAppenderSpec, LogSanitizerSpec, and
ObservabilityConfigSpec cover endpoint validation, interval clamping,
severity/MDC/exception mapping, self-diagnostic filtering, redaction, and
truncation.

Rebased onto current main.
Build on the logging foundation (PR 1) with the backend emit path: workflow
lifecycle metrics and a run-level setup trace span, plus the tracing/metrics
primitives they use. This is PR 2 of the observability stack.

Primitives (common/observability):
- TexeraTracer: lazy accessor for the process tracer off GlobalOpenTelemetry.
- SpanAttrs: typed AttributeKey constants so span attributes use standard keys
  rather than ad hoc strings.
- WorkflowMetrics: the OTel instruments (start/completion/failure/cancellation
  counters and run-duration histogram) keyed by workflow kind.
- TraceparentValidator: validates W3C traceparent headers before use.

Emit path (amber):
- WorkflowMetricsRecorder: single owner of the metric instruments. init() wires
  them once; onStart stamps a run's start; onStateChange records terminal
  counters and duration exactly once on the first transition into a terminal
  state (idempotent, safe to call on every transition).
- WorkflowService.initExecutionService runs inside a run-level setup span so
  setup-path logs carry its trace id. The span covers only the synchronous
  setup; synchronous setup failures are recorded on it in the catch block, and
  the span is ended in finally. The async errorHandler deliberately does not
  touch the span: it is invoked after setup returns and the span has ended, so
  the failure is surfaced through the metadata store instead.
- ExecutionStateStore.updateWorkflowState is the single chokepoint that feeds
  every state transition to WorkflowMetricsRecorder.onStateChange.
- ComputingUnitMaster initializes the recorder at startup.

All observability sources live under common/observability (the module from
PR 1); the tracing/metrics classes are not duplicated into common/config.

Tests: WorkflowMetricsSpec, SpanAttrsSpec, and TraceparentValidatorSpec.

Review follow-ups addressed: span attributes use SpanAttrs keys instead of
plain strings; the error handler no longer records onto a span that may have
already ended (documented and handled via the metadata store).

Stacked on PR 1 (obs/pr1/foundations).
Add the local deployment wiring that receives and stores the telemetry
emitted by PR 2: an OpenTelemetry Collector plus logs, metrics, traces,
and profiles backends, for both the single-node docker-compose stack and
the Kubernetes Helm chart. Infrastructure only; the application runs
unchanged whether or not these services are started. This is PR 3 of the
observability stack.

Single-node compose (bin/single-node, bin/observability):
- otel-collector as the single OTLP ingress, routing to VictoriaLogs
  (logs), VictoriaMetrics (metrics), and Jaeger (traces); the collector
  config caps OTLP message size and bounds memory via memory_limiter.
- Parca server + eBPF agent for continuous profiling, configured in
  bin/observability/parca.
- Each signal is its own compose profile and is OFF by default:
  COMPOSE_PROFILES is empty, so observability is opt-in. docker compose
  reads COMPOSE_PROFILES natively via the canonical launcher.
- Emission is gated by OTEL_SDK_DISABLED (default disabled); nothing is
  emitted unless an operator opts in alongside the collector profile.

Kubernetes Helm chart (bin/k8s):
- Templates for the collector, Jaeger, VictoriaLogs, and VictoriaMetrics
  under templates/base/observability, plus values.yaml knobs and an
  install.sh helper.
- Service deployments gain the OTLP endpoint wiring.

Every observability port binds to loopback (127.0.0.1) or stays on the
bridge network; no 0.0.0.0 host bindings.

Tests: ObservabilityComposeSpec and ParcaConfigSpec (string-level smoke
tests over the compose, collector, and Parca config, pinning image
versions, the opt-in default, loopback binding, and the no-high-
cardinality-label rule). All 24 pass; sbt scalafmtCheckAll is clean.

Review follow-ups addressed: observability is opt-in rather than on by
default, and the privileged eBPF agent is opt-in with its host-wide scope
documented; removed the non-canonical up.sh launcher so COMPOSE_PROFILES
is the single switch; corrected the Parca README (the agent profiles every
host process, and profiles cannot be joined to traces by trace_id); moved
the local-dev-only OTLP host ports into docker-compose.override.yml so the
base compose stays deploy-clean; the release workflow now ships
bin/observability and rewrites the mount paths so the compose works from
the released single-node bundle; and removed the inter-PR references from
the configs, comments, and specs.

Stacked on PR 2 (obs/pr2/backend-emit).
Ma77Ball and others added 3 commits October 9, 2026 23:58
Add the tenant-scoped read path the dashboard queries, plus the Angular
shell that hosts the per-signal panels. This is the first user-visible
observability surface: it enforces tenancy and rate limiting on every
query and gates each panel on signal reachability. This is PR 4 of the
observability stack.

Query gateway (amber, org.apache.texera.web.observability.gateway):
- BackendClient issues the read-only HTTP queries to the telemetry
  backends (VictoriaLogs, VictoriaMetrics, Jaeger).
- ScopeResolver derives the caller's tenant scope and constrains every
  query to it, so a user never reads another tenant's telemetry.
- RateLimiter, AuditLogger, and GatewayContext add per-request rate
  limiting, audit logging, and the shared request context.
- dtos.scala and builders.scala define the typed request objects and
  their validators (time window, page size, free text, service name).
- ObservabilityResources exposes /observability/health and is
  registered in TexeraWebApplication.
- RequestContextMdcFilter and UserContextMdcFilter inject the request
  and user context into the logging MDC.
- ObservabilityGatewayConfig plus observability-gateway.conf hold the
  gateway settings.

Dashboard shell (frontend):
- Admin-only Observability page, route, and navigation entry, plus the
  observability service, its types, and the traces-pivot service.
- Each tab is guarded by the per-signal reachability check: an
  unreachable signal renders an explicit state rather than a broken
  panel. The signal panels themselves follow in later PRs of the stack.

Tests: backend specs for the gateway core, DTO validation, scope
resolver, rate limiter, and the MDC filters, plus frontend component and
service specs. sbt scalafmtCheckAll is clean.

Stacked on PR 3 (obs/pr3/deployment).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the first complete signal: log search through the query gateway plus
the UI panel that drives it. Builds on the gateway core and dashboard
shell from PR 4. This is PR 5 of the observability stack.

Gateway (amber):
- ResponseParsers turns the log store's responses into the typed results
  the UI consumes (parseLogs, parseLogSources), applying per-field
  redaction so sensitive values never reach the client.
- builders and dtos gain the logs query builder and its validated
  request objects (time window, service, workflow, computing unit,
  execution, level, free text, page size).
- ObservabilityResources exposes the log-search and source-facets
  endpoints, registered in TexeraWebApplication.

Logs panel (frontend):
- A logs panel under the observability page with time-window, service,
  workflow, computing-unit, execution, level, and free-text filters,
  server-side paging, and sources-backed autofill.
- observability-prefs persists the panel's filter choices, and the
  observability component hosts the panel behind its reachability gate.

Tests: backend specs for the logs query builder and the response
parsers, plus the frontend logs-panel and service specs. sbt
scalafmtCheckAll is clean.

Stacked on PR 4 (obs/pr4/gateway-core).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the metrics signal: workflow metrics through the query gateway plus
an ECharts-based UI panel that renders them. Builds on the logs signal
from PR 5. This is PR 6 of the observability stack.

Gateway (amber):
- builders and parseMetrics add the metrics query builder and turn the
  metrics store's responses into typed series.
- ObservabilityResources serves a closed allowlist of named queries
  (throughput, outcome rates, and duration percentiles), enforced on the
  gateway so an arbitrary query can never be run; registered in
  TexeraWebApplication.
- WorkflowRunCounter and the GatewayContext and dtos changes supply the
  run counts and validated request objects the named queries need.

Metrics panel (frontend):
- A metrics panel under the observability page that renders each named
  query as a typed series with Apache ECharts. Chart data is bound as
  typed arrays; backend output is never interpreted as a formatter or
  template.
- The ECharts dependency is added to package.json, yarn.lock, and the
  frontend LICENSE-binary.

Tests: backend specs for the metrics query builder and the parsers, plus
the frontend metrics-panel and service specs. sbt scalafmtCheckAll is
clean.

Stacked on PR 5 (obs/pr5/logs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci changes related to CI common dependencies Pull requests that update a dependency file docs Changes related to documentations engine frontend Changes related to the frontend GUI infra platform Non-amber Scala service paths

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Observability] Workflow metrics: gateway endpoint and dashboard panel

2 participants