Repository navigation
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #5380 +/- ##
============================================
- Coverage 92.60% 91.72% -0.88%
- Complexity 5043 5103 +60
============================================
Files 1252 1287 +35
Lines 53587 55330 +1743
Branches 6673 7025 +352
============================================
+ Hits 49623 50754 +1131
- Misses 2315 2831 +516
- Partials 1649 1745 +96
☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
| config | throughput | MB/s | latency | max Δ latest / 7d | |
|---|---|---|---|---|---|
| 🔴 | bs=10 sw=10 sl=64 | 515 | 0.315 | 18,041/32,243/32,243 us | 🔴 +31.6% / 🔴 +115.9% |
| 🔴 | bs=100 sw=10 sl=64 | 1,210 | 0.738 | 82,034/94,733/94,733 us | 🔴 -5.1% / 🟢 -19.5% |
| ⚪ | bs=1000 sw=10 sl=64 | 1,411 | 0.861 | 711,418/769,232/769,232 us | ⚪ within ±5% / 🟢 +26.6% |
Baseline details
Latest main 8bfe034 from same runner
| config | metric | PR | latest main | 7d avg | Δ latest | Δ 7d |
|---|---|---|---|---|---|---|
| bs=10 sw=10 sl=64 | throughput | 515 tuples/sec | 611 tuples/sec | 850.28 tuples/sec | -15.7% | -39.4% |
| bs=10 sw=10 sl=64 | MB/s | 0.315 MB/s | 0.373 MB/s | 0.519 MB/s | -15.5% | -39.3% |
| bs=10 sw=10 sl=64 | p50 | 18,041 us | 15,730 us | 12,159 us | +14.7% | +48.4% |
| bs=10 sw=10 sl=64 | p95 | 32,243 us | 24,495 us | 14,933 us | +31.6% | +115.9% |
| bs=10 sw=10 sl=64 | p99 | 32,243 us | 24,495 us | 18,219 us | +31.6% | +77.0% |
| bs=100 sw=10 sl=64 | throughput | 1,210 tuples/sec | 1,275 tuples/sec | 1,085 tuples/sec | -5.1% | +11.5% |
| bs=100 sw=10 sl=64 | MB/s | 0.738 MB/s | 0.778 MB/s | 0.662 MB/s | -5.1% | +11.4% |
| bs=100 sw=10 sl=64 | p50 | 82,034 us | 78,188 us | 97,351 us | +4.9% | -15.7% |
| bs=100 sw=10 sl=64 | p95 | 94,733 us | 91,615 us | 103,954 us | +3.4% | -8.9% |
| bs=100 sw=10 sl=64 | p99 | 94,733 us | 91,615 us | 117,621 us | +3.4% | -19.5% |
| bs=1000 sw=10 sl=64 | throughput | 1,411 tuples/sec | 1,403 tuples/sec | 1,115 tuples/sec | +0.6% | +26.6% |
| bs=1000 sw=10 sl=64 | MB/s | 0.861 MB/s | 0.856 MB/s | 0.68 MB/s | +0.6% | +26.5% |
| bs=1000 sw=10 sl=64 | p50 | 711,418 us | 709,982 us | 955,077 us | +0.2% | -25.5% |
| bs=1000 sw=10 sl=64 | p95 | 769,232 us | 753,992 us | 1,000,375 us | +2.0% | -23.1% |
| bs=1000 sw=10 sl=64 | p99 | 769,232 us | 753,992 us | 1,025,400 us | +2.0% | -25.0% |
Raw CSV
config_idx,batch_size,schema_width,string_len,num_batches,total_ms,total_tuples,total_bytes,tuples_per_sec,mb_per_sec,lat_p50_us,lat_p95_us,lat_p99_us
0,10,10,64,20,388.07,200,128000,515,0.315,18041.05,32242.97,32242.97
1,100,10,64,20,1653.00,2000,1280000,1210,0.738,82033.97,94733.10,94733.10
2,1000,10,64,20,14175.76,20000,12800000,1411,0.861,711417.53,769232.40,769232.40
Automated Reviewer SuggestionsBased on the
|
Introduce the OpenTelemetry foundation for Texera: an SDK bootstrap, a Logback-to-OTel log bridge, and a log sanitizer, wired into every service entry point. This is PR 1 of the observability stack; it ships logging only, with the trace and metric exporters wired but not yet emitted (those arrive in follow-up PRs). New module (common/observability): - OtelInit: one-call SDK bootstrap per service. Reads OTEL_* settings from observability.conf, validates the OTLP endpoint (scheme + host allowlist) before any exporter is built, wires the span/log/metric providers explicitly (no sdk-extension-autoconfigure), clamps the metric export interval, and attaches the log appender to the Logback ROOT logger. The whole init body is guarded so a missing or malformed config returns None with one WARN instead of throwing into a service's run(), even when telemetry is disabled. - TexeraOtelLogAppender: bridges Logback events to OTel log records. Maps severity, forwards MDC, sets exception.type/message/stacktrace semantic attributes, and drops records from io.opentelemetry loggers so export-failure diagnostics are not fed back to the collector that just failed. - LogSanitizer: strips C0 control characters, redacts secrets, and caps body length (MaxBodyChars) to keep individual records bounded. Config (common/config): - observability.conf with the OTEL_* defaults, ObservabilityConfig to read it, and ENV_OTEL_* entries in EnvironmentalVariable. Wiring: - build.sbt defines the Observability module and adds dependsOn(Observability) to the eight Dropwizard services. - Each service entry point calls OtelInit.init with its own service name, and each service config gains a logging block. Deployment: - OTEL_* env entries in bin/single-node/.env, bin/k8s/values.yaml, and bin/k8s/values-development.yaml (kept a name-for-name mirror). - LICENSE-binary manifests updated with the pinned OTel 1.50.0 jars. Tests: OtelInitSpec, TexeraOtelLogAppenderSpec, LogSanitizerSpec, and ObservabilityConfigSpec cover endpoint validation, interval clamping, severity/MDC/exception mapping, self-diagnostic filtering, redaction, and truncation. Rebased onto current main.
Build on the logging foundation (PR 1) with the backend emit path: workflow lifecycle metrics and a run-level setup trace span, plus the tracing/metrics primitives they use. This is PR 2 of the observability stack. Primitives (common/observability): - TexeraTracer: lazy accessor for the process tracer off GlobalOpenTelemetry. - SpanAttrs: typed AttributeKey constants so span attributes use standard keys rather than ad hoc strings. - WorkflowMetrics: the OTel instruments (start/completion/failure/cancellation counters and run-duration histogram) keyed by workflow kind. - TraceparentValidator: validates W3C traceparent headers before use. Emit path (amber): - WorkflowMetricsRecorder: single owner of the metric instruments. init() wires them once; onStart stamps a run's start; onStateChange records terminal counters and duration exactly once on the first transition into a terminal state (idempotent, safe to call on every transition). - WorkflowService.initExecutionService runs inside a run-level setup span so setup-path logs carry its trace id. The span covers only the synchronous setup; synchronous setup failures are recorded on it in the catch block, and the span is ended in finally. The async errorHandler deliberately does not touch the span: it is invoked after setup returns and the span has ended, so the failure is surfaced through the metadata store instead. - ExecutionStateStore.updateWorkflowState is the single chokepoint that feeds every state transition to WorkflowMetricsRecorder.onStateChange. - ComputingUnitMaster initializes the recorder at startup. All observability sources live under common/observability (the module from PR 1); the tracing/metrics classes are not duplicated into common/config. Tests: WorkflowMetricsSpec, SpanAttrsSpec, and TraceparentValidatorSpec. Review follow-ups addressed: span attributes use SpanAttrs keys instead of plain strings; the error handler no longer records onto a span that may have already ended (documented and handled via the metadata store). Stacked on PR 1 (obs/pr1/foundations).
Add the local deployment wiring that receives and stores the telemetry emitted by PR 2: an OpenTelemetry Collector plus logs, metrics, traces, and profiles backends, for both the single-node docker-compose stack and the Kubernetes Helm chart. Infrastructure only; the application runs unchanged whether or not these services are started. This is PR 3 of the observability stack. Single-node compose (bin/single-node, bin/observability): - otel-collector as the single OTLP ingress, routing to VictoriaLogs (logs), VictoriaMetrics (metrics), and Jaeger (traces); the collector config caps OTLP message size and bounds memory via memory_limiter. - Parca server + eBPF agent for continuous profiling, configured in bin/observability/parca. - Each signal is its own compose profile and is OFF by default: COMPOSE_PROFILES is empty, so observability is opt-in. docker compose reads COMPOSE_PROFILES natively via the canonical launcher. - Emission is gated by OTEL_SDK_DISABLED (default disabled); nothing is emitted unless an operator opts in alongside the collector profile. Kubernetes Helm chart (bin/k8s): - Templates for the collector, Jaeger, VictoriaLogs, and VictoriaMetrics under templates/base/observability, plus values.yaml knobs and an install.sh helper. - Service deployments gain the OTLP endpoint wiring. Every observability port binds to loopback (127.0.0.1) or stays on the bridge network; no 0.0.0.0 host bindings. Tests: ObservabilityComposeSpec and ParcaConfigSpec (string-level smoke tests over the compose, collector, and Parca config, pinning image versions, the opt-in default, loopback binding, and the no-high- cardinality-label rule). All 24 pass; sbt scalafmtCheckAll is clean. Review follow-ups addressed: observability is opt-in rather than on by default, and the privileged eBPF agent is opt-in with its host-wide scope documented; removed the non-canonical up.sh launcher so COMPOSE_PROFILES is the single switch; corrected the Parca README (the agent profiles every host process, and profiles cannot be joined to traces by trace_id); moved the local-dev-only OTLP host ports into docker-compose.override.yml so the base compose stays deploy-clean; the release workflow now ships bin/observability and rewrites the mount paths so the compose works from the released single-node bundle; and removed the inter-PR references from the configs, comments, and specs. Stacked on PR 2 (obs/pr2/backend-emit).
468fc47 to
b14eab1
Compare
Add the tenant-scoped read path the dashboard queries, plus the Angular shell that hosts the per-signal panels. This is the first user-visible observability surface: it enforces tenancy and rate limiting on every query and gates each panel on signal reachability. This is PR 4 of the observability stack. Query gateway (amber, org.apache.texera.web.observability.gateway): - BackendClient issues the read-only HTTP queries to the telemetry backends (VictoriaLogs, VictoriaMetrics, Jaeger). - ScopeResolver derives the caller's tenant scope and constrains every query to it, so a user never reads another tenant's telemetry. - RateLimiter, AuditLogger, and GatewayContext add per-request rate limiting, audit logging, and the shared request context. - dtos.scala and builders.scala define the typed request objects and their validators (time window, page size, free text, service name). - ObservabilityResources exposes /observability/health and is registered in TexeraWebApplication. - RequestContextMdcFilter and UserContextMdcFilter inject the request and user context into the logging MDC. - ObservabilityGatewayConfig plus observability-gateway.conf hold the gateway settings. Dashboard shell (frontend): - Admin-only Observability page, route, and navigation entry, plus the observability service, its types, and the traces-pivot service. - Each tab is guarded by the per-signal reachability check: an unreachable signal renders an explicit state rather than a broken panel. The signal panels themselves follow in later PRs of the stack. Tests: backend specs for the gateway core, DTO validation, scope resolver, rate limiter, and the MDC filters, plus frontend component and service specs. sbt scalafmtCheckAll is clean. Stacked on PR 3 (obs/pr3/deployment). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the first complete signal: log search through the query gateway plus the UI panel that drives it. Builds on the gateway core and dashboard shell from PR 4. This is PR 5 of the observability stack. Gateway (amber): - ResponseParsers turns the log store's responses into the typed results the UI consumes (parseLogs, parseLogSources), applying per-field redaction so sensitive values never reach the client. - builders and dtos gain the logs query builder and its validated request objects (time window, service, workflow, computing unit, execution, level, free text, page size). - ObservabilityResources exposes the log-search and source-facets endpoints, registered in TexeraWebApplication. Logs panel (frontend): - A logs panel under the observability page with time-window, service, workflow, computing-unit, execution, level, and free-text filters, server-side paging, and sources-backed autofill. - observability-prefs persists the panel's filter choices, and the observability component hosts the panel behind its reachability gate. Tests: backend specs for the logs query builder and the response parsers, plus the frontend logs-panel and service specs. sbt scalafmtCheckAll is clean. Stacked on PR 4 (obs/pr4/gateway-core). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the metrics signal: workflow metrics through the query gateway plus an ECharts-based UI panel that renders them. Builds on the logs signal from PR 5. This is PR 6 of the observability stack. Gateway (amber): - builders and parseMetrics add the metrics query builder and turn the metrics store's responses into typed series. - ObservabilityResources serves a closed allowlist of named queries (throughput, outcome rates, and duration percentiles), enforced on the gateway so an arbitrary query can never be run; registered in TexeraWebApplication. - WorkflowRunCounter and the GatewayContext and dtos changes supply the run counts and validated request objects the named queries need. Metrics panel (frontend): - A metrics panel under the observability page that renders each named query as a typed series with Apache ECharts. Chart data is bound as typed arrays; backend output is never interpreted as a formatter or template. - The ECharts dependency is added to package.json, yarn.lock, and the frontend LICENSE-binary. Tests: backend specs for the metrics query builder and the parsers, plus the frontend metrics-panel and service specs. sbt scalafmtCheckAll is clean. Stacked on PR 5 (obs/pr5/logs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
b14eab1 to
3d0da09
Compare
What changes were proposed in this PR?
Workflow metrics through the gateway plus an ECharts-based UI panel.
Backend:
parseMetrics.MetricsResourceserving a closed allowlist of named queries (throughput, outcome rates, and duration percentiles), enforced on both the gateway and the client.Frontend:
Any related issues, documentation, or discussions?
Closes: #5372
Part of #4070. Stacked on #5379.
How was this PR tested?
sbt scalafmtCheckAllpasses.prettier-eslintandeslintpass.Was this PR authored or co-authored using generative AI tooling?
Co-authored with Claude Opus 4.8 in compliance with ASF