Repository navigation
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #5379 +/- ##
============================================
- Coverage 92.60% 91.89% -0.71%
- Complexity 5043 5105 +62
============================================
Files 1252 1284 +32
Lines 53587 55035 +1448
Branches 6673 6964 +291
============================================
+ Hits 49623 50575 +952
- Misses 2315 2725 +410
- Partials 1649 1735 +86
☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
| config | throughput | MB/s | latency | max Δ latest / 7d | |
|---|---|---|---|---|---|
| 🔴 | bs=10 sw=10 sl=64 | 530 | 0.323 | 18,376/22,752/22,752 us | 🟢 -12.5% / 🔴 +52.4% |
| 🔴 | bs=100 sw=10 sl=64 | 1,206 | 0.736 | 79,506/119,980/119,980 us | 🔴 +20.2% / 🟢 -18.3% |
| 🔴 | bs=1000 sw=10 sl=64 | 1,415 | 0.863 | 708,939/822,508/822,508 us | 🔴 +13.2% / 🟢 +26.9% |
Baseline details
Latest main 8bfe034 from same runner
| config | metric | PR | latest main | 7d avg | Δ latest | Δ 7d |
|---|---|---|---|---|---|---|
| bs=10 sw=10 sl=64 | throughput | 530 tuples/sec | 562 tuples/sec | 850.28 tuples/sec | -5.7% | -37.7% |
| bs=10 sw=10 sl=64 | MB/s | 0.323 MB/s | 0.343 MB/s | 0.519 MB/s | -5.8% | -37.8% |
| bs=10 sw=10 sl=64 | p50 | 18,376 us | 18,206 us | 12,159 us | +0.9% | +51.1% |
| bs=10 sw=10 sl=64 | p95 | 22,752 us | 25,988 us | 14,933 us | -12.5% | +52.4% |
| bs=10 sw=10 sl=64 | p99 | 22,752 us | 25,988 us | 18,219 us | -12.5% | +24.9% |
| bs=100 sw=10 sl=64 | throughput | 1,206 tuples/sec | 1,237 tuples/sec | 1,085 tuples/sec | -2.5% | +11.1% |
| bs=100 sw=10 sl=64 | MB/s | 0.736 MB/s | 0.755 MB/s | 0.662 MB/s | -2.5% | +11.1% |
| bs=100 sw=10 sl=64 | p50 | 79,506 us | 79,838 us | 97,351 us | -0.4% | -18.3% |
| bs=100 sw=10 sl=64 | p95 | 119,980 us | 99,841 us | 103,954 us | +20.2% | +15.4% |
| bs=100 sw=10 sl=64 | p99 | 119,980 us | 99,841 us | 117,621 us | +20.2% | +2.0% |
| bs=1000 sw=10 sl=64 | throughput | 1,415 tuples/sec | 1,440 tuples/sec | 1,115 tuples/sec | -1.7% | +26.9% |
| bs=1000 sw=10 sl=64 | MB/s | 0.863 MB/s | 0.879 MB/s | 0.68 MB/s | -1.8% | +26.8% |
| bs=1000 sw=10 sl=64 | p50 | 708,939 us | 699,503 us | 955,077 us | +1.3% | -25.8% |
| bs=1000 sw=10 sl=64 | p95 | 822,508 us | 726,288 us | 1,000,375 us | +13.2% | -17.8% |
| bs=1000 sw=10 sl=64 | p99 | 822,508 us | 726,288 us | 1,025,400 us | +13.2% | -19.8% |
Raw CSV
config_idx,batch_size,schema_width,string_len,num_batches,total_ms,total_tuples,total_bytes,tuples_per_sec,mb_per_sec,lat_p50_us,lat_p95_us,lat_p99_us
0,10,10,64,20,377.52,200,128000,530,0.323,18376.31,22752.04,22752.04
1,100,10,64,20,1658.29,2000,1280000,1206,0.736,79506.49,119980.17,119980.17
2,1000,10,64,20,14137.35,20000,12800000,1415,0.863,708939.11,822507.90,822507.90
Automated Reviewer SuggestionsBased on the
|
Introduce the OpenTelemetry foundation for Texera: an SDK bootstrap, a Logback-to-OTel log bridge, and a log sanitizer, wired into every service entry point. This is PR 1 of the observability stack; it ships logging only, with the trace and metric exporters wired but not yet emitted (those arrive in follow-up PRs). New module (common/observability): - OtelInit: one-call SDK bootstrap per service. Reads OTEL_* settings from observability.conf, validates the OTLP endpoint (scheme + host allowlist) before any exporter is built, wires the span/log/metric providers explicitly (no sdk-extension-autoconfigure), clamps the metric export interval, and attaches the log appender to the Logback ROOT logger. The whole init body is guarded so a missing or malformed config returns None with one WARN instead of throwing into a service's run(), even when telemetry is disabled. - TexeraOtelLogAppender: bridges Logback events to OTel log records. Maps severity, forwards MDC, sets exception.type/message/stacktrace semantic attributes, and drops records from io.opentelemetry loggers so export-failure diagnostics are not fed back to the collector that just failed. - LogSanitizer: strips C0 control characters, redacts secrets, and caps body length (MaxBodyChars) to keep individual records bounded. Config (common/config): - observability.conf with the OTEL_* defaults, ObservabilityConfig to read it, and ENV_OTEL_* entries in EnvironmentalVariable. Wiring: - build.sbt defines the Observability module and adds dependsOn(Observability) to the eight Dropwizard services. - Each service entry point calls OtelInit.init with its own service name, and each service config gains a logging block. Deployment: - OTEL_* env entries in bin/single-node/.env, bin/k8s/values.yaml, and bin/k8s/values-development.yaml (kept a name-for-name mirror). - LICENSE-binary manifests updated with the pinned OTel 1.50.0 jars. Tests: OtelInitSpec, TexeraOtelLogAppenderSpec, LogSanitizerSpec, and ObservabilityConfigSpec cover endpoint validation, interval clamping, severity/MDC/exception mapping, self-diagnostic filtering, redaction, and truncation. Rebased onto current main.
Build on the logging foundation (PR 1) with the backend emit path: workflow lifecycle metrics and a run-level setup trace span, plus the tracing/metrics primitives they use. This is PR 2 of the observability stack. Primitives (common/observability): - TexeraTracer: lazy accessor for the process tracer off GlobalOpenTelemetry. - SpanAttrs: typed AttributeKey constants so span attributes use standard keys rather than ad hoc strings. - WorkflowMetrics: the OTel instruments (start/completion/failure/cancellation counters and run-duration histogram) keyed by workflow kind. - TraceparentValidator: validates W3C traceparent headers before use. Emit path (amber): - WorkflowMetricsRecorder: single owner of the metric instruments. init() wires them once; onStart stamps a run's start; onStateChange records terminal counters and duration exactly once on the first transition into a terminal state (idempotent, safe to call on every transition). - WorkflowService.initExecutionService runs inside a run-level setup span so setup-path logs carry its trace id. The span covers only the synchronous setup; synchronous setup failures are recorded on it in the catch block, and the span is ended in finally. The async errorHandler deliberately does not touch the span: it is invoked after setup returns and the span has ended, so the failure is surfaced through the metadata store instead. - ExecutionStateStore.updateWorkflowState is the single chokepoint that feeds every state transition to WorkflowMetricsRecorder.onStateChange. - ComputingUnitMaster initializes the recorder at startup. All observability sources live under common/observability (the module from PR 1); the tracing/metrics classes are not duplicated into common/config. Tests: WorkflowMetricsSpec, SpanAttrsSpec, and TraceparentValidatorSpec. Review follow-ups addressed: span attributes use SpanAttrs keys instead of plain strings; the error handler no longer records onto a span that may have already ended (documented and handled via the metadata store). Stacked on PR 1 (obs/pr1/foundations).
Add the local deployment wiring that receives and stores the telemetry emitted by PR 2: an OpenTelemetry Collector plus logs, metrics, traces, and profiles backends, for both the single-node docker-compose stack and the Kubernetes Helm chart. Infrastructure only; the application runs unchanged whether or not these services are started. This is PR 3 of the observability stack. Single-node compose (bin/single-node, bin/observability): - otel-collector as the single OTLP ingress, routing to VictoriaLogs (logs), VictoriaMetrics (metrics), and Jaeger (traces); the collector config caps OTLP message size and bounds memory via memory_limiter. - Parca server + eBPF agent for continuous profiling, configured in bin/observability/parca. - Each signal is its own compose profile and is OFF by default: COMPOSE_PROFILES is empty, so observability is opt-in. docker compose reads COMPOSE_PROFILES natively via the canonical launcher. - Emission is gated by OTEL_SDK_DISABLED (default disabled); nothing is emitted unless an operator opts in alongside the collector profile. Kubernetes Helm chart (bin/k8s): - Templates for the collector, Jaeger, VictoriaLogs, and VictoriaMetrics under templates/base/observability, plus values.yaml knobs and an install.sh helper. - Service deployments gain the OTLP endpoint wiring. Every observability port binds to loopback (127.0.0.1) or stays on the bridge network; no 0.0.0.0 host bindings. Tests: ObservabilityComposeSpec and ParcaConfigSpec (string-level smoke tests over the compose, collector, and Parca config, pinning image versions, the opt-in default, loopback binding, and the no-high- cardinality-label rule). All 24 pass; sbt scalafmtCheckAll is clean. Review follow-ups addressed: observability is opt-in rather than on by default, and the privileged eBPF agent is opt-in with its host-wide scope documented; removed the non-canonical up.sh launcher so COMPOSE_PROFILES is the single switch; corrected the Parca README (the agent profiles every host process, and profiles cannot be joined to traces by trace_id); moved the local-dev-only OTLP host ports into docker-compose.override.yml so the base compose stays deploy-clean; the release workflow now ships bin/observability and rewrites the mount paths so the compose works from the released single-node bundle; and removed the inter-PR references from the configs, comments, and specs. Stacked on PR 2 (obs/pr2/backend-emit).
Add the tenant-scoped read path the dashboard queries, plus the Angular shell that hosts the per-signal panels. This is the first user-visible observability surface: it enforces tenancy and rate limiting on every query and gates each panel on signal reachability. This is PR 4 of the observability stack. Query gateway (amber, org.apache.texera.web.observability.gateway): - BackendClient issues the read-only HTTP queries to the telemetry backends (VictoriaLogs, VictoriaMetrics, Jaeger). - ScopeResolver derives the caller's tenant scope and constrains every query to it, so a user never reads another tenant's telemetry. - RateLimiter, AuditLogger, and GatewayContext add per-request rate limiting, audit logging, and the shared request context. - dtos.scala and builders.scala define the typed request objects and their validators (time window, page size, free text, service name). - ObservabilityResources exposes /observability/health and is registered in TexeraWebApplication. - RequestContextMdcFilter and UserContextMdcFilter inject the request and user context into the logging MDC. - ObservabilityGatewayConfig plus observability-gateway.conf hold the gateway settings. Dashboard shell (frontend): - Admin-only Observability page, route, and navigation entry, plus the observability service, its types, and the traces-pivot service. - Each tab is guarded by the per-signal reachability check: an unreachable signal renders an explicit state rather than a broken panel. The signal panels themselves follow in later PRs of the stack. Tests: backend specs for the gateway core, DTO validation, scope resolver, rate limiter, and the MDC filters, plus frontend component and service specs. sbt scalafmtCheckAll is clean. Stacked on PR 3 (obs/pr3/deployment). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the first complete signal: log search through the query gateway plus the UI panel that drives it. Builds on the gateway core and dashboard shell from PR 4. This is PR 5 of the observability stack. Gateway (amber): - ResponseParsers turns the log store's responses into the typed results the UI consumes (parseLogs, parseLogSources), applying per-field redaction so sensitive values never reach the client. - builders and dtos gain the logs query builder and its validated request objects (time window, service, workflow, computing unit, execution, level, free text, page size). - ObservabilityResources exposes the log-search and source-facets endpoints, registered in TexeraWebApplication. Logs panel (frontend): - A logs panel under the observability page with time-window, service, workflow, computing-unit, execution, level, and free-text filters, server-side paging, and sources-backed autofill. - observability-prefs persists the panel's filter choices, and the observability component hosts the panel behind its reachability gate. Tests: backend specs for the logs query builder and the response parsers, plus the frontend logs-panel and service specs. sbt scalafmtCheckAll is clean. Stacked on PR 4 (obs/pr4/gateway-core). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
f9713b3 to
3a5939f
Compare
What changes were proposed in this PR?
First complete signal: log search through the gateway plus a UI panel to drive it.
Backend:
parseLogs,parseLogSources), applying per-field redaction.LogsResourceexposing log search and the source-facets endpoint, registered inTexeraWebApplication.Frontend:
Any related issues, documentation, or discussions?
Closes: #5371
Part of #4070. Stacked on #5378.
How was this PR tested?
sbt scalafmtCheckAllpasses.prettier-eslintandeslintpass.Was this PR authored or co-authored using generative AI tooling?
Co-authored with Claude Opus 4.8 in compliance with ASF