Skip to content

feat(observability): deployment (OTel collector + Parca/eBPF) - #5377

Open
Ma77Ball wants to merge 3 commits into
apache:mainfrom
Ma77Ball:obs/pr3/deployment
Open

Ma77Ball wants to merge 3 commits into
apache:mainfrom
Ma77Ball:obs/pr3/deployment

Conversation

@Ma77Ball

@Ma77Ball Ma77Ball commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

What changes were proposed in this PR?

Provides the local deployment wiring that receives and stores telemetry. Off by default and isolated from the application services.

  • Adds an OpenTelemetry Collector configuration and wires it into the single-node docker-compose stack and the Kubernetes Helm chart.
  • Adds the Parca server configuration and the Parca eBPF agent for continuous profiling.
  • Adds the observability backends to the single-node .env and docker-compose as opt-in compose profiles; COMPOSE_PROFILES is empty by default, so docker compose up runs only core Texera.
  • Infrastructure only; the application runs unchanged whether or not these services are started.

Any related issues, documentation, or discussions?

Closes: #5369
Part of #4070. Stacked on #5376.

How was this PR tested?

  • Configuration-validation specs for the collector and Parca config (ObservabilityComposeSpec, ParcaConfigSpec); 24 tests pass.
  • sbt scalafmtCheckAll passes; the build runs in this PR's CI.

Was this PR authored or co-authored using generative AI tooling?

Co-authored with Claude Opus 4.8 in compliance with ASF policy.

@github-actions github-actions Bot added dependencies Pull requests that update a dependency file docs Changes related to documentations infra common labels Jun 5, 2026
@codecov-commenter

codecov-commenter commented Jun 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 80.37825% with 83 lines in your changes missing coverage. Please review.
✅ Project coverage is 92.52%. Comparing base (b143a4f) to head (3523acd).
⚠️ Report is 4 commits behind head on main.

Files with missing lines Patch % Lines
...ala/org/apache/texera/observability/OtelInit.scala 70.86% 25 Missing and 12 partials ⚠️
...rg/apache/texera/web/service/WorkflowService.scala 74.13% 11 Missing and 4 partials ⚠️
.../apache/texera/observability/WorkflowMetrics.scala 84.93% 4 Missing and 7 partials ⚠️
...e/texera/observability/TexeraOtelLogAppender.scala 78.04% 2 Missing and 7 partials ⚠️
...org/apache/texera/observability/TexeraTracer.scala 0.00% 4 Missing ⚠️
...ra/web/observability/WorkflowMetricsRecorder.scala 83.33% 2 Missing and 1 partial ⚠️
...org/apache/texera/observability/LogSanitizer.scala 92.50% 0 Missing and 3 partials ⚠️
.../scala/org/apache/texera/service/FileService.scala 0.00% 1 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##               main    #5377      +/-   ##
============================================
- Coverage     92.60%   92.52%   -0.08%     
- Complexity     5043     5061      +18     
============================================
  Files          1252     1261       +9     
  Lines         53587    53963     +376     
  Branches       6673     6744      +71     
============================================
+ Hits          49623    49930     +307     
- Misses         2315     2353      +38     
- Partials       1649     1680      +31     
Flag Coverage Δ
access-control-service 77.44% <100.00%> (+0.06%) ⬆️
agent-service 99.16% <ø> (ø)
amber 87.98% <80.33%> (-0.11%) ⬇️
computing-unit-managing-service 60.52% <100.00%> (+0.03%) ⬆️
config-service 87.50% <100.00%> (+0.12%) ⬆️
file-service 81.45% <0.00%> (-0.09%) ⬇️
frontend 96.58% <ø> (-0.01%) ⬇️
notebook-migration-service 83.77% <100.00%> (+0.03%) ⬆️
pyamber 98.52% <ø> (-0.07%) ⬇️
workflow-compiling-service 74.25% <100.00%> (+0.15%) ⬆️

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions github-actions Bot added the platform Non-amber Scala service paths label Jun 5, 2026
@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

⚠️ Benchmark changes need a look

🟢 0 better · 🔴 8 worse · ⚪ 7 noise (<±5%) · 0 without baseline

Compared against main be65bf8 benchmarked on this same runner, so the delta is largely free of cross-runner hardware noise. The "7d avg" column still reflects the gh-pages dashboard. Treat <±5% as noise unless repeated.

Dashboard · Run

config throughput MB/s latency max Δ latest / 7d
🔴 bs=10 sw=10 sl=64 359 0.219 25,463/39,456/39,456 us 🔴 +17.1% / 🔴 +168.9%
🔴 bs=100 sw=10 sl=64 786 0.48 126,112/142,588/142,588 us 🔴 +5.9% / 🔴 +39.3%
🔴 bs=1000 sw=10 sl=64 895 0.546 1,116,208/1,215,976/1,215,976 us 🔴 +8.4% / 🔴 +23.5%
Baseline details

Latest main be65bf8 from same runner

config metric PR latest main 7d avg Δ latest Δ 7d
bs=10 sw=10 sl=64 throughput 359 tuples/sec 407 tuples/sec 866.99 tuples/sec -11.8% -58.6%
bs=10 sw=10 sl=64 MB/s 0.219 MB/s 0.248 MB/s 0.529 MB/s -11.7% -58.6%
bs=10 sw=10 sl=64 p50 25,463 us 22,238 us 11,989 us +14.5% +112.4%
bs=10 sw=10 sl=64 p95 39,456 us 33,684 us 14,672 us +17.1% +168.9%
bs=10 sw=10 sl=64 p99 39,456 us 33,684 us 17,906 us +17.1% +120.4%
bs=100 sw=10 sl=64 throughput 786 tuples/sec 827 tuples/sec 1,111 tuples/sec -5.0% -29.3%
bs=100 sw=10 sl=64 MB/s 0.48 MB/s 0.505 MB/s 0.678 MB/s -5.0% -29.2%
bs=100 sw=10 sl=64 p50 126,112 us 119,115 us 95,405 us +5.9% +32.2%
bs=100 sw=10 sl=64 p95 142,588 us 146,009 us 102,332 us -2.3% +39.3%
bs=100 sw=10 sl=64 p99 142,588 us 146,009 us 115,219 us -2.3% +23.8%
bs=1000 sw=10 sl=64 throughput 895 tuples/sec 912 tuples/sec 1,141 tuples/sec -1.9% -21.5%
bs=1000 sw=10 sl=64 MB/s 0.546 MB/s 0.557 MB/s 0.696 MB/s -2.0% -21.6%
bs=1000 sw=10 sl=64 p50 1,116,208 us 1,099,146 us 938,794 us +1.6% +18.9%
bs=1000 sw=10 sl=64 p95 1,215,976 us 1,121,883 us 984,521 us +8.4% +23.5%
bs=1000 sw=10 sl=64 p99 1,215,976 us 1,121,883 us 1,008,827 us +8.4% +20.5%
Raw CSV
config_idx,batch_size,schema_width,string_len,num_batches,total_ms,total_tuples,total_bytes,tuples_per_sec,mb_per_sec,lat_p50_us,lat_p95_us,lat_p99_us
0,10,10,64,20,556.57,200,128000,359,0.219,25462.92,39456.16,39456.16
1,100,10,64,20,2545.78,2000,1280000,786,0.480,126111.86,142587.79,142587.79
2,1000,10,64,20,22349.80,20000,12800000,895,0.546,1116208.19,1215976.00,1215976.00

@github-actions

github-actions Bot commented Jun 23, 2026 •

Copy link
Copy Markdown
Contributor

Automated Reviewer Suggestions

Based on the git blame history of the changed files, we recommend the following reviewers:

  • Committers with relevant context: @parshimers
    You can request their reviews formally with /request-review @parshimers.

  • Contributors with relevant context: @bobbai00, @dzueck, @aicam
    You can notify them by mentioning @bobbai00, @dzueck, @aicam in a comment.

@zuozhiw zuozhiw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

left some comments, please fix them, I clicked "approve" so that you can feel free to merge this PR after fixing these commments, because this is a env setup PR, please make sure to carefully make AI actually test these configurations and make sure they work

apart from that I have some reservations about ebpf, it's more sensitive than application level log/traces/metrics. Furthermore, ebpf is normally used for very advanced diagnostics and profiling, and seems a bit heavy for our use case. But again feel free to try it and see if we see it being useful.

Also one more missing piece of the telemetry data is host level metrics, e.g. cpu, memory being the most important ones, please check how we can collect them, e.g. I think otel has some host metrics receiver, please check that, also check if docker / k8s expose metrics and how we can get them.

Another missing piece is database related metrics. E.g. database load, it would also be nice to consider them, but maybe in future PRs.

Comment thread bin/single-node/docker-compose.yml Outdated
sh -c 'apk add --no-cache curl jq bash > /dev/null 2>&1 && bash /examples/load-examples.sh'

# ========================================================================
# Part 5: Observability stack (PR 6).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove these descriptions about "PR 6", the code comments should be more concise and factual and it's pointless to carry such information, claude nowadays is not good at writing good and concise comments, ask claude to pay more attention to it

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the "PR 6" references and tightened the observability comments to state only what each service does.

security_opt:
- no-new-privileges:true
volumes:
- ../observability/otel-collector/config.yaml:/etc/otelcol-contrib/config.yaml:ro,z

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This path might not exist in Texera’s published Docker Compose bundle. The release workflow currently archives only bin/single-node/, sql/, and NOTICE, while this mount depends on bin/observability/otel-collector/config.yaml.

can you make sure to let AI actually run and test both single node and local dev release bundle workflows with these new files?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The release workflow now archives bin/observability and rewrites the ../observability mount paths to ./observability, so the compose resolves from the released single-node bundle.

Comment thread bin/single-node/.env Outdated
# TEXERA_OBSERVABILITY_* env vars below into the right COMPOSE_PROFILES.
# * Or edit COMPOSE_PROFILES directly here.
#
# Disable env-var conventions (consumed by up.sh):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think we should by default turn on the entire observability stack, it's more for deployments, and when we deploy, each deployment should override these configurations.

These profiles add the collector, three signal backends, Parca, and a privileged Parca agent to every single-node installation. Their configured memory limits alone total roughly 5.3 GB, while Texera documents 4 GB as the minimum for the complete single-node deployment. I think these services should be opt-in, or enabled through an explicit installation option. Remember we are open source and we might serve various users who might not need observability, unless they are hosting a service.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Observability is now opt-in: COMPOSE_PROFILES defaults to empty, so no observability service starts unless the operator enables the profiles.

Comment thread bin/single-node/.env Outdated
# TEXERA_OBSERVABILITY_METRICS=disabled drops victoriametrics
# TEXERA_OBSERVABILITY_TRACES=disabled drops jaeger
# TEXERA_OBSERVABILITY_PROFILES=disabled drops parca + parca-agent
# TEXERA_OBSERVABILITY_COLLECTOR=disabled drops the otel-collector (rare)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These variables might not be consumed by Texera’s official bin/single-node.sh path. That entry point delegates to bin/single-node/main.sh, which does not inspect any TEXERA_OBSERVABILITY_* variables.

The referenced up.sh seems to no longer be the canonical launcher, so commands such as TEXERA_OBSERVABILITY_TRACES=disabled bin/single-node.sh up do not disable the profile as documented. Please integrate this behavior into the canonical launcher and test it there. Please double check, I'm not very familiar with the current launching process

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dropped the TEXERA_OBSERVABILITY_* / up.sh convention and removed up.sh; COMPOSE_PROFILES, which docker compose reads natively through the canonical launcher, is now the single switch.

Comment thread bin/single-node/docker-compose.yml Outdated
# IntelliJ) can emit telemetry. In compose-only deploys, services
# talk to otel-collector:4317/:4318 via the bridge network.
ports:
- "127.0.0.1:4317:4317"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldn't local dev override config be in the local dev docker override file?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Moved the OTLP host-port publishing into bin/single-node/docker-compose.override.yml, auto-loaded for local dev; the base compose no longer publishes them.

# posture. The privileged + bind-mount block here is the ONLY
# observability service that requires elevated permissions; the
# surface is documented and reviewed.
parca-agent:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ebpf is much more sensitive than logs/metrics/traces, because those are application level telemetry, ebpf is host level telemetry and involves much high privileges. What do you mean by "the surface is documented and reviewed"? documented where and how is it reviewed?

also we are turning it on by default, and the comment below explicitly say that it requires sys admin class privilege and macos/windows cannot run, so I really don't think we need to turn it on by default, this should be an opt-in feature and document the security implications before the user enables it.

also make sure we have proper default overrides in the local dev overrides

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Parca and the eBPF agent are now opt-in behind the observability-profiles profile, off by default; removed the "documented and reviewed" wording and documented the host-wide scope and privilege/platform implications before enabling.

privileged: true
pid: "host"
# Bind-mount kernel state read-only — the agent reads /proc and
# /sys for stack-trace symbolization but cannot write to either.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you ask more AI, maybe different models to double check and fact check and think harder on this statement?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reworded the comment to state only the kernel-enforced read-only (:ro) behavior of the /proc and /sys mounts.

Comment thread bin/observability/parca/README.md Outdated
`parca-agent.env`:

- `deployment=texera`
- `cluster=local` (override per env)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

make sure that parca really only collects texera process and not other processes in the host.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Corrected the README: the agent profiles every process on the host, not only Texera; labels only filter what a query displays, not what is collected.

Comment thread bin/observability/parca/README.md Outdated
- `deployment=texera`
- `cluster=local` (override per env)

When the PR 7 Texera query gateway runs Parca queries, it filters on

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

again the readme file should not contain any thing about "pr5, pr6, pr7", make absolutely sure to press claude to carefully inspect the writing style of these readmes and code comments! this is an important global comment for all the PRs!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed all inter-PR references from the README.

Comment thread bin/observability/parca/README.md Outdated
We do **not** label profiles with `workflow.id` / `execution.id`. As
with metrics, those are unbounded identifiers and would blow up
Parca's storage cardinality. Per-execution profile views are reached
by joining on `trace_id` at query time (the Parca query API supports

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please fact check this statement that the ebpf collections can join with trace_id, how does it know our application level trace id? make sure test it

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Corrected the README: the agent samples stack traces and does not read application-level OpenTelemetry trace ids, so profiles cannot be joined to individual traces by trace_id.

Introduce the OpenTelemetry foundation for Texera: an SDK bootstrap, a
Logback-to-OTel log bridge, and a log sanitizer, wired into every service
entry point. This is PR 1 of the observability stack; it ships logging only,
with the trace and metric exporters wired but not yet emitted (those arrive in
follow-up PRs).

New module (common/observability):
- OtelInit: one-call SDK bootstrap per service. Reads OTEL_* settings from
  observability.conf, validates the OTLP endpoint (scheme + host allowlist)
  before any exporter is built, wires the span/log/metric providers explicitly
  (no sdk-extension-autoconfigure), clamps the metric export interval, and
  attaches the log appender to the Logback ROOT logger. The whole init body is
  guarded so a missing or malformed config returns None with one WARN instead
  of throwing into a service's run(), even when telemetry is disabled.
- TexeraOtelLogAppender: bridges Logback events to OTel log records. Maps
  severity, forwards MDC, sets exception.type/message/stacktrace semantic
  attributes, and drops records from io.opentelemetry loggers so export-failure
  diagnostics are not fed back to the collector that just failed.
- LogSanitizer: strips C0 control characters, redacts secrets, and caps body
  length (MaxBodyChars) to keep individual records bounded.

Config (common/config):
- observability.conf with the OTEL_* defaults, ObservabilityConfig to read it,
  and ENV_OTEL_* entries in EnvironmentalVariable.

Wiring:
- build.sbt defines the Observability module and adds dependsOn(Observability)
  to the eight Dropwizard services.
- Each service entry point calls OtelInit.init with its own service name, and
  each service config gains a logging block.

Deployment:
- OTEL_* env entries in bin/single-node/.env, bin/k8s/values.yaml, and
  bin/k8s/values-development.yaml (kept a name-for-name mirror).
- LICENSE-binary manifests updated with the pinned OTel 1.50.0 jars.

Tests: OtelInitSpec, TexeraOtelLogAppenderSpec, LogSanitizerSpec, and
ObservabilityConfigSpec cover endpoint validation, interval clamping,
severity/MDC/exception mapping, self-diagnostic filtering, redaction, and
truncation.

Rebased onto current main.
Build on the logging foundation (PR 1) with the backend emit path: workflow
lifecycle metrics and a run-level setup trace span, plus the tracing/metrics
primitives they use. This is PR 2 of the observability stack.

Primitives (common/observability):
- TexeraTracer: lazy accessor for the process tracer off GlobalOpenTelemetry.
- SpanAttrs: typed AttributeKey constants so span attributes use standard keys
  rather than ad hoc strings.
- WorkflowMetrics: the OTel instruments (start/completion/failure/cancellation
  counters and run-duration histogram) keyed by workflow kind.
- TraceparentValidator: validates W3C traceparent headers before use.

Emit path (amber):
- WorkflowMetricsRecorder: single owner of the metric instruments. init() wires
  them once; onStart stamps a run's start; onStateChange records terminal
  counters and duration exactly once on the first transition into a terminal
  state (idempotent, safe to call on every transition).
- WorkflowService.initExecutionService runs inside a run-level setup span so
  setup-path logs carry its trace id. The span covers only the synchronous
  setup; synchronous setup failures are recorded on it in the catch block, and
  the span is ended in finally. The async errorHandler deliberately does not
  touch the span: it is invoked after setup returns and the span has ended, so
  the failure is surfaced through the metadata store instead.
- ExecutionStateStore.updateWorkflowState is the single chokepoint that feeds
  every state transition to WorkflowMetricsRecorder.onStateChange.
- ComputingUnitMaster initializes the recorder at startup.

All observability sources live under common/observability (the module from
PR 1); the tracing/metrics classes are not duplicated into common/config.

Tests: WorkflowMetricsSpec, SpanAttrsSpec, and TraceparentValidatorSpec.

Review follow-ups addressed: span attributes use SpanAttrs keys instead of
plain strings; the error handler no longer records onto a span that may have
already ended (documented and handled via the metadata store).

Stacked on PR 1 (obs/pr1/foundations).
@Ma77Ball
Ma77Ball marked this pull request as ready for review October 9, 2026 10:58
@Ma77Ball
Ma77Ball force-pushed the obs/pr3/deployment branch 2 times, most recently from a28973f to eea0a7a Compare October 9, 2026 18:00
Add the local deployment wiring that receives and stores the telemetry
emitted by PR 2: an OpenTelemetry Collector plus logs, metrics, traces,
and profiles backends, for both the single-node docker-compose stack and
the Kubernetes Helm chart. Infrastructure only; the application runs
unchanged whether or not these services are started. This is PR 3 of the
observability stack.

Single-node compose (bin/single-node, bin/observability):
- otel-collector as the single OTLP ingress, routing to VictoriaLogs
  (logs), VictoriaMetrics (metrics), and Jaeger (traces); the collector
  config caps OTLP message size and bounds memory via memory_limiter.
- Parca server + eBPF agent for continuous profiling, configured in
  bin/observability/parca.
- Each signal is its own compose profile and is OFF by default:
  COMPOSE_PROFILES is empty, so observability is opt-in. docker compose
  reads COMPOSE_PROFILES natively via the canonical launcher.
- Emission is gated by OTEL_SDK_DISABLED (default disabled); nothing is
  emitted unless an operator opts in alongside the collector profile.

Kubernetes Helm chart (bin/k8s):
- Templates for the collector, Jaeger, VictoriaLogs, and VictoriaMetrics
  under templates/base/observability, plus values.yaml knobs and an
  install.sh helper.
- Service deployments gain the OTLP endpoint wiring.

Every observability port binds to loopback (127.0.0.1) or stays on the
bridge network; no 0.0.0.0 host bindings.

Tests: ObservabilityComposeSpec and ParcaConfigSpec (string-level smoke
tests over the compose, collector, and Parca config, pinning image
versions, the opt-in default, loopback binding, and the no-high-
cardinality-label rule). All 24 pass; sbt scalafmtCheckAll is clean.

Review follow-ups addressed: observability is opt-in rather than on by
default, and the privileged eBPF agent is opt-in with its host-wide scope
documented; removed the non-canonical up.sh launcher so COMPOSE_PROFILES
is the single switch; corrected the Parca README (the agent profiles every
host process, and profiles cannot be joined to traces by trace_id); moved
the local-dev-only OTLP host ports into docker-compose.override.yml so the
base compose stays deploy-clean; the release workflow now ships
bin/observability and rewrites the mount paths so the compose works from
the released single-node bundle; and removed the inter-PR references from
the configs, comments, and specs.

Stacked on PR 2 (obs/pr2/backend-emit).
@Ma77Ball
Ma77Ball force-pushed the obs/pr3/deployment branch from eea0a7a to 3523acd Compare October 9, 2026 18:40

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci changes related to CI common dependencies Pull requests that update a dependency file docs Changes related to documentations engine infra platform Non-amber Scala service paths

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Observability] Deploy OpenTelemetry Collector and Parca/eBPF profiling backends

3 participants