Skip to content

statistics: browser-safe subpath, interpolating median, explicit empty-input policy #758

Description

@drewstone

Ask

Expose the descriptive statistics through a browser-safe subpath (for example @tangle-network/agent-eval/statistics), and give it an interpolating median and a stated empty-input policy, so a consumer can replace its local statistics module with the package instead of keeping a copy.

Where this comes from

tangle-network/blueprint-agent upgraded agent-eval 0.145.14 -> 0.182.0 (ticket tangle-network/blueprint-agent#2463) with the rule "replace a local duplicate with the upstream export only after a parity test passes". Two local modules could not be replaced:

  • scripts/experiments/lib/stats.ts (Node side): mean and wilson now delegate to weightedMean and wilson, whose semantics matched on every input class. median stays local.
  • packages/bench-report/src/stats.ts (isomorphic; bundled into the web client): stays local entirely.

What blocks parity, measured on 0.182.0

  1. No browser-safe entry. package.json exports has 28 subpaths and no statistics entry. Every statistics function ships only through the root barrel, and dist/index.js imports node:fs, node:path, node:crypto and node:child_process directly (plus chunks such as bounded-process-*.js, llm-judge-*.js). The chunk that holds weightedMean, confidenceInterval and summarizeNumberSeries (dist/descriptive-*.js) imports no Node builtin, and the chunk holding wilson (dist/paired-arms-*.js) only reaches errors-*, internal-* and paired-tests-*, so the code is already browser-clean; only the public path is not.
  2. No interpolating median. The only median-like export is summarizeNumberSeries(values).p50, defined as the nearest-rank order statistic (ceil(0.5 * n)-th value). For [1, 2, 3, 4] it answers 2; the conventional median is 2.5. It also sorts non-finite members in rather than dropping them.
  3. Empty-input answers differ across the exports. weightedMean([]) is 0, confidenceInterval([]) is { mean: 0, lower: 0, upper: 0 }, summarizeNumberSeries([]) is null, iqr([]) is 0, wilson(0, 0) is { 0, 0, 0 }. A zero can read as a measured all-zero series (the summarizeNumberSeries docstring says exactly this). A consumer that needs "no data" to stay distinguishable from a real zero has to keep its own convention, which is what both blueprint-agent modules do (0/null in one, NaN in the other).
  4. No plain mean, median, quantile, stddev or cv exports exist; weightedMean stands in for mean.

The parity test that fails today

Strict form of the test that now lives in blueprint-agent as scripts/experiments/lib/__tests__/stats-parity.test.ts (there it pins the divergences so it stays green; here it asserts parity and fails on 0.182.0):

import { describe, it } from 'node:test'
import assert from 'node:assert/strict'
import { confidenceInterval, summarizeNumberSeries, weightedMean } from '@tangle-network/agent-eval'

const EVEN = [1, 2, 3, 4]

describe('agent-eval statistics parity', () => {
  it('median of an even-length series interpolates the two middle values', () => {
    assert.equal(summarizeNumberSeries(EVEN)?.p50, 2.5) // actual: 2 (nearest rank)
  })
  it('median ignores non-finite members', () => {
    assert.equal(summarizeNumberSeries([1, Number.NaN, 3])?.p50, 2) // actual: NaN sorts in
  })
  it('an empty series has no mean and no interval, not a zero one', () => {
    assert.ok(Number.isNaN(weightedMean([]))) // actual: 0
    const empty = confidenceInterval([])
    assert.ok(Number.isNaN(empty.mean) && Number.isNaN(empty.lower) && Number.isNaN(empty.upper)) // actual: { 0, 0, 0 }
  })
  it('the statistics resolve without Node builtins', async () => {
    // A `@tangle-network/agent-eval/statistics` entry does not exist on 0.182.0.
    await assert.doesNotReject(import('@tangle-network/agent-eval/statistics'))
  })
})

Run against @tangle-network/agent-eval@0.182.0 with tsx --test: 4 failing, 0 passing.

Proposed shape

  • ./statistics subpath re-exporting weightedMean, confidenceInterval, summarizeNumberSeries, iqr, wilson, pairedBootstrap, pearsonR, spearmanR, weightedComposite, mulberry32, plus a median(values, { empty?: 'nan' | 'null' }) and quantile(values, q) with linear interpolation.
  • One documented empty-input policy per function, or an empty option where callers need a sentinel.

Happy to send the PR if the shape is accepted.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions