Skip to content

serving: inference miners serve full systems, validators probe them on the trace, and systems are held to a measured latency ceiling - #135

Merged
BANADDA merged 1 commit into
mainfrom
serve-systems
Sep 30, 2026
Merged

BANADDA merged 1 commit into
mainfrom
serve-systems

Conversation

@BANADDA

@BANADDA BANADDA commented Sep 30, 2026

Copy link
Copy Markdown
Member

Makes the posted framework true end to end: inference miners run full systems for customers, not just single models, and speed now costs rewards.

Serving systems

  • SystemEngine runs a certified full system (small model, harness, router, escalation) behind a catalogue model name. mt operator run --serve MODEL=DIGEST@system:/path --escalation-url URL --gpu-layers -1 serves it through the normal agent path, so customer traffic through the gateway is unchanged.
  • Responses carry the answer, usage including escalation tokens, and a trace signed by the operator.
  • open_system accepts an archive record or an artifact directory, checked against its manifest. The small model can offload to the GPU for serving, while validators still replay on CPU.

Probing system operators

  • judge_full_system checks that the trace is signed by the operator and names the certified system, and that the returned answer is the traced final answer. It then runs the step 5 checks on the certified archive: small model replay, router recompute, declared escalation, harness rerun.
  • The probe loop uses it for any artifact that is a full system, with no statistical calibration needed.

Latency counts

  • run_live records the validator's own round-trip time for every task. A full system whose end to end p95 is over the arena's latency ceiling scores zero for the round.
  • The p95 and each request's measured time reach the dashboard rows.

Checked on a real Qwen3 0.6B: an operator served a certified system through serve_taken (answer "blue", usage counted, signed trace attached), and the probe passed. The same operator running Qwen3.5 weights under the certified system's name was judged a cheat, with a worst margin of 22.45 logits.

Docs: miner_setup.md describes the full system miner, inference_miner.md gains section 7a (serving full systems) and the trace based verification table, and the playbook adds the latency ceiling and serving.

🤖 Generated with Claude Code

…n the trace, and systems are held to a measured latency ceiling
@BANADDA
BANADDA merged commit fffdadf into main Sep 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant