Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/full_system_playbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@ The router is what turns confidence into money. It sends only the requests the s
6. **Calibrate and quantise.** Fold temperature into the file (`scripts/fold_temperature.py`) so the confidence validators recompute is already calibrated. Quantise with the output layer kept at Q8 (`--output-tensor-type q8_0`, or `--token-embedding-type q8_0` for tied embeddings).
7. **Simulate the whole system.** `mt miner simulate --escalation-url URL` runs the small model, harness, router and escalation locally over the train split and prints every trace with the scores validators will compute: end to end quality, the small model alone, escalation rate, waste, misses, calibration and cost.
8. **Package and host.** Write `system.json` at schema version 2 (model, harness with its runtime, router with its features, escalation model pinned to a revision, endpoint), then serve it through the dial out agent. No public IP is needed.
9. **Serve it to customers.** Once certified, inference miners serve your system under the arena's catalogue name on GPUs (`inference_miner.md`, section 7a), and we can serve it from the archive after you stop.

## What validators check

Expand All @@ -32,6 +33,7 @@ The router is what turns confidence into money. It sends only the requests the s
- **Your escalation is honest.** Only your declared, allowlisted model, charged at its published price.
- **The archive reproduces the live run.** Your archived harness, router and small model are rerun on sampled tasks; prompts, tokens, decisions and final answers must match.
- **Your harness stays inside its package.** No URLs, no network or process modules, no `eval`.
- **Your system is fast enough.** Validators time every request themselves; a system whose end to end p95 is over the arena's latency ceiling earns nothing that round.

A system that fails any check is not certified and earns nothing that round.

Expand Down
41 changes: 41 additions & 0 deletions docs/inference_miner.md
Original file line number Diff line number Diff line change
Expand Up @@ -373,6 +373,33 @@ up. Reconnect is exponential to a 60 second ceiling.

---

## 7a · Serving full systems

A certified model can be a full system: a small specialist, its harness, a
router and an escalation model. You serve it exactly like a model, under its
catalogue name, and every request runs the whole system on your machine.

```bash
mt operator run \
--serve mt/invoice-4g=sha256:<system digest>@system:/srv/systems/invoice-4g \
--escalation-url http://127.0.0.1:18090 \
--gpu-layers -1
```

- `system:<path>` points at the certified system, either an archive record or an
artifact directory with its `manifest.json`. It is checked against the
manifest before it serves anything.
- `--escalation-url` is an OpenAI compatible server holding our mirror of the
system's escalation model (the name is `microtensor-archive/<model>`). Run it
on the same card with SGLang, or on another card in the same box.
- `--gpu-layers -1` puts the small model on the GPU. Validators replay it on CPU;
the tolerance covers the difference.

Each response carries the answer, the usage including escalation tokens, and a
trace signed by your hotkey: the small model's answer, confidence and tokens,
the router's features and decision, harness steps, and the escalation answer if
there was one. Clients see a single answer.

## 8 · How you are verified

You send tokens. You build no proof and compute no commitment. The validator
Expand All @@ -398,6 +425,20 @@ sampled statistics that disagree. It is counted and never gates.
Admission is a sequential test, not a fixed count. A clean miner is usually
admitted in well under a hundred probes. One validator rejecting keeps you out.

**Full systems are judged on the trace.** The validator checks that the trace is
signed by you and names the certified system, that the answer you returned is
the traced final answer, and then:

| check | how |
|---|---|
| small model | the traced tokens are replayed on the certified archive on CPU; a token more than 0.5 logits below the model's own choice, or a confidence off by more than 0.02, is a cheat |
| router | its features are recomputed from the replayed small model and its declared rule must give the same decision |
| escalation | only the system's declared model |
| harness | the certified harness is rerun on the same request and must return the same answer |

Serving other weights under a certified system's name fails the first check at
once: a swapped small model shows margins of tens of logits.

```bash
mt operator status
```
Expand Down
11 changes: 11 additions & 0 deletions docs/miner_setup.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,16 @@ artifact, and commit a pointer on chain once per round. Validators fetch it and
run it on their own certified hardware, so your machine is busy only while you
are training.

In the full system arenas you submit more than a model. A system is four parts
built for one task: the specialist small model with calibrated confidence, the
harness around it (prompts, tools, checks and output templates), a router that
reads the small model's confidence and decides whether to answer, and an
escalation model from the arena's allowlist of open models that takes over only
when the router sends a request up. You host the system through the dial out
agent, validators test it live on withheld tasks, and it is ranked on end to end
quality against total cost within a latency ceiling. Start with
[full_system_playbook.md](full_system_playbook.md).

**You have a GPU and want it earning continuously without training.** You are
an inference miner. You post collateral, run a stock engine on certified
artifacts, and answer live traffic. You do not choose which models exist and you
Expand Down Expand Up @@ -148,6 +158,7 @@ and nothing more. You start the next one clean.
| | |
|---|---|
| [system_miner.md](system_miner.md) | Train, package, publish and commit, round by round |
| [full_system_playbook.md](full_system_playbook.md) | Build a full system: small model, harness, router and escalation |
| [inference_miner.md](inference_miner.md) | Register, post collateral, serve and get verified |
| [compute_miner.md](compute_miner.md) | Enrol a GPU machine in the pool |
| [validator_setup.md](validator_setup.md) | Evaluate submissions and verify serving |
Expand Down
45 changes: 41 additions & 4 deletions microtensor/cli/operator.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,15 @@
from microtensor.serving import audit as serving_audit
from microtensor.serving import client, plan, supervise
from microtensor.serving import loop as probe_loop
from microtensor.serving.agent import AgentError, Pool, Served, Settings, run, served
from microtensor.serving.agent import (
AgentError,
HttpEngine,
Pool,
Served,
Settings,
run,
served,
)
from microtensor.serving.client import ServerError

log = logging.getLogger("microtensor.cli.operator")
Expand Down Expand Up @@ -253,7 +261,7 @@ def _run(args: argparse.Namespace) -> int:
except (OSError, ValueError) as exc:
return fail(f"could not read the pool: {exc}")

pool = Pool(settings.serves)
pool = Pool(settings.serves, build=_engines(args, wallet))
try:
asyncio.run(_serve(settings, pool, systems))
except KeyboardInterrupt:
Expand All @@ -263,6 +271,33 @@ def _run(args: argparse.Namespace) -> int:
return 0


SYSTEM_SCHEME = "system:"


def _engines(args: argparse.Namespace, wallet: Any) -> Any:
def build(url: str) -> Any:
if not url.startswith(SYSTEM_SCHEME):
return HttpEngine(url)
from microtensor.harness.sdk import openai_escalation
from microtensor.miner.host import wallet_signer
from microtensor.serving.archived import SystemEngine, open_system, runtime_for

if not args.escalation_url:
raise AgentError("a served system escalates to our mirrors; pass --escalation-url")
restored = open_system(Path(url[len(SYSTEM_SCHEME) :]), Path(args.restore_dir))
system = restored.manifest.system
if system is None or system.escalation is None:
raise AgentError(f"{url} is not a full system")
mirrored = f"{args.mirror_org}/{system.escalation.model.split('/', 1)[-1]}"
escalate = openai_escalation(args.escalation_url, mirrored)
runtime = runtime_for(
restored, escalate, hotkey=hotkey_address(wallet), gpu_layers=args.gpu_layers
)
return SystemEngine(runtime, wallet_signer(wallet), url)

return build


def _archived(args: argparse.Namespace, wallet: Any) -> dict[str, Any]:
from microtensor.harness.sdk import openai_escalation
from microtensor.miner.host import wallet_signer
Expand Down Expand Up @@ -311,10 +346,12 @@ def _verify(args: argparse.Namespace) -> int:
calibrations = probe_loop.load_calibrations(args.calibrations)
if not artifacts:
return fail(f"no artifact paths in {args.artifacts}")
if not calibrations:
if not calibrations and not all(probe_loop.is_system(p) for p in artifacts.values()):
return fail(f"no calibrations in {args.calibrations}")

missing = sorted(set(artifacts) - set(calibrations))
missing = sorted(
m for m in set(artifacts) - set(calibrations) if not probe_loop.is_system(artifacts[m])
)
if missing:
return fail(f"no calibration for {', '.join(missing)}; an uncalibrated model cannot judge")

Expand Down
7 changes: 5 additions & 2 deletions microtensor/harness/engines/gguf.py
Original file line number Diff line number Diff line change
Expand Up @@ -239,9 +239,12 @@ def __call__(self, input_ids: Any, logits: Any) -> bool:
class GgufEngine:
format = ArtifactFormat.GGUF

def __init__(self, *, threads: int = THREADS, validate: bool = True) -> None:
def __init__(
self, *, threads: int = THREADS, validate: bool = True, gpu_layers: int = GPU_LAYERS
) -> None:
self._threads = threads
self._validate = validate
self._gpu_layers = gpu_layers
self._model: Any = None
self._manifest: LoadManifest | None = None
self._answer_ids: dict[str, int] = {}
Expand All @@ -267,7 +270,7 @@ def load(self, artifact: Path, manifest: LoadManifest) -> None:
n_ctx=context,
n_threads=self._threads,
n_threads_batch=self._threads,
n_gpu_layers=GPU_LAYERS,
n_gpu_layers=self._gpu_layers,
seed=SEED,
logits_all=False,
embedding=False,
Expand Down
4 changes: 4 additions & 0 deletions microtensor/serving/agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -527,6 +527,10 @@ async def serve_taken(
answer["answered_by"] = str(found.get("answered_by", "front"))
if found.get("router_features"):
answer["router_features"] = dict(found["router_features"])
if found.get("trace"):
answer["trace"] = dict(found["trace"])
if found.get("usage"):
answer["usage"] = dict(found["usage"])
state.release(model, ok=True)
return answer
except asyncio.CancelledError:
Expand Down
72 changes: 70 additions & 2 deletions microtensor/serving/archived.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@
from microtensor.harness.engines.router import load_router
from microtensor.harness.sdk import EscalationModel, Runtime, engine_small
from microtensor.registry.manifest import ArtifactManifest, verify_tree
from microtensor.serving.agent import Engine

INTAKE = "intake.json"

Expand Down Expand Up @@ -65,13 +66,80 @@ def restore(record: Path, workdir: Path) -> Restored:
return Restored(root=target, manifest=manifest, record=intake)


def runtime_for(restored: Restored, escalate: EscalationModel, *, hotkey: str) -> Runtime:
def open_system(path: Path, workdir: Path) -> Restored:
if (path / INTAKE).is_file():
return restore(path, workdir)
manifest = ArtifactManifest.from_json((path / "manifest.json").read_bytes())
ok, reason = verify_tree(path, manifest)
if not ok:
raise ArchiveError(f"the system files do not match their manifest: {reason}")
if manifest.system is None or not manifest.system.full:
raise ArchiveError(f"{path} is not a full system")
return Restored(root=path, manifest=manifest, record={})


def prompt_of(request: Mapping[str, Any]) -> str:
prompt = str(request.get("prompt", ""))
if prompt:
return prompt
for message in reversed(list(request.get("messages") or [])):
if isinstance(message, Mapping) and message.get("role") == "user":
return str(message.get("content", ""))
raise ArchiveError("the request carries no prompt")


class SystemEngine(Engine):
def __init__(
self, runtime: Runtime, sign: Callable[[Mapping[str, Any]], str], url: str = "system:"
) -> None:
super().__init__(url)
self.runtime = runtime
self.sign = sign

async def generate(self, request: Mapping[str, Any]) -> dict[str, Any]:
import asyncio

trace = await asyncio.to_thread(
self.runtime.run,
0,
str(request.get("request_id") or "request"),
prompt_of(request),
dict(request.get("inputs") or {}),
)
signed = trace.signed_with(self.sign(trace.body()))
escalation = trace.escalation
return {
"text": str(trace.final),
"prompt_tokens": [],
"completion_tokens": list(trace.small.tokens),
"finish_reason": "stop",
"escalated": trace.escalated,
"answered_by": "escalation" if trace.escalated else "small",
"router_features": dict(trace.router.features),
"trace": signed.to_dict(),
"usage": {
"prompt_tokens": trace.small.prompt_tokens
+ (escalation.prompt_tokens if escalation else 0),
"completion_tokens": len(trace.small.tokens)
+ (escalation.completion_tokens if escalation else 0),
},
}

async def stream(self, request: Mapping[str, Any], on_delta: Any) -> dict[str, Any]:
found = await self.generate(request)
await on_delta(found["text"])
return found


def runtime_for(
restored: Restored, escalate: EscalationModel, *, hotkey: str, gpu_layers: int = 0
) -> Runtime:
from microtensor.harness.engines.gguf import GgufEngine

system = restored.manifest.system
if system is None or system.harness is None or system.escalation is None:
raise ArchiveError("the archived submission is not a full system")
engine = GgufEngine()
engine = GgufEngine(gpu_layers=gpu_layers)
engine.load(restored.root, restored.manifest.load)
return Runtime(
restored.root / system.harness.path,
Expand Down
44 changes: 40 additions & 4 deletions microtensor/serving/loop.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
from pathlib import Path
from typing import Any

from microtensor.core.system import FULL_SYSTEM
from microtensor.serving import probe, verify
from microtensor.serving.client import ServerError

Expand Down Expand Up @@ -94,11 +95,12 @@ def cycle(
continue
calibration = calibrations.get(wanted)
artifact = artifacts.get(wanted)
if calibration is None or artifact is None:
system = artifact is not None and is_system(artifact)
if artifact is None or (calibration is None and not system):
log.info("no calibrated artifact for %s; skipping", wanted)
continue
if wanted not in opened:
opened[wanted] = probe.artifact_model(artifact)
opened[wanted] = open_system(artifact) if system else probe.artifact_model(artifact)

for _ in range(per_operator):
verdict = _one(
Expand All @@ -112,6 +114,7 @@ def cycle(
model=wanted,
prompt=probe.prompt_for(chance),
tally=tally,
system=system,
)
if verdict is None:
break
Expand All @@ -131,11 +134,12 @@ def _one(
credential: str,
wallet: Any,
engine: Any,
calibration: verify.Calibration,
calibration: verify.Calibration | None,
hotkey: str,
model: str,
prompt: str,
tally: Counted,
system: bool = False,
) -> dict[str, Any] | None:
try:
answered = probe.ask(gateway, credential, hotkey=hotkey, model=model, prompt=prompt)
Expand All @@ -144,7 +148,10 @@ def _one(
log.info("operator %s did not answer: %s", hotkey[:12], exc)
return None

found = probe.judge(engine, answered, calibration)
if system or calibration is None:
found = probe.judge_full_system(engine, answered, prompt)
else:
found = probe.judge(engine, answered, calibration)
tally.probed += 1

if found.verdict == verify.CHEAT:
Expand Down Expand Up @@ -191,6 +198,35 @@ def _one(
return reported


def is_system(path: Path) -> bool:
if (path / "intake.json").is_file():
return True
manifest = path / "manifest.json"
if not manifest.is_file():
return False
try:
return (
int(
dict(json.loads(manifest.read_text(encoding="utf-8")).get("system") or {}).get(
"schema_version", 1
)
)
>= FULL_SYSTEM
)
except (ValueError, TypeError):
return False


def open_system(path: Path) -> tuple[Any, Any]:
from microtensor.harness.engines.gguf import GgufEngine
from microtensor.serving.archived import open_system as restore

restored = restore(path, path.parent / f".{path.name}-restored")
engine = GgufEngine()
engine.load(restored.root, restored.manifest.load)
return restored, engine


def load_calibrations(path: Path) -> dict[str, verify.Calibration]:
if not path.exists():
return {}
Expand Down
Loading
Loading