Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/full_system_playbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ The router is what turns confidence into money. It sends only the requests the s
5. **Fit the router.** `scripts/train_router.py` records your small model's own results on the train split with every permitted feature (confidence, entropy, input length, how typical the input is, harness errors) and fits a short rule that trades quality against escalations. Typicality is computed out of fold, so the router learns what an unusual input looks like.
6. **Calibrate and quantise.** Fold temperature into the file (`scripts/fold_temperature.py`) so the confidence validators recompute is already calibrated. Quantise with the output layer kept at Q8 (`--output-tensor-type q8_0`, or `--token-embedding-type q8_0` for tied embeddings).
7. **Simulate the whole system.** `mt miner simulate --escalation-url URL` runs the small model, harness, router and escalation locally over the train split and prints every trace with the scores validators will compute: end to end quality, the small model alone, escalation rate, waste, misses, calibration and cost.
8. **Package and host.** Write `system.json` at schema version 2 (model, harness with its runtime, router with its features, escalation model pinned to a revision, endpoint), then serve it through the dial out agent. No public IP is needed.
8. **Package and host.** Write `system.json` at schema version 2 (model, harness with its runtime, router with its features, escalation model pinned to a revision, endpoint), then serve it with `mt miner host` on a Bittensor axon at an IP and port validators can reach, from your commit until the round settles.
9. **Serve it to customers.** Once certified, inference miners serve your system under the arena's catalogue name on GPUs (`inference_miner.md`, section 7a), and we can serve it from the archive after you stop.

## What validators check
Expand Down
4 changes: 2 additions & 2 deletions docs/full_system_rollout.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ Current arenas keep running unchanged. A system arena opens only at a round boun
- `target_quality` and `target_model` for the outside model line.
5. **Mirror every escalation model once:** `mt archive mirror --model org/name --revision <sha>`.
6. **Corpus.** Withheld routine and unusual tasks, with at least 20% unusual, passing `mt corpus check`, and teacher data published for the train split.
7. **Validators.** Set `MT_GATEWAY_URL` and `MT_GATEWAY_SECRET`, and confirm the jail runs a GGUF replay on the Linux host.
7. **Validators.** Live tests go from the validator's dendrite, signed by its wallet, to each miner's axon read from the metagraph; no gateway credential is needed. Confirm the jail runs a GGUF replay on the Linux host.
8. **Coordinator box.** Run `mt archive intake --db <coordinator.sqlite> --watch 300 --reveals --org <private org>` so every submission is copied as it is committed.
9. **Announce one round ahead** (draft below), then open the round the usual way: server open with the config, coordinator open, anchor.

Expand All @@ -50,7 +50,7 @@ Current arenas keep running unchanged. A system arena opens only at a round boun

> **Full system arenas open at round N.**
>
> From round N, the new system arenas take a full system, not a single model: your small specialist model, the harness around it, a router, and an escalation model chosen from the arena's allowlist. You host it through the dial out agent; no public IP is needed.
> From round N, the new system arenas take a full system, not a single model: your small specialist model, the harness around it, a router, and an escalation model chosen from the arena's allowlist. You host it with `mt miner host` on a Bittensor axon at a public IP and port, registered on chain, from your commit until the round settles.
>
> Validators send withheld tasks to your system during the round and check every answer: your small model is replayed from the archive, your router's decisions are recomputed from its declared rule, and escalations are charged at the published price. You are ranked on end to end quality against total cost, and a system that fails a check earns nothing that round.
>
Expand Down
5 changes: 3 additions & 2 deletions docs/miner_setup.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,8 +32,9 @@ built for one task: the specialist small model with calibrated confidence, the
harness around it (prompts, tools, checks and output templates), a router that
reads the small model's confidence and decides whether to answer, and an
escalation model from the arena's allowlist of open models that takes over only
when the router sends a request up. You host the system through the dial out
agent, validators test it live on withheld tasks, and it is ranked on end to end
when the router sends a request up. You host the system on a Bittensor axon at
a public IP and port, validators test it live on withheld tasks from their
dendrites, and it is ranked on end to end
quality against total cost within a latency ceiling. Start with
[full_system_playbook.md](full_system_playbook.md).

Expand Down
16 changes: 11 additions & 5 deletions docs/system_miner.md
Original file line number Diff line number Diff line change
Expand Up @@ -862,13 +862,19 @@ carries `system.json`.
### Host it through the round

```bash
mt miner host --escalation-url http://127.0.0.1:18090 --gpu-layers -1
mt miner host --escalation-url http://127.0.0.1:18090 --port 8091 --gpu-layers -1
```

keeps your system online through the dial out agent under your hotkey. No
inbound port and no public address. **Keep it running from your commit until
the round settles.** During the round validators send it every withheld task,
in the same format real traffic uses, and time each request themselves. Every
serves your system on a Bittensor axon and registers its IP and port on chain
under your hotkey, so validators find it in the metagraph. The port must be
reachable from the internet: open it in your firewall, and pass
`--external-ip` and `--external-port` if you are behind NAT or a forwarded
port. Only validators holding a permit get through; anyone else is refused
before your system runs, and requests are served in order of the caller's
stake. **Keep it running from your commit until the round settles.** During the
round validators call it from their dendrites with every withheld task, signed
by their hotkeys, in the same format real traffic uses, and time each request
themselves. Every
answer goes back with a trace signed by your hotkey: the small model's answer,
confidence and tokens, the router's features and decision, harness steps, the
escalation answer if there was one, and the final answer. A task your system
Expand Down
25 changes: 25 additions & 0 deletions microtensor/chain/synapse.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
from typing import Any

_SYSTEM_TASK: Any = None


def system_task() -> Any:
global _SYSTEM_TASK
if _SYSTEM_TASK is None:
import bittensor as bt
from pydantic import Field

class SystemTask(bt.Synapse): # type: ignore[misc]
system: str = ""
round_index: int = 0
task_ref: str = ""
prompt: str = ""
inputs: dict[str, Any] = Field(default_factory=dict)
trace: dict[str, Any] | None = None
failure: str = ""

def deserialize(self) -> dict[str, Any] | None:
return self.trace

_SYSTEM_TASK = SystemTask
return _SYSTEM_TASK
52 changes: 37 additions & 15 deletions microtensor/cli/miner.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,6 @@
from microtensor.core.constants import (
COORDINATOR_URL,
CORPUS_VERSION,
GATEWAY_URL,
GENESIS_BLOCK,
PROVENANCE_REQUIRED,
PUBLIC_SERVER_URL,
Expand Down Expand Up @@ -182,15 +181,19 @@ def register(subparsers: argparse._SubParsersAction[argparse.ArgumentParser]) ->
serve.set_defaults(handler=_serve)

host = inner.add_parser(
"host", help="keep your full system online for live testing until the round settles"
"host", help="serve your full system on your axon until the round settles"
)
_add_settings_arguments(host)
host.add_argument(
"--escalation-url",
required=True,
help="OpenAI compatible server running your allowlisted escalation model",
)
host.add_argument("--gateway", default=GATEWAY_URL, help="where the dial out agent connects")
host.add_argument("--port", type=int, default=8091, help="the axon port validators reach")
host.add_argument(
"--external-ip", default="", help="the public IP to register, if not detected"
)
host.add_argument("--external-port", type=int, default=0, help="the public port, if forwarded")
host.add_argument(
"--gpu-layers", type=int, default=-1, help="small model layers on the GPU, -1 for all"
)
Expand Down Expand Up @@ -272,13 +275,16 @@ def _wallet_from_saved(args: argparse.Namespace, home: Path) -> None:
setattr(args, attr, stored[attr])


VALIDATOR_REFRESH_SECONDS = 300.0


def _host(args: argparse.Namespace) -> int:
import asyncio
import time

from microtensor.chain.wallet import hotkey_address
from microtensor.harness.sdk import openai_escalation
from microtensor.miner.axon import AxonUnavailable, serve_system
from microtensor.miner.host import system_handler, wallet_signer
from microtensor.serving.agent import AgentError, Pool, Settings, run
from microtensor.serving.archived import ArchiveError, open_system, runtime_for

try:
Expand All @@ -298,20 +304,36 @@ def _host(args: argparse.Namespace) -> int:
gpu_layers=args.gpu_layers,
)
handlers = {system.endpoint.name: system_handler(runtime, wallet_signer(wallet))}
client = open_client(config.chain, wallet)
refreshed = [0.0]
cached: dict[str, float] = {}

def validators() -> dict[str, float]:
if time.monotonic() - refreshed[0] > VALIDATOR_REFRESH_SECONDS:
neurons = client.snapshot(refresh=True).neurons
cached.clear()
cached.update({n.hotkey: float(n.stake) for n in neurons if n.validator_permit})
refreshed[0] = time.monotonic()
return dict(cached)

try:
settings = Settings(
gateway=args.gateway,
hotkey=hotkey,
serves=(),
worker=system.endpoint.worker,
systems=(system.endpoint.name,),
serve_system(
wallet,
client.subtensor,
config.chain.netuid,
handlers,
validators,
port=args.port,
external_ip=args.external_ip,
external_port=args.external_port,
)
except AgentError as exc:
except AxonUnavailable as exc:
return fail(str(exc))
print(f"hosting {system.endpoint.name} ({system.digest()[:19]}) as {hotkey[:12]}")
print("keep this running until your round settles; validators test it live")
print(f"serving {system.endpoint.name} ({system.digest()[:19]}) on axon port {args.port}")
print("your axon is registered on chain; keep this running until your round settles")
try:
asyncio.run(run(settings, Pool(()), systems=handlers))
while True:
time.sleep(60)
except KeyboardInterrupt:
print("stopped")
return 0
Expand Down
73 changes: 73 additions & 0 deletions microtensor/miner/axon.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
from __future__ import annotations

import logging
import typing
from collections.abc import Callable, Mapping
from dataclasses import dataclass
from typing import Any

Expand Down Expand Up @@ -76,3 +78,74 @@ def answer(synapse: TrainingStatus) -> TrainingStatus:
axon.start()
log.info("serving training status on port %d", port)
return axon


DEFAULT_SYSTEM_PORT = 8091


def serve_system(
wallet: Any,
subtensor: Any,
netuid: int,
handlers: Mapping[str, Callable[[Mapping[str, Any]], dict[str, Any]]],
validators: Callable[[], Mapping[str, float]],
*,
port: int = DEFAULT_SYSTEM_PORT,
external_ip: str = "",
external_port: int = 0,
) -> Any:
try:
import bittensor as bt
except ImportError as exc:
raise AxonUnavailable("the system axon needs bittensor: pip install \".[miner]\"") from exc
from microtensor.chain.synapse import system_task

task = system_task()

def forward(synapse: Any) -> Any:
handler = handlers.get(synapse.system)
if handler is None:
synapse.failure = f"this miner hosts no system named {synapse.system!r}"
return synapse
try:
synapse.trace = handler(
{
"round_index": synapse.round_index,
"task_ref": synapse.task_ref,
"prompt": synapse.prompt,
"inputs": dict(synapse.inputs or {}),
}
)
except Exception as exc:
synapse.failure = f"{type(exc).__name__}: {exc}"
return synapse

def caller(synapse: Any) -> str:
return str(getattr(getattr(synapse, "dendrite", None), "hotkey", "") or "")

def blacklist(synapse: Any) -> tuple[bool, str]:
hotkey = caller(synapse)
if not hotkey:
return True, "the request carries no hotkey"
if hotkey not in validators():
return True, "only validators with a permit test this system"
return False, "validator"

def priority(synapse: Any) -> float:
return float(validators().get(caller(synapse), 0.0))

forward.__annotations__ = {"synapse": task, "return": task}
blacklist.__annotations__ = {"synapse": task, "return": typing.Tuple[bool, str]} # noqa: UP006
priority.__annotations__ = {"synapse": task, "return": float}

options: dict[str, Any] = {"wallet": wallet, "port": port}
if external_ip:
options["external_ip"] = external_ip
if external_port:
options["external_port"] = external_port
axon = (getattr(bt, "Axon", None) or bt.axon)(**options)
axon.attach(forward_fn=forward, blacklist_fn=blacklist, priority_fn=priority)
axon.serve(netuid=netuid, subtensor=subtensor)
axon.start()
log.info("serving %s on axon port %d", ", ".join(sorted(handlers)), port)
return axon
17 changes: 9 additions & 8 deletions microtensor/validator/evaluate.py
Original file line number Diff line number Diff line change
Expand Up @@ -649,19 +649,20 @@ def _evaluate_full(
) -> Evaluation:
from microtensor.scoring.metrics import score_task
from microtensor.scoring.system import COST_UNITS_PER_USD, score_system
from microtensor.validator.live import GatewaySystemClient, chain_verifier, run_live
from microtensor.validator.live import AxonSystemClient, chain_verifier, run_live
from microtensor.validator.verify_system import jailed_verify

config = context.config
system = participant.system
if not config.gateway_url or not config.gateway_secret or system.endpoint is None:
raise Abstain(
"full systems are tested live through the gateway; "
"set MT_GATEWAY_URL and MT_GATEWAY_SECRET"
)
client = GatewaySystemClient(
config.gateway_url, config.gateway_secret, participant.hotkey, system.endpoint.worker
if context.wallet is None:
raise Abstain("full systems are tested over the miner's axon; this validator has no wallet")
neuron = next(
(n for n in context.client.snapshot().neurons if n.hotkey == participant.hotkey), None
)
if neuron is None or not neuron.address or neuron.port <= 0:
log.info("%s scored zero: no axon is registered for its system", participant.hotkey)
return _evaluation(participant, tasks, measured=measured)
client = AxonSystemClient(context.wallet, participant.hotkey, neuron.address, neuron.port)
live = run_live(
client,
system,
Expand Down
41 changes: 41 additions & 0 deletions microtensor/validator/live.py
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,47 @@ def task(self, name: str, request: Mapping[str, Any]) -> Mapping[str, Any]:
return dict(payload["trace"])


class AxonSystemClient:
def __init__(
self, wallet: Any, hotkey: str, address: str, port: int, timeout: float = 900.0
) -> None:
import bittensor as bt

self.dendrite = (getattr(bt, "Dendrite", None) or bt.dendrite)(wallet=wallet)
self.axon = bt.AxonInfo(
version=0,
ip=address,
port=port,
ip_type=6 if ":" in address else 4,
hotkey=hotkey,
coldkey="",
)
self.timeout = timeout

def task(self, name: str, request: Mapping[str, Any]) -> Mapping[str, Any]:
from microtensor.chain.synapse import system_task

synapse = system_task()(
system=name,
round_index=int(request["round_index"]),
task_ref=str(request["task_ref"]),
prompt=str(request.get("prompt", "")),
inputs=dict(request.get("inputs") or {}),
)
found = self.dendrite.query(
axons=[self.axon], synapse=synapse, timeout=self.timeout, deserialize=False
)
answer = found[0] if isinstance(found, list) else found
if getattr(answer, "failure", ""):
raise LiveError(str(answer.failure))
trace = getattr(answer, "trace", None)
if not trace:
status = getattr(getattr(answer, "dendrite", None), "status_code", None)
message = getattr(getattr(answer, "dendrite", None), "status_message", "") or ""
raise LiveError(f"the miner's axon returned no trace ({status} {message})".strip())
return dict(trace)


@dataclass(frozen=True, slots=True)
class LiveRun:
traces: tuple[Trace, ...]
Expand Down
Loading