A universe where legal agents learn, act, and are evaluated.
Quick Start · Benchmarks · Outputs · Architecture
LexVerse is an evolving runtime for evaluating legal language models and agents. It brings legal benchmarks into a shared workflow for task preparation, execution, and result management while preserving their original task formats, interaction harnesses, and evaluators. We are continuing to integrate more benchmarks and develop environment capabilities for richer agent interactions.
| Benchmark | Coverage | Interaction | Verification |
|---|---|---|---|
| LexEval | 23 tasks | Single response | Upstream task evaluator |
| LawBench | 20 tasks | Single response | Upstream task evaluator |
| J1Bench | CI, CR, KQ, LC, CD, DD | Official multi-role harness | Upstream scenario evaluator |
Use Python 3.10 or newer. Install LexVerse with the dependencies for all three benchmarks:
python -m pip install -e ".[lexeval,lawbench,j1bench,openai]"
python -m lexverse --helpCompatibility note One environment can run all these benchmarks. LexEval and LawBench use different Rouge packages, so some Rouge-based scores may differ slightly from upstream results in a combined installation.
Create local configuration files from the tracked *.example.yaml templates.
Configure at least one OpenAI-compatible connection in
configs/secrets.local.yaml.
The benchmark YAML selects the connection with model.profile and the actual
model with model.name. J1Bench also has evaluation.model, because its
official evaluator calls a scoring model. Model names belong in benchmark
configuration, not in the secrets file.
prepare resolves the selected upstream cases and freezes them into a
manifest. It does not call the model. run performs generation, official
task-level evaluation, and aggregation.
python -m lexverse prepare \
--config configs/benchmarks/lexeval.local.yaml \
--output .lexverse/prepared/lexeval-all.json
LEXEVAL_RUN="runs/lexeval-all-$(date +%Y%m%d-%H%M%S)"
python -m lexverse run \
--prepared .lexverse/prepared/lexeval-all.json \
--output-root "$LEXEVAL_RUN"python -m lexverse prepare \
--config configs/benchmarks/lawbench.local.yaml \
--output .lexverse/prepared/lawbench-all.json
LAWBENCH_RUN="runs/lawbench-all-$(date +%Y%m%d-%H%M%S)"
python -m lexverse run \
--prepared .lexverse/prepared/lawbench-all.json \
--output-root "$LAWBENCH_RUN"If a run is interrupted after model generation, set execution.resume: true
in configs/benchmarks/lawbench.local.yaml, run prepare again, and reuse the
same --output-root. Completed trials will be reused instead of calling the
model again.
J1Bench uses gated data. Accept the terms on the J1-Eval dataset page and authenticate once:
hf auth loginThe first prepare downloads the selected J1-Eval_<SCENARIO>.jsonl files to
.lexverse/datasets/j1bench/. Later runs reuse this cache and do not modify the
downloaded files.
python -m lexverse prepare \
--config configs/benchmarks/j1bench.local.yaml \
--output .lexverse/prepared/j1bench-all.json
J1BENCH_RUN="runs/j1bench-all-$(date +%Y%m%d-%H%M%S)"
python -m lexverse run \
--prepared .lexverse/prepared/j1bench-all.json \
--output-root "$J1BENCH_RUN"J1Bench runs the official multi-role conversation for every case and invokes
the official evaluator once per scenario. It is slower and more expensive than
the two single-response benchmarks. A successful run has already completed
evaluation; there is no separate eval command.
Each --output-root contains manifest.json, summary.json, trial records,
and official evaluation artifacts. LexEval and LawBench report task-level
scores in evaluation_result.csv; J1Bench keeps per-case intermediate results
and a scenario-level final result under trials/<scenario>/verifier/.
collect rebuilds summary.json from existing Trial records and the existing
official evaluation_result.csv:
python -m lexverse collect --run-dir "$LEXEVAL_RUN"For a cheaper smoke test, change only execution.limit in a local benchmark
YAML. Run prepare again after every YAML change; run rejects configuration
drift by design.
LexVerse connects benchmark preparation, execution, evaluation, and result management in one workflow.
Open the interactive architecture map →
- Do not edit files under
.lexverse/datasets/j1bench/. - Re-run
preparewhenever a local YAML changes. - Upstream repositories are pinned in
lexverse/runtime/upstream.py; patches make them relocatable without replacing evaluator logic. - Benchmark and dataset licenses remain governed by their original sources.
LexVerse is released under the Apache License 2.0.
For issues and feature requests, use GitHub Issues. You can also email xieh@tsinghua.edu.cn.