Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
80 changes: 65 additions & 15 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@

---

Generate high-quality QA pairs and evaluation datasets from any source documents. YourBench transforms your PDFs, Word docs, and text files into structured benchmark datasets with configurable output formats. Appearing at COLM 2025. **100% free and open source.**
Generate QA pairs and evaluation datasets from source documents. YourBench transforms PDFs, Word documents, and text into structured benchmark datasets with configurable output formats. The library is open source; hosted model calls may incur provider charges.

## Features

Expand All @@ -34,28 +34,64 @@ Generate high-quality QA pairs and evaluation datasets from any source documents
- **Custom Output Schemas** – Define your own Pydantic models for question/answer format
- **Multi-Model Support** – Use different LLMs for different pipeline stages
- **HuggingFace Integration** – Push datasets directly to the Hub or save locally
- **Quality Filtering** – Citation scoring and deduplication built-in
- **Reviewable Outputs** – Source references, citation scores, and exact normalized-question deduplication

See the [redesign and migration guide](docs/MIGRATION.md) for new interfaces, breaking changes, and verification.

## Quick Start

Use [uv](https://docs.astral.sh/uv/getting-started/installation/) to run the packaged CLI directly:
Describe the evaluation you want and point YourBench at your documents:

```bash
pip install -e .
yourbench create "Test understanding of policy exceptions and difficult customer questions" \
--source ./documents --model YOUR_MODEL_ID --output ./benchmark
```

For an OpenAI-compatible endpoint:

```bash
uvx --from yourbench yourbench run example/default_example/config.yaml --debug
yourbench create "Build questions about policy exceptions" \
--source ./documents --model YOUR_MODEL_ID --output ./benchmark \
--base-url http://localhost:8000/v1 --api-key-env MODEL_API_KEY
```

The example config works out-of-the-box with env vars from `.env` (see `.env.template`).
Set `MODEL_API_KEY` in your environment, or omit `--api-key-env` for an unauthenticated local endpoint. Hugging Face providers use `HF_TOKEN` when available.

Install locally if you prefer:
YourBench interprets the brief, saves `plan.json` and `config.yaml`, then generates local datasets and JSONL under the output directory. Add `--plan-only` to inspect the interpretation first (this still makes a model call). Rerun a saved recipe with:

```bash
uv pip install yourbench
yourbench run example/default_example/config.yaml
yourbench run ./benchmark
yourbench inspect ./benchmark
```

The brief can specify domain, audience, language, difficulty, and question style. Exact counts, dollar budgets, conversational tasks, and executable evaluators are currently unsupported and should be reported by the planner. Generated answers still require evaluation of their quality; schema validation checks structure, not factual correctness.

Use `--max-tokens 4000 --concurrency 2` to bound each response and simultaneous requests, including planning. These are not total cost or question-count limits.

The natural-language frontend defaults to local output. YAML configurations remain supported for explicit stage/model settings. See [CLI reference](docs/CLI.md), [configuration changes](docs/CONFIGURATION.md#configuration-changes-in-the-natural-language-redesign), and [schema/export contracts](docs/CUSTOM_SCHEMAS.md#validation-and-export-contracts).

## Use from Python

```python
from yourbench import create, load_result

result = create(
"Test understanding of policy exceptions",
source="./documents", output="./benchmark", model="MODEL_ID",
base_url="http://localhost:8000/v1", max_tokens=4000, concurrency=2,
)
questions = result.load_dataset()

# Later, without model credentials or another inference call:
print(load_result("./benchmark").summary())
```

For an authenticated endpoint, set the key in the environment and pass `api_key_env="MODEL_API_KEY"`. See the [Python API guide](docs/PYTHON_API.md) for planning, rerunning, reading subsets, and notebook usage.

## Installation

Requires **Python 3.12+**.
Requires **Python 3.12**.

```bash
# With uv (recommended)
Expand All @@ -80,9 +116,13 @@ pip install -e .
```yaml
hf_configuration:
hf_dataset_name: my-benchmark
push_to_hub: false
upload_card: false
export_jsonl: true

model_list:
- model_name: openai/gpt-4o-mini
- model_name: MODEL_ID
base_url: https://api.openai.com/v1
api_key: $OPENAI_API_KEY

pipeline:
Expand Down Expand Up @@ -122,10 +162,12 @@ YourBench provides several CLI commands:

| Command | Description |
|---------|-------------|
| `yourbench run <config>` | Run the full pipeline |
| `yourbench create "brief" --source DIR --model MODEL --output DIR` | Interpret an objective and generate a local benchmark |
| `yourbench run <config-or-output>` | Run enabled stages from a saved recipe |
| `yourbench inspect <config-or-output> [--json]` | Read local status and subset sizes without inference |
| `yourbench validate <config>` | Check config without running |
| `yourbench estimate <config>` | Estimate token usage |
| `yourbench init` | Generate starter config interactively |
| `yourbench init` | Generate a local starter config |
| `yourbench stages` | List available pipeline stages |
| `yourbench version` | Show version |

Expand All @@ -135,6 +177,7 @@ See [CLI Reference](./docs/CLI.md) for full documentation.

| Guide | Description |
|-------|-------------|
| [Python API](./docs/PYTHON_API.md) | Create, run, and read local results from Python |
| [Configuration](./docs/CONFIGURATION.md) | Full config reference with all options |
| [Custom Schemas](./docs/CUSTOM_SCHEMAS.md) | Define your own output formats |
| [How It Works](./docs/PRINCIPLES.md) | Pipeline architecture and stages |
Expand All @@ -156,17 +199,24 @@ No installation needed:
The `example/` folder contains ready-to-use configurations:

- `default_example/` – Basic setup with sample documents
- `harry_potter_quizz/` – Generate quiz questions from books
- `harry_potter_quizz/` – Multiple-choice quiz with a replaceable sample corpus
- `custom_prompts_demo/` – Custom prompts for domain-specific questions
- `local_vllm_private_data/` – Use local models for private data
- `rich_pdf_extraction_with_gemini/` – LLM-based PDF extraction for charts/figures
- `rich_pdf_extraction_with_gemini/` – PDF extraction using a compatible vision model

Run any example:
Set the endpoint variables used by the examples (the sample documents are included):

```bash
export YOURBENCH_MODEL=MODEL_ID
export YOURBENCH_BASE_URL=http://localhost:8000/v1
# For an authenticated endpoint, set YOURBENCH_API_KEY in your environment.
# For an unauthenticated local endpoint, use a nonempty placeholder:
export YOURBENCH_API_KEY=not-needed
yourbench run example/default_example/config.yaml
```

See the [examples guide](example/README.md) for the six recipes and their required model capabilities.

## API Keys

Set in environment or `.env` file:
Expand Down
Loading
Loading