Leanest does not predict which tests will fail. It determines which tests are safe enough not to run.
Local-first test selection using semantic judgments (classifier.dev by default, or Jev/Laya). Leanest sits in front of your existing test runner and runs only the tests that matter for a given code change. Everything else it skips, on purpose, out loud.
Works with npm, pnpm, yarn, or Bun.
npm install -D leanest
# or: bun add -D leanestWorks with no setup: leanest defaults to classifier.dev, a free, no-auth judge. Switch to Jev if you want it by exporting TYPESAFE_API_KEY and setting LEANEST_PROVIDER=jev (see Judge provider).
npx leanest playwright # select + actually run the affected e2e tests
npx leanest playwright --base origin/main # diff against a specific base
npx leanest select playwright # just show the selection, don't run anything
npx leanest inspect playwright # rank every test by relevance, for debuggingRepository
|
+-- Git change (base...head)
+-- Discovered test files
|
v
Leanest
|
+-- Change resolver (git diff)
+-- Test discovery (respects the framework's own config, e.g. playwright.config.ts testDir)
+-- Whole-suite rules (runner setup changed: run all; only Markdown changed: skip all)
+-- Context builder (packages the diff + each test's source for the judge)
+-- Judge (classifier.dev by default: one semantic judgment per test)
+-- Selection policy (RUN / SKIP, fail-open on low confidence)
|
v
Selected test files
|
v
Your existing runner (playwright / vitest), unmodified
For each discovered test, Leanest asks:
Could the current code change affect behavior verified by this test?
Tests that are confidently irrelevant get skipped. Everything else runs through your existing runner exactly as it would outside Leanest: same reporter, same exit code, same flags.
- Fail open: uncertainty means RUN. A missing API key, an API timeout, or a malformed response always falls back to running the full suite, loudly (
⚠ Judge unavailable (...), running the full suite.). Finding no tests at all is an error (exit 1), not a silent pass. - Deterministic overrides: whatever the judge says, these run. A test whose own file changed always runs, as does one that statically imports a changed file, or that navigates a route a changed file's own path names (e.g.
page.goto("/admin/users")against a changedroutes/admin/users.tsx) -- a heuristic that catches e2e route coupling no import graph can see, since a browser test never imports the page it drives. - Whole-suite rules: before any per-test decision, a change to the runner's own setup (
playwright.config.*,vitest.config.*,vite.config.*,package.json, a lockfile, or anything in.github/workflows/) runs every test, and a change that only touches Markdown files skips every test, except one a deterministic override forces (say, a test that imports the changed.mdfile). Neither asks the judge, so every shard of a matrix gets the same answer. - Leanest doesn't run tests itself: it selects file paths and hands them to your actual runner (
playwright test <paths>,vitest run <paths>). It leaves reporters, retries, sharding, and CI-required-check behavior alone. Anything after--goes straight to the runner:leanest playwright -- --shard=1/3. - Static checks are out of scope on purpose: lint/format/typecheck are already fast at full scope, and semantic per-rule selection would add latency for no real payoff. Leanest spends its judge calls only on suites that are expensive to run in full: e2e today, more later.
| Framework | Status | Command |
|---|---|---|
| Playwright | First-class | npx leanest playwright |
| Vitest | First-class | npx leanest vitest |
| Jest | Planned | — |
| Pytest | Planned | — |
npx leanest playwrightnpx leanest playwright --changednpx leanest playwright --base mainnpx leanest playwright --dir ~/projects/my-appnpx leanest playwright --jsonnpx leanest select playwright --base origin/mainnpx leanest playwright --fullRuns every test for real (it skips nothing), in two batches: the selected tests, then the ones Leanest would have skipped. If the skipped batch fails, it prints Shadow mode: MISS, so after a few weeks you know how often selection alone would have let a failure through:
npx leanest playwright --shadownpx leanest inspect playwrightRUN tests/e2e/admin-users-export.pw.ts
RUN tests/e2e/downgrade.pw.ts
SKIP tests/e2e/qr-generator.pw.ts
SKIP tests/e2e/avatar.pw.ts
...
Leanest loads .env for local convenience. The API key is never persisted or logged.
TYPESAFE_API_KEY=...Framework choice, base ref, and target directory are all CLI flags (--base, --dir), so there's nothing else to set up per project.
Leanest's selection judgment is pluggable. Pick a provider with LEANEST_PROVIDER:
| Provider | How | API key needed |
|---|---|---|
classifier-dev (default) |
classifier.dev, a free zero-shot classifier | none |
jev |
TypeSafe's Jev, over HTTP | TYPESAFE_API_KEY |
laya |
Laya, self-hosted, runs in-process via ONNX Runtime (bun add @receptron/laya) |
none |
LEANEST_PROVIDER=jev npx leanest playwrightWhere your code goes: classifier-dev and jev send the diff and the source of each candidate test to an outside service (classifier.dev or TypeSafe). On a private repo, check that's acceptable first, or use laya, which runs in-process.
permissions:
contents: read
pull-requests: write # for the report comment
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: baronunread/leanest@v0.2.7
with:
framework: playwrightThis installs the leanest version matching the Action's ref with the runner's Node (it doesn't touch your Bun), and replaces your existing "run e2e tests" step: same reporter output, same exit code, just fewer tests executed. No secret required — the default classifier-dev provider needs no API key, which also means forked-repo PRs can use it without access to your repo's secrets. Pass provider: jev and typesafe-api-key: ${{ secrets.TYPESAFE_API_KEY }} to use Jev instead.
On pull requests it diffs against the PR's base branch; on push, against the previous commit. Override with base:. Pass runner flags with args:, for example args: --shard=${{ matrix.shard }}/3. Each run writes a job summary listing every test file, whether it ran, and why. On pull requests it also posts that report as a PR comment and edits the same comment on later pushes:
### leanest: 5 of 8 playwright test files selected
2 touch the change directly, the judge wasn't sure enough to skip 2 and it thinks 1 is affected.
| Test | Decision | Reason |
| --- | --- | --- |
| `tests/e2e/checkout.pw.ts` | RUN | test file changed |
| `tests/e2e/cart.pw.ts` | RUN | imports a changed file |
| `tests/e2e/login.pw.ts` | RUN | judge p=0.34 c=0.18 |
| … | | |
▸ 3 skipped (collapsed, each with the judge's p and c)If the judge is down, the comment opens with a warning that gives its error and says the full suite ran.
Turn the comment off with comment: false. In a sharded matrix only shard 1 posts it. With other matrix axes (browsers, OSes), set comment: false on all but one job, or they'll overwrite each other. Without pull-requests: write, and on fork PRs (which get a read-only token), posting logs a warning and the tests' result stands.
npm install -g leanest
leanest playwright --base origin/mainWorks anywhere you can run a shell command and set an env var: GitLab CI, CircleCI, Buildkite.
bun install
bun run check # lint + format check + typecheck + test
bun test # just the test suiteSee LEANEST_SPEC.md for the design rationale behind the selection policy.
MIT. See LICENSE.