Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 1 addition & 8 deletions .github/workflows/docs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,11 +9,4 @@ on:

jobs:
Docs:
#Disabled until major code changes completed.
# uses: tskit-dev/.github/.github/workflows/docs.yml@v15

# Placeholder
runs-on: ubuntu-latest
steps:
- name: Placeholder
run: echo "Docs broken, fix when code is ready!"
uses: tskit-dev/.github/.github/workflows/docs.yml@v15
7 changes: 6 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,9 @@ valgrind --leak-check=full --error-exitcode=1 ./build/tests
Tests are in `lib/tests/tests.c` using the CUnit framework. The build uses
`-Wall -Wextra -Werror -Wpedantic` and other strict warnings.

Ensure that all new C code is covered by tests in the C test suite by running
tests with coverage.

## Architecture

**tsinfer** infers tree sequences from genetic variation data stored in VCZ (Variant Call Zarr) format.
Expand Down Expand Up @@ -77,7 +80,9 @@ The public API is in `tsinfer/__init__.py`, exposing three main functions from `
Source in `lib/`. Three main classes exposed to Python:
- `AncestorBuilder` — builds inferred ancestors from genotype data
- `AncestorMatcher` — Li & Stephens HMM matching algorithm
- `TreeSequenceBuilder` — constructs tree sequences incrementally

When changes are made to the C library, ensure that the ``_tskit`` module is rebuilt
before running Python tests.

Vendored dependencies in `lib/subprojects/`: tskit C library and kastore.

Expand Down
3 changes: 2 additions & 1 deletion docs/_config.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,12 +32,13 @@ sphinx:
- sphinx.ext.viewcode
- sphinx.ext.intersphinx
- sphinx_issues
- sphinxarg.ext
- sphinx_click
- IPython.sphinxext.ipython_console_highlighting

config:
html_theme: sphinx_book_theme
html_theme_options:
announcement: "⚠ This documentation is under active development. The API is not yet stable and many elements are out of date."
pygments_dark_style: monokai
navigation_with_keys: false
logo:
Expand Down
2 changes: 0 additions & 2 deletions docs/_static/example_ancestral_state.fa

This file was deleted.

27 changes: 27 additions & 0 deletions docs/_static/example_config.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Example tsinfer config for a small dataset.
#
# Run:
# vcf2zarr convert example_data.vcf.gz example_data.vcz
# tsinfer run example_config.toml -v

[[source]]
name = "example"
path = "example_data.vcz"

[ancestral_state]
path = "example_data.vcz"
field = "variant_AA"

[[ancestors]]
name = "ancestors"
path = "example_ancestors.vcz"
sources = ["example"]

[match]
output = "example_output.trees"

[match.sources.ancestors]
node_flags = 0
create_individuals = false

[match.sources.example]
Binary file added docs/_static/example_data.vcf.gz
Binary file not shown.
6 changes: 2 additions & 4 deletions docs/_toc.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,18 +10,16 @@ parts:
- file: installation
- caption: Usage
chapters:
- file: quickstart
- file: config
- file: usage
- caption: Inference
chapters:
- file: inference
- file: large_scale
- caption: Interfaces
chapters:
- file: api
- file: cli
- caption: File Formats
chapters:
- file: file_formats
- caption: Miscellaneous
chapters:
- file: development
Expand Down
89 changes: 0 additions & 89 deletions docs/api.rst

This file was deleted.

49 changes: 9 additions & 40 deletions docs/cli.rst
Original file line number Diff line number Diff line change
@@ -1,49 +1,18 @@
.. _sec_cli:
.. _sec_cli_reference:

======================
Command line interface
======================

.. warning::

The command line interface only supports the deprecated SampleData format
used in tsinfer<0.4.0.

The command line interface in ``tsinfer`` is intended to provide a convenient
interface to the high-level :ref:`API functionality <sec_api>`. There are two
equivalent ways to invoke this program:

.. code-block:: bash

$ tsinfer

or
The ``tsinfer`` command line interface runs the inference pipeline using a
TOML configuration file. See the :ref:`quickstart <sec_quickstart>` for an
introduction and the :ref:`config reference <sec_config_reference>` for all
available options.

.. code-block:: bash

$ python3 -m tsinfer

The first form is more intuitive and works well most of the time. The second
form is useful when multiple versions of Python are installed or if the
:command:`tsinfer` executable is not installed on your path.

The :command:`tsinfer` program has five subcommands: :command:`list` prints a
summary of the data held in one of tsinfer's :ref:`file formats <sec_file_formats>`;
:command:`infer` runs the complete :ref:`inference process <sec_inference>` for a given
input SampleData file; and
:command:`generate-ancestors`, :command:`match-ancestors` and
:command:`match-samples` run the three parts of this inference
process as separate steps. Running the inference as separate steps like this
is recommended for large inferences as it allows for greater control over
the inference process.

++++++++++++++++
Argument details
++++++++++++++++

.. argparse::
:module: tsinfer
:func: get_cli_parser
:prog: tsinfer
:nodefault:
$ tsinfer run config.toml --threads 4 -v

.. click:: tsinfer.cli:main
:prog: tsinfer
:nested: full
116 changes: 116 additions & 0 deletions docs/config.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
(sec_config_reference)=

# Configuration reference

_Tsinfer_ is configured via a TOML file passed to the CLI. Paths in the config
are resolved relative to the config file's directory.

A complete annotated example is in
[example_config.toml](https://github.com/tskit-dev/tsinfer/blob/main/example_config.toml).


## `[[source]]`

Each `[[source]]` block defines a named view over a VCZ store. The same store
can appear multiple times with different filters.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `name` | string | (required) | Unique name for this source |
| `path` | string | (required) | Path to VCZ store |
| `include` | string | — | bcftools include expression (e.g. `"TYPE='snp'"`) |
| `exclude` | string | — | bcftools exclude expression |
| `samples` | string | — | Sample filter (comma-separated; prefix `^` to exclude) |
| `regions` | string | — | Genomic region, half-open (e.g. `"chr20:1000-50000"`) |
| `targets` | string | — | Exact target positions |
| `sample_time` | various | — | Per-sample times: constant, field name, or `{path, field}` dict |


## `[ancestral_state]`

Specifies where to read the ancestral allele for each variant position.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `path` | string | (required) | Path to VCZ containing ancestral alleles |
| `field` | string | (required) | Array name in the store (e.g. `"variant_AA"`) |


## `[[ancestors]]`

Controls the ancestor-generation step (`infer-ancestors`). At least one
`[[ancestors]]` block is required unless `[match]` specifies a `reference_ts`.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `name` | string | (required) | Unique ancestor set name |
| `path` | string | (required) | Output VCZ path |
| `sources` | list[str] | (required) | Source names to build ancestors from |
| `max_gap_length` | int | 500,000 | Split intervals at gaps wider than this (bp) |
| `samples_chunk_size` | int | 100 | Zarr chunk size (ancestor dimension) |
| `variants_chunk_size` | int | 50,000 | Zarr chunk size (site dimension) |
| `compressor` | string | `"zstd"` | Blosc compressor name |
| `compression_level` | int | 7 | Compression level (0–9) |
| `genotype_encoding` | string | `"eight_bit"` | `"one_bit"` uses ~8x less memory (biallelic only) |


## `[match]`

Controls the HMM matching step.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `output` | string | (required) | Output `.trees` file path |
| `path_compression` | bool | `true` | Enable Viterbi path compression |
| `reference_ts` | string | — | Reference tree sequence (skip ancestor generation) |
| `workdir` | string | — | Checkpoint directory (enables resume) |
| `keep_intermediates` | bool | `false` | Keep per-group checkpoint files |


## `[match.sources.<name>]`

Per-source parameters. Every source that should appear in the output tree
sequence needs an entry here.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `node_flags` | int | 1 | tskit node flags (`1` = `NODE_IS_SAMPLE`, `0` for ancestors) |
| `create_individuals` | bool | `true` | Group sample nodes into tskit individuals |


## `[post_process]`

Optional cleanup applied after matching.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `split_ultimate` | bool | `true` | Split virtual root into per-tree roots |
| `erase_flanks` | bool | `true` | Erase ancestry outside informative sites |


## `[augment_sites]`

Place non-inference sites via parsimony.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `sources` | list[str] | (required) | Source names for parsimony placement |


## `[individual_metadata]`

Map VCZ sample-dimensioned arrays into tskit individual metadata.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `population` | string | — | VCZ array whose unique values become tskit populations |

### `[individual_metadata.fields]`

Each key becomes a tskit metadata field; the value names the VCZ array.

```toml
[individual_metadata.fields]
name = "sample_id"
sex = "sample_sex"
```
Loading
Loading