Simple Topic Modeling finds the themes in a collection of documents. It runs fully in your browser.
The app is a marimo notebook. The export runs it with Pyodide, which is Python compiled to WebAssembly. A static web server is enough to host it. No Python server is needed.
Open the app at https://maehr.github.io/simple-topic-modeling/. Nothing to install.
Version 1 was a Streamlit app that ran through stlite. Version 2 shares no code with it. The app is
now a marimo notebook, it draws every chart with Altair, and it ships one wheel that hatchling
builds.
Version 1 receives no further work. Its code stays at the tag
v1.0.0 and the branch
legacy/streamlit.
The app opens with two demos, so you can see a real result before you load your own documents. Newspaper Corpus demo fits the newspaper demo. Single Book demo fits the book demo.
The newspaper demo holds 295 articles from the Journal de Genève and the Gazette de Lausanne of 1914. Each article carries its publication date and the section heading that the newspaper printed above it. Six sections give six clear themes: the military chronicle, sport, finance, the weather, book reviews, and the courts. A machine read the articles from a scan, so some words carry errors. A real archive looks like this.
The Digital Humanities Laboratory of the EPFL digitised the historical archive of
Le Temps. It published the year 1914 under CC BY 4.0, for the
2015 Swiss Open Cultural Data Hackathon. The articles
are anonymous newspaper text from 1914, so they left copyright in 1985. NOTICE holds
the full statement.
scripts/build_demo_corpus.py rebuilds the corpus from the archive. It checks the archive against
a known hash and records each source URL and each SHA-256 in
scripts/demo_corpus.provenance.json.
The book demo is Mary Wollstonecraft, A Vindication of the Rights of Woman (1792), in English. The app splits the book into 784 paragraphs and keeps their order, in Ordered text mode. The results show where each topic occurs, from the dedication to the last chapter.
The text comes from Project Gutenberg #3420. The author
died in 1797, so the work is in the public domain. scripts/build_demo_text.py rebuilds the file.
It removes the Project Gutenberg header, footer, and licence, the table of contents, and the bare
chapter and section numbers. It checks the source against a known hash and records each SHA-256 in
scripts/demo_text.provenance.json.
Every run uses a fixed random seed. The same settings on the same corpus give the same topics.
- Download
config.jsonin Step 4. It records the app version, the language, the stop words, every model parameter, including the seed, and the split of one text. - Download the demo file in Step 1, or keep your own corpus.
- Send both files to your reader.
- The reader loads
config.jsonin Step 2. The app restores each setting.
Your documents stay on your computer. The app sends no text to an analysis server. The app calls no model API. The app builds every download in your browser.
- Add a corpus. Drag files in, paste text, or load a demo. Select Newspaper Corpus demo or Single Book demo to fit a model on a demo in one step.
- Choose a language. Edit the stop words when you need to.
- Choose a model and its parameters.
- Select Run model. The app fits the model in your browser.
- Explore the topics in the five tabs.
- Rename a topic with the Rename button in its header on the Topics tab. The app does not refit the model.
- Download the results.
Step 2 shows the model and the number of topics. The other parameters wait in a closed Advanced settings panel.
The app fits a model only when you select Run model, Newspaper Corpus demo, Single Book demo, or Rerun. A changed setting does not start a new fit. The app keeps the last result until a new fit succeeds.
The app reads two kinds of file.
Use a structured file when one row holds one document plus its metadata:
| Format | Extension |
|---|---|
| Comma-separated values | .csv |
| Tab-separated values | .tsv, .tab |
| JSON | .json |
| JSON Lines | .jsonl, .ndjson |
Use a text file when the file itself is one document. The app reads any UTF-8 text file. It removes the tags from HTML and XML. It removes the common syntax from Markdown.
The app reads a text-based PDF file. It extracts the text in your browser with pypdf. One PDF opens as an ordered text. A PDF rarely holds blank lines, so each page usually becomes one segment.
- A scanned PDF holds images, not text. The app does no OCR. Convert the scan to a searchable PDF or to a text file first.
- The app skips a page without text, and it names the skipped pages.
- The app does not open a PDF that needs a password.
The app rejects any other binary file, such as a DOCX file, an image, or a ZIP archive.
A book, a thesis, or a transcript is one long text. The app can show where each topic occurs in it.
- Load one text file or one PDF file, or paste one text.
- Set Analyse as to Ordered text.
- Keep Split on blank lines. The app then makes one segment from each paragraph.
- Select Run model.
Each segment is one document for the model. Each segment keeps its position in the text:
- Topics shows where the selected topic occurs, and the representative passages.
- Documents shows the topic share through the whole text, as a heatmap.
documents_topics.csvholdsparent_document_id,segment_index, andsegment_number, so you can rebuild the order of the text.
Independent segments is the default for a text file. Ordered text is the default for a PDF.
Independent segments splits the text the same way, but it treats the segments as
unrelated documents. A line under the input summary says what the model sees, for example
1 file → 327 paragraphs → analysed as an ordered text.
The app ships a stop-word list for each language below. The language changes the stop words only. It adds no stemming and no grammar analysis.
| Language | Code |
|---|---|
| English | en |
| German | de |
| French | fr |
| Italian | it |
| Spanish | es |
A corpus with more than one language can turn the base list off and add its own words.
| Model | When to use it |
|---|---|
| NMF, with TF-IDF | Start here. It is fast, and it suits a short or medium document. |
| LDA, with counts | Use it for a probabilistic topic mixture, or for a long document. |
The explorer holds five tabs:
- Overview shows the topic cards, the topic map, the prevalence bars, and the similarity heatmap.
- Topics shows the top terms, the word cloud, and the representative documents of one topic. An ordered text also shows where the topic occurs.
- Documents shows the document map, a text search, a topic filter, and a score filter. A long document also shows the topic share through the text.
- Metadata shows the topic mix per group and the topic share over time. It needs a group column or a date column.
- Diagnostics shows the descriptive numbers and the friendly notices. These numbers are not a quality score.
You can download six files:
| File | Content |
|---|---|
documents_topics.csv |
One row per document, with its topic scores |
topics.csv |
One row per topic, with its name and prevalence |
topic_terms.csv |
One row per topic and term, in long format |
topic_similarity.csv |
The cosine similarity of each topic pair |
config.json |
The settings that produced the result |
project.zip |
The five files above, plus a README.txt |
Install uv first. Then install the dependencies:
uv syncuv run marimo edit app.pyThis command runs local CPython. It does not test the browser build.
Build the package wheel first. The browser installs the app from this file.
uv build --wheel -o public/wheelsCheck the browser compatibility:
uv run marimo check app.py --select MWExport the app:
uv run marimo export html-wasm app.py -o dist --mode runServe the output over HTTP:
python -m http.server --directory dist 8000Then open http://localhost:8000.
Caution: the app cannot start from a file:// address. The browser blocks the worker that
Pyodide needs.
.
├── app.py the user interface, the state, and the downloads
├── simple_topic_modeling/ the logic, which the tests cover
│ ├── config.py the settings models
│ ├── errors.py one friendly message per failure
│ ├── exports.py the CSV, JSON, and ZIP files
│ ├── io.py the import, the PDF text, the split, and the corpus statistics
│ ├── metrics.py the similarity, the diversity, and the notices
│ ├── modeling.py the vectorizer and the model fit
│ ├── plots.py the chart data and the chart specifications
│ ├── preprocess.py the text cleaning
│ ├── result.py the result object
│ ├── stopwords.py the packaged stop-word lists
│ └── data/ the stop-word lists and the two demos
├── public/wheels/ the built wheel, which the export copies
├── scripts/ the scripts that build the two demos
└── tests/
app.py holds the user interface only. simple_topic_modeling/ holds the logic. The tests cover the
logic, not the notebook.
Run the full gate before each commit:
uv run ruff check . && uv run ruff format --check .
uv run ty check
uv run pytest --doctest-modules --cov --cov-fail-under=100Check the notebook after you edit it. marimo check does not find a name that two cells both define.
uv run python -c "
import importlib.util, sys
spec = importlib.util.spec_from_file_location('notebook_app', 'app.py')
module = importlib.util.module_from_spec(spec); sys.modules['notebook_app'] = module
spec.loader.exec_module(module); module.app.run(); print('cells ok')
"Build the export and open it in a browser before you call the work done. The notebook and the export are two targets. A change can pass one and fail the other.
Write each commit message as a Conventional Commit. Then update the changelog:
git-cliff --tag v2.0.0-alpha -o CHANGELOG.mdAGENTS.md holds the rules for this repository. Read it before you change the code.
| File | Content |
|---|---|
AGENTS.md |
The rules for an agent or a contributor |
SPECS.md |
The product specification |
CHANGELOG.md |
The record of every notable change |
CONTRIBUTING.md |
How to set up, to check, and to open a pull request |
CODE_OF_CONDUCT.md |
Contributor Covenant 3.0 |
SECURITY.md |
How to report a security issue |
NOTICE |
The origin and the licence of the demo corpus and the stop-word lists |
LICENSE |
The full licence text |
Moritz Mähr wrote and maintains Simple Topic Modeling.
- Report a problem or ask for a feature.
- Read the contribution guidelines before you open a pull request.
- Read the code of conduct.
- Read how to report a security issue.
AGPL-3.0-only. See LICENSE.
The stop-word lists come from spaCy under the MIT licence. The demo corpus
comes from the historical archive of Le Temps, under CC BY 4.0. See NOTICE.