Skip to content

Repository files navigation

wikitermbase

Table of Contents

Overview

Wiki Term Base is a tool designed to standardise terminology used on Arabic Wikipedia and accelerate vocabulary translation.

ℹ For functional documentation, please check the dedicated Wikipedia page مسرد الويكي (in Arabic).

🌐 The website is available at: https://wikitermbase.toolforge.org

It is hosted on Toolforge, as a Python ASGI application built with the FastAPI framework (served by gunicorn with uvicorn workers via the Toolforge Build Service), using a MariaDB relational database.

The website's frontend is built with React framework.

The Wikipedia gadget is built with OOUI (loaded on demand) and can be enabled in Arabic Wikipedia's user preferences.

Wiki Gadget

The Wikipedia gadget can be activated in user preferences -> "مسرد الويكي".

The deployed version in Arabic Wikipedia:

Files in gadget/:

  • Gadget-WikiTerm.js and Gadget-WikiTerm.css are the gadget, copied verbatim to the MediaWiki: pages above.
  • SearchTerm.js is the user script variant used for development: the same body as the gadget wrapped in mw.loader.using( [ 'mediawiki.util' ] ) (only the first and last lines differ). Regenerate it after editing the gadget so both stay in sync.

Design constraints (the gadget is meant to be enabled by default, see the default-gadget criteria):

  • The only page-load dependency is mediawiki.util. OOUI (about 90 KB gzipped) and the dialog are loaded on the first click via mw.loader.using(); no request reaches the WikiTermBase API until the user submits a search.
  • Entry points: an icon button in the header on Vector 2022 (and its sticky header), on Minerva and in the Content Translation tool (Special:ContentTranslation has its own skin); an item in the page-actions menu ("المزيد") on Vector legacy, MonoBook, Timeless and any other skin, via mw.util.addPortletLink(). Users of those skins who prefer the top personal toolbar can set window.wikiTermConfig = { placement: 'personal' }; in their common.js.
  • Only ES2015 syntax (MediaWiki's Grade A baseline is ES2019, and requiresES6 cannot be combined with default). No console.* calls.
  • Results are fetched 30 groups at a time (limit / offset on /api/v1/search/aggregated); "show more" requests the next window. Broad terms have thousands of groups and multi-megabyte full responses. A new request aborts the one in flight.

Recommended gadget definition (registered users only, testable with ?withgadget=WikiTerm before enabling it by default):

* WikiTerm [default |rights=minoredit |supportsUrlLoad |dependencies=mediawiki.util] |WikiTerm.js |WikiTerm.css

Testing the gadget

Tooling lives in gadget/package.json (ESLint with the Wikimedia config, Playwright for a browser matrix):

cd gadget && npm install
npm run lint                              # eslint-config-wikimedia: client/es6 + mediawiki + jquery
npx playwright install firefox webkit     # once; Chrome uses the installed Google Chrome
npm run matrix                            # Chrome/Firefox/WebKit × Vector 2022 (light+night)/Vector 2010/MonoBook/Timeless/Minerva/Content Translation
npm run summary                           # Markdown table from tests/out/matrix_results.json (screenshots in tests/out/shots/)
BROWSERS=chrome SKINS=vector npm run matrix   # subset

The matrix opens a real ar.wikipedia article (logged out), injects the working-tree gadget, and drives it end to end: entry point → dialog (lazy OOUI load, bytes and time recorded) → search → expand → citation copy → "show more" → close, failing on any uncaught JavaScript error. npm run check is network-free: it verifies SearchTerm.js is in sync with the gadget, enforces a gzipped size budget (10 KB JS, 3 KB CSS) and rejects console.* calls.

The same three commands run in GitHub Actions (gadget.yml) on every pull request touching gadget/, on pushes to main, and weekly. The run's job summary shows the results table and the screenshots + JSON are attached as a downloadable artifact, so the numbers can be checked and re-run by anyone from the Actions tab.

To try the working-tree version on-wiki without deploying anything, disable the WikiTerm gadget in your preferences, open any page and paste in the browser console (replace main with your branch):

const base = 'https://raw.githubusercontent.com/forzagreen/wikitermbase/main/gadget/';
fetch(base + 'Gadget-WikiTerm.css').then(r => r.text()).then(css => mw.util.addCSS(css));
fetch(base + 'Gadget-WikiTerm.js').then(r => r.text()).then(js => $.globalEval(js));

Or install it as a user script: copy gadget/SearchTerm.js to User:You/SearchTerm.js, the CSS to User:You/SearchTerm.css, and load both from your common.js.

Reproducible footprint checks anyone can run in the browser console on ar.wikipedia:

  • mw.loader.getState('oojs-ui-core') — registered means OOUI is not loaded; the old definition makes it ready on every page, the new one only after the first click.
  • mw.loader.inspect() — MediaWiki's own per-module size report; look for ext.gadget.WikiTerm and the oojs-ui-* rows.
  • DevTools → Network, filter load.php: with the new gadget nothing is fetched from wikitermbase.toolforge.org until a search is submitted.

Once the gadget definition carries supportsUrlLoad, external tools can A/B the page-load impact on the same URL with and without ?withgadget=WikiTerm (e.g. Lighthouse in Chrome DevTools, PageSpeed Insights, WebPageTest). Note ?withgadget= only works for users the gadget is registered for, so a rights= restriction hides it from logged-out tools.

Local Setup

Please note that the database content is managed in the project arabterm.

Clone the arabterm repository, and start the MariaDB database in a Docker container:

make init
make init_mariadb  # start or create container
make delete_mariadb  # delete database if exists
make migrate_to_mariadb  # migrate the SQLite content to MariaDB

Then from wikitermbase repository, install python dependencies (requires uv):

make init

Create a file at ./var/local.cnf with (adapt values):

[client]
user = MyUserName
password = MyTestPassword

Start the application:

make run

You can then open the web application at http://127.0.0.1:5001/

Backend

Python version: 3.13

API

Interactive OpenAPI docs (Swagger UI) are available at /docs — and at /redoc for the ReDoc rendering. These are auto-generated from the FastAPI route signatures and let you try every endpoint from the browser.

  • Aggregated search (results are groupped by the arabic term):
GET /api/v1/search/aggregated?q=magnetoscope
GET /api/v1/search/aggregated?q=اشتقاق

As a result, we get a JSON. An example can found at gadget/response.json

  • Raw search (without groupping):
GET /api/v1/search?q=magnetoscope
GET /api/v1/search?q=اشتقاق

API on Toolforge (Build Service)

ASGI applications cannot run on Toolforge's legacy python3.13 uWSGI webservice — they require the Build Service backend, which uses Cloud Native Buildpacks to build a container image directly from the public GitHub repo and runs it according to the Procfile. Frontend assets (backend/frontend/dist/) are committed to git so the Python buildpack alone is sufficient — no Node.js step in the build pipeline.

Refs:

Initial Setup

DB credentials don't need to be configured: Toolforge auto-injects TOOL_REPLICA_USER and TOOL_REPLICA_PASSWORD into Build Service containers (same as for the legacy uWSGI webservice). The app reads them directly from os.environ.

ssh toolforge
become wikitermbase

# Stop the legacy webservice if it was previously running on python3.13
toolforge webservice --backend=kubernetes python3.13 stop || true

# Build the image from the public GitHub repo
toolforge build start https://github.com/forzagreen/wikitermbase
toolforge build show   # wait until status is ok(Succeeded)

# Start the Build Service webservice
toolforge webservice buildservice start --mount=none

Test: https://wikitermbase.toolforge.org/api/v1/stats. Logs: toolforge webservice buildservice logs -f.

Updating the Codebase

Code deploys are automated. On push to main, the deploy-code job in .github/workflows/ci.yml SSHs into the bastion and runs toolforge build start + toolforge webservice buildservice restart. Markdown-only and data-only changes skip the rebuild. Manual re-deploy: Actions tab → "CI" → "Run workflow" on main.

Include any frontend rebuild in the commit (make build_frontend && git add backend/frontend/dist && git commit). The Python buildpack auto-detects uv.lock and installs deps with uv sync, so committing changes to pyproject.toml + uv.lock is all that's needed when adding dependencies.

Verify the gadget on Arabic Wikipedia still works after each deploy.

Manual fallback (if GitHub Actions is down):

ssh toolforge && become wikitermbase
toolforge build start https://github.com/forzagreen/wikitermbase
toolforge build show   # wait until status is ok(Succeeded)
toolforge webservice buildservice restart

Database: MariaDB

Data lives in forzagreen/arabterm — that's the source of truth and where dictionary edits happen. When a PR touching db/mariadb/arabterm.sql.gz is merged to arabterm's main, the cross-repo CI flow auto-opens a PR here with the regenerated db/arabterm.sql; merging that PR triggers the production DB import (see "Updating the Database" below). For the upstream dump-generation workflow (make init_mariadb, make migrate_to_mariadb, make dump), see arabterm's README.

MariaDB on Toolforge

Initial Setup

Ref: https://wikitech.wikimedia.org/wiki/Help:Toolforge/Database#User_databases

  • ssh toolforge and become wikitermbase
  • Find out your user in $HOME/replica.my.cnf
  • Create the database:
    • Open the SQL console: sql tools
    • Create the database: MariaDB [(none)]> CREATE DATABASE s55953__arabterm;

Updating the Database

DB imports are automated. The flow is:

  1. Update data in forzagreen/arabterm and merge to main. When db/mariadb/arabterm.sql.gz changes, arabterm's notify-wikitermbase.yml dispatches an event to this repo.
  2. wikitermbase's refresh-dump.yml runs make download_dump && make fix_dump and opens a PR titled chore: refresh DB dump from arabterm@<sha>.
  3. Review the diff to db/arabterm.sql and merge. CI's deploy-db job SSHs into the bastion and runs mariadb ... < db/arabterm.sql automatically.
  4. CI's wikidata-stats job then copies the new counts from /api/v1/stats to the Wikidata item Q133800945 (P4876 = terms, P2670 dictionary + P1114 = dictionaries), which the project page displays. It needs the WIKIDATA_USERNAME / WIKIDATA_PASSWORD secrets, a bot password with the "Edit existing pages" grant, stored in the wikidata environment (Settings → Environments), which only main may use. Preview the edit locally with uv run python backend/wikidata_stats.py --dry-run.

Manual triggers:

  • Re-run the dump regeneration: Actions tab → "Refresh DB dump from arabterm" → "Run workflow".

  • Re-sync Wikidata: Actions tab → "CI" → "Run workflow" on main (also redeploys the code). It edits nothing when the counts already match.

  • Re-import without a code change:

    ssh toolforge && become wikitermbase
    cd ~/wikitermbase
    mariadb --defaults-file=$HOME/replica.my.cnf -h tools.db.svc.wikimedia.cloud s55953__arabterm < db/arabterm.sql

Troubleshooting

All these issues are fixed by running make fix_dump

  • https://jira.mariadb.org/browse/MDEV-34183 drop the line /*!999999\- enable the sandbox mode */ or /*M!999999\- enable the sandbox mode */
  • ERROR 1273 (HY000) at line 25: Unknown collation: 'utf8mb4_uca1400_ai_ci', replace it with utf8mb4_unicode_520_ci

References

About

Standardise terminology used on Arabic Wikipedia and accelerate vocabulary translation

Topics

Resources

Stars

15 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages