Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Text Analysis Course

Digital Humanities Master. University of Bern.

Instructor: Elena Spadini.

Content

Date Topic Notes Corpus
18.09 Intro and regex TGG
25.09 Voyant Tools With Ursula Loosli (Unibe) TGG+
2.10 Lexical diversity TGG+
9.10 Word embeddings TGG+
16.10 Tokenization, lemmatization, and POS tagging TGG
23.10 Case study 1: the workflow Frankestein
30.10 Case study 1: categorising variants Frankestein
6.11 Network analysis TGG
13.11 Network analysis With José Luis Losada Palenzuela (Unibas) TGG+
20.11 Infoclio Conference (see below)
27.11 Data visualisations To be rescheduled (see below)
4.12 Case study 2: the workflow TGG
11.12 Case study 2: gender agency in dialogues TGG
18.12 Presentations

Calendar

Infoclio Conference 2026 On the 20th of November, we go to the https://www.infoclio.ch/en/programm-tagung2026. Please confirm to me you will be able to attend 10:00-12:00. If you want to stay for the whole day, you will have to pay the registration fee of 15 CHF. This is not mandatory.

Reschedule On the 27th of November, our lesson is cancelled. This lesson is rescheduled to … 23rd or 30th of November ?

Corpus

FITZ-TGG-25 or TGG = Francis Scott Fitzgerald, The Great Gatsby*, 1925 (~47'000 words)

CATH-ALL-23 = Willa Cather, A Lost Lady, 1923 (~37'500 words)

HEMI-TSAR-26 = Ernest Hemingway, The Sun Also Rises, 1926 (~67'500 words)

TGG+ = FITZ-TGG-25 + CATH-ALL-23 + HEMI-TSAR-26.

These texts are available via the Project Gutenberg, e.g. https://www.gutenberg.org/ebooks/64317. We will use the Plain Text format (see Download options).

Frankestein = materials from https://shelleygodwinarchive.org/.

Create a folder data in your computer, where you collect the texts to work with. A similar folder data is available here.

Assignment

For the assignment, you create your own analysis of one or more texts.

Two deliverables:

  • An oral presentation, during the last session
  • A Jupyter Notebook

Methods: you can choose between

  • word embeddings
  • network analysis
  • a mixed workflow including LLMs and traditional NLP approaches (lexical diversity, lemmatisation, POS tagging).

Corpus: one or more texts of your choice or ask for it.

Let's discuss it before. Write an email or ask for an appointment to briefly present your plan by November 20 at the latest. You should indicate the chosen corpus and method(s).

Using LLMs and AI assistants

You may use LLMs and AI assistants for this course and for the assignment under two conditions:

  • Document it. State which tools you used and for which steps.
  • Check it. You are responsible for every claim and every reference.

Tools and more

Markdown

XML

Text and code editor

Up-to-date browser

  • Recommended: Firefox or Chrome/Chromium.

Jupyter Notebooks

Regular expressions

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages